Mutual inductor defect detection method and system based on multi-layer and multi-mode fusion
Through the multi-layer multimodal fusion method, deep feature fusion of camera and lidar data is solved, and the problems of perspective differences and heterogeneity are achieved, high-precision and high-efficiency power equipment detection are achieved to meet the needs of real-time and complex scenarios.
Patent Information
- Application Number
- CN202510043226.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-27
AI Technical Summary
The existing single-stage multimodal fusion method is difficult to effectively deal with the difference in perspective and heterogeneity between camera and lidar data, resulting in limited detection effect and efficiency, especially in complex power scenarios, which is difficult to meet the requirements of real-time and efficientness.
The multi-layer multimodal fusion method is used to map the camera 2D image information data into the point cloud data of the lidar, and the multimodal data is fused before feature extraction and fusion generation, and deep feature fusion is used to use the detection backbone network and self-attention mechanism to generate, and a 3D detection box is generated through geometric and semantic consistency analysis, and the defects of the target transformer are finally detected.
It realizes efficient integration of camera and lidar data, improves the accuracy and efficiency of power equipment detection, can meet the requirements of real-time and efficient in complex scenarios, and reduces the risks and costs of manual inspections.
Smart Images

Figure CN120047390A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of transformer detection, and more specifically, to a method for detecting transformer defects based on multi-layer and multi-modal fusion. Background Art
[0002] In the power scenario, with the development of intelligent and unmanned inspection technologies, environmental perception technology has gradually become an important means to ensure the safety and efficiency of power equipment. Power inspection work is an important part of power equipment maintenance and fault prevention, and the inspection quality directly affects the stability and reliability of the power system. Traditional power inspections mostly rely on manual operations, with complex working environments and long inspection cycles, posing certain safety hazards, especially for the inspection of high-voltage transmission lines, substations, and wind power generation equipment. Therefore, in recent years, automated devices such as drones and robots have been widely used and combined with various sensors, such as visual sensors (cameras) and lidar, in order to better identify and locate the key components and potential defects of power equipment.
[0003] Visual sensors and lidar are the two most commonly used sensor types in power inspections, and they are significantly complementary in the power scenario. Cameras can capture high-definition image data, providing a high-resolution visual representation of the detailed information of the equipment, especially suitable for identifying fine features such as the appearance structure, surface damage, and corrosion marks of power equipment. However, due to the imaging principle of cameras, the accuracy of their depth information is poor, and usually only two-dimensional plane analysis can be performed through the image content, and they are not sensitive enough to the three-dimensional spatial position and distance information of power equipment. In addition, visual sensors are highly dependent on lighting conditions, and the imaging quality will significantly decline under lighting changes, shadows, strong light, or low-light environments at night, affecting the stability of the detection results.
[0004] As another important sensor, lidar can directly collect three-dimensional spatial information, providing accurate position and shape data of power equipment, and is almost unaffected by lighting conditions, with strong environmental adaptability. This feature makes it very suitable for applications in scenarios that require precise measurement, such as measuring the height of transmission lines, the inclination angle of utility poles, and the distance between equipment and the ground or other objects. In addition, the point cloud data generated by lidar can present the overall shape contour of the equipment, providing effective spatial geometric information. However, lidar data is relatively sparse in long-distance scenarios, and as the distance increases, the density of the point cloud rapidly decreases. Especially during the inspection of elevated lines, the point cloud of the target object is sparse, affecting the data resolution. At the same time, lidar cannot provide rich surface detail information such as color and texture, resulting in deficiencies in identifying surface damage or fine defects of the equipment.
[0005] To overcome the limitations of a single sensor, object detection technology based on multi-modal data fusion has become an effective means to improve the effectiveness of power inspection. By fusing the data of cameras and lidar, the high-resolution image information of cameras and the precise distance information of lidar can be utilized simultaneously, achieving complementary advantages between the two, and thus enabling the accurate identification and positioning of power equipment in various complex environments. Specifically, multi-modal data fusion can integrate visual and depth information, retaining both the detailed features such as the appearance and color of power equipment and having accurate three-dimensional spatial position data, which helps the inspection system to accurately detect and analyze the status of power equipment in different scenarios.
[0006] However, current multi-modal fusion research mainly focuses on single-stage fusion algorithms, mainly including early fusion, late fusion, or shallow feature fusion. Although these methods can improve the detection performance to a certain extent, there are still deficiencies. First, single-stage fusion algorithms cannot completely eliminate the perspective differences between camera and lidar data. The perspective misalignment problem caused by this difference affects the accurate registration of target features and reduces the fusion effect. Second, the heterogeneity of camera and lidar data is relatively strong, that is, camera data is two-dimensional high-resolution dense information, while lidar point clouds are sparse three-dimensional data. In the process of feature fusion, due to the different dimensions and structures of the data, it is difficult for existing single-stage methods to achieve deep feature fusion, resulting in inaccurate matching between image details and point cloud spatial information. Finally, when dealing with complex power scenarios, single-stage fusion methods often require high computing resources and have a long processing flow, making it difficult to meet the requirements of real-time and efficiency in the power inspection process.
[0007] Although existing single-stage multi-modal fusion methods have achieved multi-sensor data fusion to a certain extent, there are still some significant technical limitations in practical applications, which limit their detection effect and efficiency in complex scenarios. Specifically, these methods usually have difficulty effectively dealing with the perspective differences of multi-modal data, resulting in deviations in the spatial alignment of image and point cloud information. In addition, due to the heterogeneity of camera and lidar data in format and feature dimensions, traditional single-stage fusion methods often have difficulty achieving deep feature fusion, affecting the accuracy and integrity of detection. Moreover, the calculation process of single-stage fusion methods is relatively long, making it difficult to meet the efficiency requirements in power inspection tasks that require fast real-time processing. Generally speaking, these technical limitations pose great challenges in multi-modal fusion applications.
[0008] The first is the problem of perspective misalignment. Due to the different installation positions and working principles of cameras and lidar, the perspectives of the two are not exactly the same, making it difficult to directly match their data in space. Existing algorithms usually project and transform the point cloud to align it with the image information to alleviate the perspective difference. However, this transformation process will cause the blurring of depth information, resulting in a decrease in the registration accuracy between the image and the point cloud, thus affecting the detection effect and the accuracy of the final recognition.
[0009] The second is the difficulty of heterogeneous data fusion. The images obtained by cameras are high-resolution dense data, containing rich detail information such as colors and textures, which helps to identify tiny defects on the surface of the device; while lidar data is sparse three-dimensional point clouds, providing spatial information such as the position and contour of the device. Due to the differences in data formats and dimensions, it is difficult to directly fuse the features of the two. Especially in multi-modal feature matching, it is difficult to effectively retain the precise correspondence between the image texture information and the point cloud spatial structure, resulting in the loss or redundancy of the fused feature information. This information mismatch problem reduces the recognition accuracy of multi-modal fusion methods in the actual detection process.
[0010] The third is the low computational efficiency. Some existing single-stage fusion methods such as PointPainting and Frustum-PointNets adopt single-stage or decision-level fusion strategies, usually fusing in the initial stage of image preprocessing or point cloud processing. The processing flow is complex and lengthy, making it difficult to meet the real-time requirements in practical applications. In large-scale inspection tasks in the power scenario, especially in scenarios with strong real-time detection and feedback requirements, such methods are difficult to operate efficiently. In addition, due to the single-stage nature of their fusion methods, it is impossible to extract information from different data in stages, resulting in a waste of computing resources and an insufficient fusion of multi-modal information. Summary of the Invention
[0011] To address the above problems, the present invention proposes a method for detecting transformer defects based on multi-layer multi-modal fusion, including:
[0012] Mapping the camera 2D image information data for the detection of the target transformer into the point cloud data of the lidar for the detection of the target transformer to generate pre-fusion multi-modal data;
[0013] Performing feature extraction on the pre-fusion multi-modal data to obtain multi-modal features, and mapping the multi-modal features into spatial features;
[0014] Identifying the spatial features, obtaining the detection results of the camera and lidar for the transformer, and fusing the detection results to generate a 3D detection box;
[0015] Detecting the 3D detection box to detect the defects of the target transformer.
[0016] Optionally, adopt conical view area point cloud color smear coding to map the camera 2D image information data for the target instrument transformer detection to the point cloud data of the lidar for the target instrument transformer detection, including:
[0017] Using a 2D image detection algorithm, generate candidate detection boxes for the point cloud data of the lidar, perform local coding on the point cloud in each candidate detection box, and map the RGB color channel information of the camera 2D image information data to the encoded point cloud reflection intensity channel.
[0018] Optionally, based on the detection backbone network, perform feature extraction on the pre-fused multi-modal data to obtain multi-modal features, and map the multi-modal features to spatial features.
[0019] Optionally, the detection backbone network is a detection network based on PointPillars, and the detection network has a self-attention mechanism;
[0020] Capture the global information of the pre-fused multi-modal data through the self-attention mechanism to generate a pseudo-image feature map, and map the multi-modal features to spatial features based on the pseudo-image feature map.
[0021] Optionally, fuse the detection results to generate a 3D detection box, including:
[0022] For the detection results, generate 2D candidate boxes and 3D candidate boxes respectively, determine the semantic consistency of the 2D candidate boxes and 3D candidate boxes, match the semantic consistency, and generate a 3D detection box.
[0023] On the other hand, the present invention also proposes an instrument transformer defect detection system based on multi-layer multi-modal fusion, including:
[0024] A pre-fusion unit for mapping the camera 2D image information data for the target instrument transformer detection to the point cloud data of the lidar for the target instrument transformer detection to generate pre-fused multi-modal data;
[0025] A detection unit for performing feature extraction on the pre-fused multi-modal data to obtain multi-modal features, and mapping the multi-modal features to spatial features;
[0026] A fusion unit for identifying the spatial features, obtaining the detection results of the camera and the lidar for the instrument transformer, fusing the detection results, and generating a 3D detection box;
[0027] An identification unit for detecting the defects of the target instrument transformer by detecting the 3D detection box.
[0028] Optionally, the front fusion unit uses conical view area point cloud color smear coding to map the camera 2D image information data for target current transformer detection into the point cloud data of the lidar for target current transformer detection, including:
[0029] Using a 2D image detection algorithm, for the point cloud data of the lidar, generate candidate detection boxes, perform local coding on the point cloud in each candidate detection box, and map the RGB color channel information of the camera 2D image information data into the encoded point cloud reflection intensity channel.
[0030] Optionally, the detection unit performs feature extraction on the front fusion multimodal data based on a detection backbone network, obtains multimodal features, and maps the multimodal features into spatial features.
[0031] Optionally, the detection backbone network is a detection network based on PointPillars, and the detection network has a self-attention mechanism;
[0032] Capture the global information of the front fusion multimodal data through the self-attention mechanism to generate a pseudo-image feature map, and map the multimodal features into spatial features based on the pseudo-image feature map.
[0033] Optionally, fuse the detection results to generate a 3D detection box, including:
[0034] For the detection results, generate 2D candidate boxes and 3D candidate boxes respectively, determine the semantic consistency of the 2D candidate boxes and 3D candidate boxes, and match the semantic consistency to generate a 3D detection box.
[0035] On the other hand, the present invention also provides a computing device, including: one or more processors;
[0036] The processor is used to execute one or more programs;
[0037] When the one or more programs are executed by the one or more processors, the method as described above is implemented.
[0038] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, the method as described above is implemented.
[0039] Compared with the prior art, the beneficial effects of the present invention are:
[0040] The present invention provides a method for detecting defects of transformers based on multi-layer multi-modal fusion, including: mapping the camera 2D image information data for detecting the target transformer into the point cloud data of the lidar for detecting the target transformer to generate pre-fusion multi-modal data; extracting features from the pre-fusion multi-modal data to obtain multi-modal features, and mapping the multi-modal features into spatial features; identifying the spatial features, obtaining the detection results of the camera and the lidar for the transformer, and fusing the detection results to generate a 3D detection frame; detecting the 3D detection frame to detect the defects of the target transformer. The present invention can provide a high-precision and high-efficiency detection solution for the automatic inspection of power equipment, reduce the risks and costs of manual inspection, and meet the requirements of real-time performance and reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a flowchart of the method of the present invention;
[0042] Figure 2 is a structural diagram of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] Now, exemplary embodiments of the present invention will be described with reference to the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely, and to fully convey the scope of the present invention to those skilled in the art. The terms in the exemplary embodiments shown in the drawings are not intended to limit the present invention. In the drawings, the same unit / element is denoted by the same reference numeral.
[0044] Unless otherwise specified, the terms (including scientific and technical terms) used herein have the ordinary meaning understood by those skilled in the art. In addition, it can be understood that the terms defined in the commonly used dictionary should be understood to have a meaning consistent with the context of the relevant field, and should not be understood as idealized or overly formal meanings.
[0045] Embodiment 1:
[0046] The present invention proposes a method for detecting defects of transformers based on multi-layer multi-modal fusion, as Figure 1 shown, including:
[0047] Step 1, mapping the camera 2D image information data for detecting the target transformer into the point cloud data of the lidar for detecting the target transformer to generate pre-fusion multi-modal data;
[0048] Step 2, extracting features from the pre-fusion multi-modal data to obtain multi-modal features, and mapping the multi-modal features into spatial features;
[0049] Step 3: Identify the spatial features, obtain the detection results of the camera and lidar for the instrument transformer, and fuse the detection results to generate a 3D detection box;
[0050] Step 4: Detect the 3D detection box to detect the defects of the target instrument transformer.
[0051] Among them, using the cone view area point cloud color smear coding, the camera 2D image information data for the detection of the target instrument transformer is mapped to the point cloud data of the lidar for the detection of the target instrument transformer, including:
[0052] Using the 2D image detection algorithm, for the point cloud data of the lidar, generate candidate detection boxes, perform local coding on the point cloud in each candidate detection box, and map the RGB color channel information of the camera 2D image information data to the coded point cloud reflection intensity channel.
[0053] Among them, based on the detection backbone network, the pre-fused multimodal data is subjected to feature extraction to obtain multimodal features, and the multimodal features are mapped to spatial features.
[0054] Among them, the detection backbone network is a detection network based on PointPillars, and the detection network has a self-attention mechanism;
[0055] Through the self-attention mechanism, capture the global information of the pre-fused multimodal data to generate a pseudo-image feature map, and map the multimodal features to spatial features based on the pseudo-image feature map.
[0056] Among them, fusing the detection results to generate a 3D detection box includes:
[0057] For the detection results, generate 2D candidate boxes and 3D candidate boxes respectively, determine the semantic consistency of the 2D candidate boxes and 3D candidate boxes, and match the semantic consistency to generate a 3D detection box.
[0058] The following further illustrates the present invention in conjunction with specific embodiments:
[0059] Embodiment of the present invention: It mainly includes three steps: the pre-fusion stage, the detection backbone network stage, and the post-fusion stage, which are specifically as follows:
[0060] First, in the pre-fusion stage, using the cone view area point cloud color smear coding method, map the 2D image information collected by the camera to the point cloud data of the lidar, as Figure 2As shown. Specifically, the system uses a 2D image detection algorithm (such as YOLO) to generate candidate detection boxes for power equipment, and then locally encodes the point cloud in each candidate detection box, mapping the RGB color channel information into the reflection intensity channel of the point cloud, so as to provide richer feature information for subsequent 3D detection. Among them, the problem of alignment error between the image and the point cloud caused by the perspective difference is effectively alleviated through local area encoding, enabling accurate mapping of multi-modal data in the pre-fusion stage.
[0061] In the detection backbone network stage, the multi-modal data after pre-fusion processing is further subjected to feature extraction. For this purpose, the system adds a self-attention mechanism to the detection backbone network to improve the ability to extract spatial features. The multi-modal data is used as input and enters a detection network based on PointPillars. By means of the self-attention mechanism, global context information is captured, enhancing the understanding and utilization of the point cloud spatial information, and enabling the fused features to have stronger representation capabilities. This process generates a pseudo-image feature map, mapping the multi-modal features into spatial features, and further improving the accuracy of device recognition.
[0062] Finally, in the post-fusion stage, the detection results of the camera and the lidar are fused through geometric and semantic consistency analysis to generate more accurate 3D detection boxes. At this stage, the system matches the 2D and 3D candidate boxes through geometric consistency analysis (such as IoU) and semantic consistency analysis (based on class information) to ensure the accuracy of the fusion result. Before the final output, the system removes redundant boxes and false alarm boxes through non-maximum suppression (NMS) to ensure the accuracy and reliability of the detection boxes.
[0063] The present invention realizes the gradual fusion of camera and lidar data to improve the accuracy and efficiency of power equipment detection, effectively overcoming the deficiencies of existing single-stage fusion technologies. In traditional power inspection tasks, due to the inconsistent perspectives of the camera and the lidar, the image and point cloud data cannot be accurately matched, affecting the detection accuracy. In addition, existing methods are difficult to achieve deep fusion of image and point cloud features, with low computational efficiency and insufficient to meet real-time requirements.
[0064] To solve these problems, the present invention introduces a color smear coding method for the cone visual area in the pre-fusion stage, maps the 2D image information of the camera into the point cloud data, ensures the consistency of geometric and semantic information, and effectively alleviates the problem of perspective misalignment. Specifically, the system first uses a 2D image detection algorithm to generate candidate detection boxes for power equipment, and then maps the RGB information to the reflection intensity channel of the point cloud, thereby enhancing the integrity and accuracy of 3D detection features. In the detection backbone network stage, the fused multi-modal data is processed through a self-attention mechanism, enabling the network to capture context information and achieve deep-level fusion of features, thus more accurately identifying power equipment, especially significantly improving the recognition effect in long-distance or sparse scenarios.
[0065] In the post-fusion stage, the present invention further fuses the detection results of the camera and lidar through geometric and semantic consistency analysis, and removes redundant and false detection boxes through non-maximum suppression (NMS). This not only optimizes the detection accuracy but also significantly improves the computational efficiency, enabling the system to meet the real-time requirements in power inspection tasks.
[0066] The multi-level fusion framework of the present invention successfully solves the problems of perspective misalignment and heterogeneous data fusion in the prior art by gradually optimizing the data features at each stage, and significantly improves the computational efficiency of detection. In practical applications, the present invention can provide a high-precision and high-efficiency detection solution for the automated inspection of power equipment, reduce the risks and costs of manual inspection, and meet the real-time and reliability requirements.
[0067] Embodiment 2:
[0068] The present invention also proposes an instrument transformer defect detection system 200 based on multi-layer multi-modal fusion, as Figure 2 shown, including:
[0069] A pre-fusion unit 201 for mapping the camera 2D image information data for the detection of the target instrument transformer into the point cloud data of the lidar for the detection of the target instrument transformer to generate pre-fusion multi-modal data;
[0070] A detection unit 202 for extracting features from the pre-fusion multi-modal data, obtaining multi-modal features, and mapping the multi-modal features into spatial features;
[0071] A fusion unit 203 for identifying the spatial features, obtaining the detection results of the camera and lidar for the instrument transformer, and fusing the detection results to generate a 3D detection box;
[0072] An identification unit 204 for detecting the defects of the target instrument transformer by detecting the 3D detection box.
[0073] Among them, the pre-fusion unit 201 adopts conical viewing area point cloud color smear coding to map the camera 2D image information data for the detection of the target current transformer into the point cloud data of the lidar for the detection of the target current transformer, including:
[0074] Using a 2D image detection algorithm, for the point cloud data of the lidar, generate candidate detection boxes, perform local coding on the point cloud in each candidate detection box, and map the RGB color channel information of the camera 2D image information data into the encoded point cloud reflection intensity channel.
[0075] Among them, the detection unit 202 performs feature extraction on the pre-fused multi-modal data based on the detection backbone network, obtains multi-modal features, and maps the multi-modal features into spatial features.
[0076] Among them, the detection backbone network is a detection network based on PointPillars, and the detection network has a self-attention mechanism;
[0077] Capture the global information of the pre-fused multi-modal data through the self-attention mechanism to generate a pseudo-image feature map, and map the multi-modal features into spatial features based on the pseudo-image feature map.
[0078] Among them, fusing the detection results to generate a 3D detection box includes:
[0079] For the detection results, generate 2D candidate boxes and 3D candidate boxes respectively, determine the semantic consistency of the 2D candidate boxes and 3D candidate boxes, and match the semantic consistency to generate a 3D detection box.
[0080] The present invention can provide a high-precision and high-efficiency detection solution for the automated inspection of power equipment, reduce the risks and costs of manual inspection, and meet the requirements of real-time performance and reliability.
[0081] Example 3:
[0082] Based on the same inventive concept, the present invention also provides a computer device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of the method in the above embodiments.
[0083] Embodiment 4:
[0084] Based on the same inventive concept, the present invention also provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The one or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the steps of the method in the above embodiments.
[0085] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code. The solutions in the embodiments of the present invention can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript, etc.
[0086] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0087] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0089] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.
[0090] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.
Claims
1. A transformer defect detection method based on multi-layer multi-modal fusion, characterized in that: include: Mapping the camera 2D image information data for target mutual inductor detection to the point cloud data of the laser radar for target mutual inductor detection to generate pre-fusion multimodal data; Performing feature extraction on the pre-fused multimodal data to obtain multimodal features, and mapping the multimodal features into spatial features; Identify the spatial features, obtain detection results of the mutual inductor by the camera and the laser radar, fuse the detection results, and generate a 3D detection frame; The defects of the target transformer are detected by detecting the 3D detection frame.
2. The method according to claim 1, characterized in that The cone view point cloud color smear coding is used to map the camera 2D image information data for the target mutual inductor detection to the laser radar point cloud data for the target mutual inductor detection, including: Using the 2D image detection algorithm, candidate detection frames are generated for the lidar point cloud data, the point cloud in each candidate detection frame is locally encoded, and the RGB color channel information of the camera's 2D image information data is mapped to the encoded point cloud reflection intensity channel.
3. The method according to claim 1, characterized in that Based on the detection backbone network, feature extraction is performed on the pre-fused multimodal data to obtain multimodal features, and the multimodal features are mapped into spatial features.
4. The method according to claim 3, characterized in that The detection backbone network is a PointPillars-based detection network, and the detection network has a self-attention mechanism; The global information of the pre-fused multimodal data is captured through a self-attention mechanism to generate a pseudo image feature map, and the multimodal features are mapped into spatial features based on the pseudo image feature map.
5. The method according to claim 1, characterized in that: The fusing the detection results to generate a 3D detection frame includes: Based on the detection results, a 2D candidate box and a 3D candidate box are generated respectively, and the semantic consistency of the 2D candidate box and the 3D candidate box is determined, and the semantic consistency is matched to generate a 3D detection box.
6. A transformer defect detection system based on multi-layer multi-modal fusion, characterized in that: include: A front fusion unit, used to map the camera 2D image information data for target mutual inductor detection to the point cloud data of the laser radar for target mutual inductor detection, and generate front fusion multimodal data; A detection unit, configured to perform feature extraction on the pre-fused multimodal data, obtain multimodal features, and map the multimodal features into spatial features; A fusion unit, used to identify the spatial features, obtain detection results of the camera and the laser radar for the mutual inductor, fuse the detection results, and generate a 3D detection frame; An identification unit is used to detect defects of the target mutual inductor by detecting the 3D detection frame.
7. The system according to claim 6, characterized in that The front fusion unit uses cone view point cloud color smear coding to map the camera 2D image information data for target mutual inductor detection to the laser radar point cloud data for target mutual inductor detection, including: Using the 2D image detection algorithm, candidate detection frames are generated for the lidar point cloud data, the point cloud in each candidate detection frame is locally encoded, and the RGB color channel information of the camera's 2D image information data is mapped to the encoded point cloud reflection intensity channel.
8. The system according to claim 6, characterized in that The detection unit performs feature extraction on the pre-fusion multimodal data based on the detection backbone network, obtains multimodal features, and maps the multimodal features into spatial features.
9. The system according to claim 8, characterized in that The detection backbone network is a PointPillars-based detection network, and the detection network has a self-attention mechanism; The global information of the pre-fused multimodal data is captured through a self-attention mechanism to generate a pseudo image feature map, and the multimodal features are mapped into spatial features based on the pseudo image feature map.
10. The system according to claim 6, characterized in that The fusing the detection results to generate a 3D detection frame includes: Based on the detection results, a 2D candidate box and a 3D candidate box are generated respectively, and the semantic consistency of the 2D candidate box and the 3D candidate box is determined, and the semantic consistency is matched to generate a 3D detection box.
11. A computer device, characterized in that: include: one or more processors; a processor for executing one or more programs; When the one or more programs are executed by the one or more processors, the method according to any one of claims 1 to 5 is implemented.
12. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed, the method according to any one of claims 1 to 5 is implemented.