Automatic disassembling method, system and equipment for main transformer equipment and medium
The main transformer image is multimodally fused through the chessboard method and convolutional neural network model, and the prototype mask is generated by combining the attention instance segmentation model, which solves the problems of position mapping deviation and high cost in the automatic disassembly of the main transformer and realizes an efficient and safe disassembly process.
Patent Information
- Application Number
- CN202510571052.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-26
AI Technical Summary
The existing main transformer automatic disassembly technology has problems such as component position mapping deviation caused by differences in multimodal sensor perspectives and timestamps, high computational complexity and difficulty in real-time disassembly, poor environmental robustness and high cost.
The chessboard method is used to align image data, combined with the convolutional neural network model for multimodal fusion, and the attention instance segmentation model is used to extract features and generate prototype masks for automatic disassembly.
The accuracy and efficiency of main transformer disassembly are improved, labor costs are reduced, environmental pollution risks are reduced, and the environmental adaptability and robustness of the system are enhanced.
Smart Images

Figure CN120707811A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of transformers, and in particular to an automatic disassembly method, method, equipment and medium for main transformer equipment. Background Art
[0002] As the power industry continues to develop, main transformers, as core equipment in power systems, are increasing in number as the scale of power grids expands. After long-term operation, main transformers are often retired due to aging, technological upgrades, and other factors. Proper disassembly and processing of these retired main transformers not only recovers high-value materials such as copper and silicon steel sheets, achieving a metal recovery rate exceeding 90% and effectively reducing mineral resource depletion, but also prevents environmental pollution from toxic substances such as polychlorinated biphenyls (PCBs), meeting the strategic requirements of sustainable industrial development. With the rapid development of automation technology, automated disassembly of main transformers has become a key development direction in the industry. Automated disassembly improves disassembly efficiency, reduces labor costs, and avoids the risk of insulating oil leaks associated with manual disassembly, significantly enhancing operator safety and protecting the soil and water environment. Therefore, research on automated disassembly technology for main transformers is of great practical significance.
[0003] However, current automated main transformer disassembly technology still faces numerous challenges in component identification and decomposition. Among existing main transformer component identification methods, multimodal sensor fusion, deep learning-driven visual algorithms, and physical property analysis each have their applications, but all suffer from significant drawbacks. From an environmental adaptability perspective, oily, rusty, or reflective surfaces can severely impact visual recognition accuracy, leading to deviations in the identification of key components. Regarding multimodal collaboration, spatiotemporal synchronization is difficult when visual, thermal, and mechanical data are integrated, and deviations between robotic arm operations and visual positioning can easily lead to misoperations. Regarding real-time performance and cost, high-precision point cloud processing relies on expensive GPUs, making it difficult for small and medium-sized recycling companies to bear the associated costs. These issues severely restrict the promotion and development of automated main transformer disassembly technology. Innovative methods are urgently needed to improve the environmental robustness, collaborative efficiency, and real-time performance of automated disassembly systems, reduce costs, and drive the automated main transformer disassembly industry toward greater efficiency, environmental friendliness, and economics. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides an automatic disassembly method for main transformer equipment to solve the problems of component position mapping deviation caused by differences in multimodal sensor viewing angles and timestamps during automatic disassembly of the main transformer, high computational complexity and difficulty in real-time disassembly, as well as poor environmental robustness of existing algorithms, low multimodal collaboration efficiency and high cost.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for automatically disassembling a main transformer device, comprising:
[0008] Obtain the original image of the main transformer equipment;
[0009] Performing data alignment on the original image of the main transformer device using a chessboard method to obtain an image of the main transformer device;
[0010] Performing multimodal fusion on the main transformer device image through a convolutional neural network model to obtain a main transformer device fusion image;
[0011] Inputting the main transformer device fusion image into the attention instance segmentation model for feature extraction to obtain the main transformer device fusion image features;
[0012] Based on the fused image features of the main transformer device, the main transformer device is automatically disassembled.
[0013] As a preferred solution of the automatic disassembly method of the main transformer device of the present invention, wherein: using the chessboard method to align the data of the original image of the main transformer device to obtain the main transformer device image, the method includes:
[0014] Acquire different coordinate system data of the original image of the main transformer device in combination with shooting device parameters;
[0015] Coordinate conversion is performed according to the different coordinate system data to obtain the main transformer equipment image.
[0016] The beneficial effect of this preferred technical solution is: by combining the shooting equipment parameters to obtain data in different coordinate systems and perform coordinate conversion, it can effectively solve the component position mapping deviation problem caused by the difference in multimodal sensor viewing angle and timestamp, and provide accurate image data for subsequent image fusion and disassembly operations.
[0017] As a preferred solution of the automatic disassembly method of the main transformer equipment of the present invention, wherein: performing multimodal fusion on the main transformer equipment image through a convolutional neural network model to obtain a fused image of the main transformer equipment includes:
[0018] The convolutional neural network model includes an input layer, a gradient residual block layer, a convolution layer and an output layer;
[0019] Performing preliminary feature extraction on the main transformer device image through the input layer to obtain a feature map;
[0020] The feature map is input into the gradient residual block layer and the convolution layer to further process the features, and the main transformer device fusion image is output through the output layer.
[0021] The beneficial effects of this preferred technical solution are: by gradually processing the image features using the various layers of the convolutional neural network model, it is possible to fully integrate different modal information, enhance the expressiveness of image features, and improve the quality of the fused image of the main transformer equipment.
[0022] As a preferred solution of the automatic disassembly method of the main transformer equipment of the present invention, it further includes:
[0023] Designing a loss function to train the convolutional neural network model and optimize the convolutional neural network model; inputting the original image of the main transformer device into the trained convolutional neural network model to obtain a fused image of the main transformer device;
[0024] The fused image of the main transformer equipment is compared with the original image of the main transformer equipment, and manual evaluation is performed to determine whether the fused image of the main transformer equipment includes visible light modalities and infrared light modalities.
[0025] As a preferred solution of the automatic disassembly method of the main transformer device of the present invention, the main transformer device fusion image is input into the attention instance segmentation model for feature extraction to obtain the main transformer device fusion image features, including:
[0026] The attention instance segmentation model uses hierarchical and progressive convolutional layers to obtain local features of the fused image of the main transformer device, and then further extracts low-dimensional features;
[0027] The attention mechanism is used to obtain the target area of the main transformer device fusion image and obtain the main transformer device fusion image features.
[0028] As a preferred solution of the automatic disassembly method of the main transformer device of the present invention, the main transformer device is automatically disassembled based on the fusion image features of the main transformer device, including:
[0029] Extract features of the main transformer device fusion image through a feature pyramid to generate a prototype mask of the main transformer device;
[0030] The main transformer device is identified and disassembled according to the main transformer device prototype mask.
[0031] The beneficial effects of this preferred technical solution are: by extracting features through the feature pyramid to generate a prototype mask of the main transformer equipment, the various components of the main transformer can be identified more accurately, and then precise disassembly can be performed based on the mask, thereby improving the efficiency and accuracy of disassembly.
[0032] As a preferred solution of the automatic disassembly method of the main transformer device of the present invention, a prototype mask of the main transformer device is generated, which is expressed by the formula:
[0033]
[0034] Among them, M instance is the instance mask, c i is the i-th coefficient, which represents the contribution weight of the prototype to the current instance, P i is the i-th prototype mask, and σ is the sigmoid function.
[0035] In a second aspect, the present invention provides an automatic disassembly system for a main transformer device, comprising: a data acquisition module for acquiring an original image of the main transformer device;
[0036] A preprocessing module, configured to perform data alignment on the original image of the main transformer device using a chessboard method to obtain an image of the main transformer device;
[0037] A feature extraction module is used to perform multimodal fusion on the main transformer device image through a convolutional neural network model to obtain a main transformer device fusion image; the main transformer device fusion image is input into an attention instance segmentation model for feature extraction to obtain main transformer device fusion image features;
[0038] The identification and disassembly module is used to automatically disassemble the main transformer device based on the fusion image features of the main transformer device.
[0039] In a third aspect, the present invention provides an electronic device, comprising:
[0040] A memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the automatic disassembly method of the main transformer device.
[0041] In a fourth aspect, the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the method for automatically disassembling the main transformer device.
[0042] Compared with the existing technology, the present invention has the following advantages: It utilizes a checkerboard method to align raw image data and combines it with camera parameters to transform coordinates, thus resolving position mapping deviations caused by multimodal sensors and ensuring accurate image data. The convolutional neural network model is trained and optimized to fuse information from different modalities. Manual judgment ensures that the fused image incorporates both visible and infrared light features, enhancing environmental adaptability and robustness. The attention instance segmentation model accurately extracts key features through layered convolution and attention mechanisms. Prototype masks are generated based on a feature pyramid to accurately identify and disassemble components, improving disassembly accuracy and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0044] Figure 1 This is a schematic diagram of the framework of each module of the automatic disassembly method of the main transformer equipment according to one embodiment of the present invention.
[0045] Figure 2 A schematic structural diagram of an attention mechanism for an automatic disassembly method of a main transformer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0046] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0047] Example 1, with reference to Figure 1 , as one embodiment of the present invention, provides an automatic disassembly method of a main transformer device, comprising: S100: acquiring an original image of the main transformer device;
[0048] S102: using a chessboard method to align the data of the original image of the main transformer device to obtain an image of the main transformer device;
[0049] S104: performing multimodal fusion on the main transformer device image through a convolutional neural network model to obtain a fused image of the main transformer device;
[0050] S106: Input the fusion image of the main transformer device into the attention instance segmentation model for feature extraction to obtain the fusion image features of the main transformer device; based on the fusion image features of the main transformer device, automatically disassemble the main transformer device.
[0051] It should be noted that manual disassembly of the main transformer equipment is inefficient and consumes a lot of manpower and time. The insulating oil of the main transformer may contain toxic substances such as polychlorinated biphenyls, and manual disassembly is prone to leakage, causing serious soil and water pollution. In the exploration of automated disassembly, although multimodal sensors are used to detect structures, differences in perspective and timestamps lead to deviations in component position mapping, and multimodal data are separately subjected to deep learning model inference and then fused, which is computationally complex and difficult to meet real-time disassembly requirements. Therefore, the present invention uses the chessboard method to align the data of the original image of the main transformer equipment, unify the coordinates of the images collected by multiple cameras, and eliminate perspective differences; then, the image is multimodally fused through a convolutional neural network model, combining the advantages of visible light and infrared light images to enhance feature expression. The fused image is input into the attention instance segmentation model to extract features, accurately locate components, and automatically disassemble the main transformer equipment based on the features.
[0052] This method improves the automation level of main transformer disassembly, increasing disassembly efficiency and reducing labor costs. It also effectively prevents insulating oil leakage and reduces environmental pollution risks. Furthermore, multimodal data alignment and fusion technology enhances adaptability to complex environments, improves component identification accuracy, and ensures a safe and accurate disassembly process.
[0053] Example 2, reference Figure 1-Figure 2 , which is an embodiment of the present invention, provides an automatic disassembly method for a main transformer device based on the above embodiment.
[0054] In the embodiment of the present invention, in the above step S100, an infrared camera and a visible light camera are used to obtain an infrared image and a visible light image of the main transformer device, that is, an original image of the main transformer device, and the infrared image and the visible light image of the main transformer device are preprocessed;
[0055] In an optional embodiment, the preprocessing includes image data denoising and upsampling. Specifically, images with large position differences and large noise in the infrared image and the visible light image are deleted to ensure the high quality of the training set; since the resolution of the infrared camera and the visible light camera is different, the infrared image is upsampled to keep the resolution of the infrared image and the visible light image the same.
[0056] In the embodiment of the present invention, in the above step S102, the original image of the main transformer device is aligned using the chessboard method to obtain the main transformer device image, including steps A1-A2;
[0057] A1: Combine the shooting equipment parameters to obtain different coordinate system data of the original image of the main transformer equipment;
[0058] A2: Perform coordinate conversion based on data from different coordinate systems to obtain the image of the main transformer equipment.
[0059] The chessboard method includes four related coordinate systems: pixel coordinate system, image coordinate system, camera coordinate system and world coordinate system; let the world coordinate be (X w ,Y w ,Z w ); the camera coordinate is (X c ,Y c ,Z c ); image coordinates are (x, y); pixel coordinates are (u, v);
[0060] In the embodiment of the present invention, the camera intrinsic parameter matrix is assumed to be:
[0061]
[0062] Among them, f x is the focal length of the camera on the x-axis, f y is the focal length of the camera on the y-axis, c x is the x-axis coordinate of the optical center, c y is the y-axis coordinate of the optical center;
[0063] For example, in combination with the shooting device parameters, the conversion formula between pixel coordinates and image coordinates is:
[0064]
[0065] It should be noted that the purpose of converting pixel coordinates to image coordinates is to remove the "barrel distortion" or "pincushion distortion" caused by the shape of the lens, where the light is bent more at the edge than at the center.
[0066] For example, combined with the depth value z c , the image coordinate system is converted to the camera coordinate system as follows:
[0067]
[0068] For example, let R be a 3×3 rotation matrix from the world coordinate system to the camera coordinate system, and t be a 3×1 translation vector from the world coordinate system to the camera coordinate system; then the conversion from the camera coordinate system to the world coordinate system is expressed as:
[0069]
[0070] It should be noted that the high resolution of visible light can capture surface texture, color, and geometric details, making it suitable for identifying components with obvious appearance features, such as sleeves and heat sinks. Infrared light can sense the temperature distribution of the equipment and identify possible abnormal hot spots, such as local overheating of the winding or oil leaks. It can also penetrate surface oil stains and rust interference to reveal hidden structures such as internal connection points. By combining the detailed information of visible light with the thermal characteristics of infrared, component recognition in the dual dimensions of "appearance + function" is achieved. The system can perceive device information more comprehensively and improve recognition accuracy. In addition, visible light is easily interfered by surface stains, resulting in errors such as misjudgment of silicon steel sheet adhesion due to edge detection failure. Infrared light can penetrate surface contamination through temperature distribution and accurately locate the internal core or winding position, so the fused data can avoid cutting into hidden support structures. Through the above conversion, the present invention unifies the coordinate values of the infrared camera and the visible light camera, providing high-quality image data for subsequent image fusion.
[0071] In an optional implementation, data alignment can be performed using a feature matching method to obtain an image of the main transformer device. Specifically, feature points such as scale-invariant feature transformation and accelerated robust features are first extracted from the infrared image and the visible light image, respectively. These feature points are then described using feature descriptors, and the feature points are matched by calculating the distance between the feature descriptors. Finally, a transformation matrix is calculated based on the matched feature point pairs, and one of the images is transformed, thereby achieving data alignment of the two images.
[0072] In another optional implementation, data alignment can be performed using the mutual information method to obtain images of the main transformer equipment. Specifically, mutual information is a measure of the statistical correlation between two random variables. The mutual information between the infrared image and the visible light image is calculated. By continuously adjusting the position and orientation of the images, the mutual information is maximized. The corresponding transformation parameters at this point are the optimal alignment parameters. Data alignment can be achieved by transforming the images according to these parameters.
[0073] In the embodiment of the present invention, in the above step S104, multimodal fusion of the main transformer device image is performed through the convolutional neural network model to obtain the main transformer device fusion image, which includes sub-steps B1-B4:
[0074] B1: The convolutional neural network model includes an input layer, a gradient residual block layer, a convolution layer, and an output layer. The input layer performs preliminary feature extraction on the main transformer device image to obtain a feature map. The feature map is input into the gradient residual block layer and the convolution layer for further feature processing, and the output layer outputs the main transformer device fusion image.
[0075] B2: Design a loss function to train the convolutional neural network model and optimize the convolutional neural network model.
[0076] B3: Input the original image of the main transformer equipment into the trained convolutional neural network model to obtain a fused image of the main transformer equipment; compare the fused image of the main transformer equipment with the original image of the main transformer equipment, and manually judge whether the fused image of the main transformer equipment includes visible light mode and infrared light mode.
[0077] In an embodiment of the present invention, the input layer is composed of a point convolution layer, which mainly performs preliminary extraction of image features of the two modalities and inputs the feature map into the gradient residual block layer. The gradient residual block layer replaces the original convolution layer with a lightweight residual block to reduce the number of model parameters and solve the problem of gradient disappearance or gradient explosion.
[0078] In the embodiment of the present invention, the loss function is mainly the content loss L content , the formula is:
[0079] L content =L int +αL texture
[0080] Among them, L int is the strength loss, L texture represents texture loss, and α represents a balance between the two losses.
[0081] Specifically, the intensity loss is the loss that constrains the pixel-level difference between infrared and visible light images, and the formula is:
[0082]
[0083] Where H and W are the height and width of the image, respectively, ||·||1 represents l1-norm and max(·) represents the element-wise maximum selection, I ir is the infrared image, I vi is a visible light image, I f Fuse images for the main transformer device.
[0084] Specifically, the texture loss formula is:
[0085]
[0086] Where H and W are the height and width of the image, respectively, ||·||1 represents l1-norm and max(·) represents the element-wise maximum selection, I ir is the infrared image, I vi is a visible light image, I f Fuse images for the main transformer device.
[0087] In an optional embodiment, the loss function of the convolutional neural network model can be a structural similarity index loss, etc.; this loss function measures the similarity of images from three aspects: brightness, contrast, and structure, and can better reflect the human eye's perception of image quality. Compared with the loss function that only considers pixel-level differences, it can make the generated fused image closer to the real scene in visual effect, thereby enhancing the visual quality of the fused image.
[0088] In an optional embodiment, a multimodal fusion of the main transformer equipment image can be performed using a wavelet transform-based method to obtain a fused image of the main transformer equipment. Specifically, the infrared image and the visible light image are first subjected to wavelet transforms respectively and decomposed into sub-band coefficients of different frequencies. The low-frequency sub-band reflects the general outline of the image, and the high-frequency sub-band contains the detailed information of the image. For the low-frequency sub-band, a weighted average method can be used for fusion, and the weight is determined based on the energy or variance of the low-frequency sub-bands of the two modal images. For the high-frequency sub-band, by comparing the absolute values of the high-frequency sub-band coefficients of the different modal images, the coefficient with the larger absolute value is selected as the fused high-frequency sub-band coefficient. Finally, the fused wavelet coefficients are subjected to an inverse wavelet transform to obtain a fused image of the main transformer equipment.
[0089] It should be noted that the present invention utilizes the point convolution layer of the convolutional neural network model input layer to efficiently and preliminarily extract infrared and visible light image features. The gradient residual block layer replaces the convolution layer with a lightweight residual block, which reduces the number of model parameters, solves the problem of gradient disappearance or explosion, ensures training stability and improves performance. The loss function combines intensity loss and texture loss, balances the relationship between the two through parameters, constrains image fusion from the two aspects of pixel-level differences and texture features, and ensures the quality of the fused image. Finally, manual evaluation is performed to determine whether the fused image contains two modes, further ensuring the validity of the fused image from a subjective perspective, and providing accurate and comprehensive image information for the subsequent automatic disassembly of the main transformer.
[0090] In the embodiment of the present invention, in the above step S106, the attention instance segmentation model is used to obtain the fusion image features of the main transformer device, and the automatic disassembly of the main transformer device includes sub-steps C1-C3:
[0091] C1: The attention instance segmentation model uses hierarchical and progressive convolutional layers to obtain local features of the fused image of the main transformer device, and then further extracts low-dimensional features;
[0092] C2: Use the attention mechanism to obtain the target area of the main transformer device fusion image and obtain the main transformer device fusion image features.
[0093] C3: The main transformer equipment is fused with image features through feature pyramid for feature extraction to generate a prototype mask of the main transformer equipment; the main transformer equipment is identified and disassembled based on the prototype mask of the main transformer equipment.
[0094] In the embodiment of the present invention, the attention instance segmentation model adopts the YOLACT algorithm (You Only Look At Coefficients) based on the attention mechanism;
[0095] In the embodiment of the present invention, the feature extraction using the attention instance segmentation model mainly includes:
[0096] ① Use hierarchical and progressive convolutional layers to extract low-level features of the image. These convolutional layers mainly use small-sized convolution kernels to scan the image, capturing local features such as edges and textures. Then, after the activation function, the pooling layer downsamples the features to retain the significant features.
[0097] ② Use larger convolution kernels to further extract low-dimensional feature maps. This process mainly increases the receptive field of the model and further extracts the feature maps.
[0098] ③Use the convolutional attention mechanism to make the model focus more on the target area and suppress background interference. The structure of the convolutional attention mechanism is as follows: Figure 2 shown.
[0099] In the embodiment of the present invention, a prototype mask of the main transformer device is generated, which is expressed by the formula:
[0100]
[0101] Among them, M instance is the instance mask, c i is the i-th coefficient, which represents the contribution weight of the prototype to the current instance, P i is the i-th prototype mask, and σ is the sigmoid function.
[0102] In an optional implementation, a region-based convolutional neural network with masking is used to extract features from the fused image of the main transformer device through backbone networks such as residual networks and visual geometry group networks, thereby obtaining a feature map with rich semantic information. Next, a region proposal network is used to generate a series of candidate regions that may contain main transformer components. Region of interest pooling is then performed on these candidate regions, mapping candidate regions of varying sizes into fixed-size feature vectors. A fully connected layer is then used to classify the feature vectors and perform bounding box regression to determine the component category and location. Simultaneously, a mask branch is used to predict a segmentation mask for each candidate region, thereby obtaining features from the fused image of the main transformer device. Finally, the main transformer device is identified and disassembled based on these features, and the specific shape and boundaries of each component are determined according to the predicted segmentation mask.
[0103] In another optional embodiment, a transformer-based instance segmentation model is used to obtain the main transformer device fusion image features and automatically disassemble the main transformer device. Specifically, the main transformer device fusion image is input into the transformer encoder, which encodes the global information of the image to obtain a feature representation containing rich contextual information. Then, a fixed set of prediction targets is generated by the decoder, each of which corresponds to a component of the main transformer. The decoder uses a multi-head attention mechanism to interact with the encoder output to learn the relationship between different regions. During the prediction process, the model directly outputs the category and segmentation mask of each component to obtain the main transformer device fusion image features. Finally, based on these features, the main transformer device is identified and disassembled according to the predicted category and mask.
[0104] It should be noted that this method uses a layered and progressive convolutional layer. Small-sized convolution kernels first capture local low-level features such as image edges and textures. Activation functions and pooling layers then downsample to retain significant features. Subsequently, larger convolution kernels are used to further extract low-dimensional feature maps, effectively increasing the model's receptive field and enabling more comprehensive and in-depth exploration of image features. The introduction of a convolutional attention mechanism allows the model to precisely focus on the target area, suppress background interference, and improve the recognition accuracy of the main transformer component. A feature pyramid is used to re-extract the fused image features and generate a prototype mask. The contribution weight coefficient in the formula reasonably measures the contribution of each prototype to the instance, making the generated mask more accurate to the actual component. Finally, the main transformer device is identified and disassembled based on the prototype mask, improving both accuracy and efficiency.
[0105] In Example 3, the above is a schematic diagram of a method for automatically disassembling a main transformer device. It should be noted that the technical solution of this AI-based communication method and the technical solution of the aforementioned method for automatically disassembling a main transformer device are based on the same concept. For details not described in detail in the technical solution of the very short-term photovoltaic power generation prediction system in this example, please refer to the description of the technical solution of the aforementioned method for automatically disassembling a main transformer device.
[0106] This embodiment also provides an automatic disassembly system for a main transformer device, comprising:
[0107] A data acquisition module is used to obtain the original image of the main transformer equipment;
[0108] A preprocessing module is used to align the original image of the main transformer device using a chessboard method to obtain an image of the main transformer device;
[0109] A feature extraction module is used to perform multimodal fusion of the main transformer device image through a convolutional neural network model to obtain a main transformer device fusion image; the main transformer device fusion image is input into the attention instance segmentation model for feature extraction to obtain the main transformer device fusion image features;
[0110] The identification and disassembly module is used to automatically disassemble the main transformer equipment based on the fusion image features of the main transformer equipment.
[0111] This embodiment also provides an electronic device suitable for automatic disassembly of a main transformer device, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the method for automatic disassembly of the main transformer device proposed in the above embodiment.
[0112] This embodiment further provides a storage medium storing a computer program, which, when executed by a processor, implements the method for automatically disassembling a main transformer device as proposed in the above embodiment.
[0113] The storage medium proposed in this embodiment and the method for realizing automatic disassembly of the main transformer device proposed in the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0114] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented with the help of software and necessary general hardware, and of course can also be implemented by hardware. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0115] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for automatically disassembling a main transformer device, characterized in that: include: Obtain the original image of the main transformer equipment; Performing data alignment on the original image of the main transformer device using a chessboard method to obtain an image of the main transformer device; Performing multimodal fusion on the main transformer device image through a convolutional neural network model to obtain a main transformer device fusion image; Inputting the main transformer device fusion image into the attention instance segmentation model for feature extraction to obtain the main transformer device fusion image features; Based on the fused image features of the main transformer device, the main transformer device is automatically disassembled.
2. The automatic disassembly method of the main transformer equipment according to claim 1, characterized in that: The main transformer device original image is data aligned using a chessboard method to obtain a main transformer device image, including: Acquire different coordinate system data of the original image of the main transformer device in combination with shooting device parameters; Coordinate conversion is performed according to the different coordinate system data to obtain the main transformer equipment image.
3. The automatic disassembly method of the main transformer equipment according to claim 2, characterized in that: Performing multimodal fusion on the main transformer device image through a convolutional neural network model to obtain a main transformer device fusion image, including: The convolutional neural network model includes an input layer, a gradient residual block layer, a convolution layer and an output layer; Performing preliminary feature extraction on the main transformer device image through the input layer to obtain a feature map; The feature map is input into the gradient residual block layer and the convolution layer to further process the features, and the main transformer device fusion image is output through the output layer.
4. The automatic disassembly method of the main transformer equipment according to claim 3, characterized in that: Also includes: Designing a loss function to train the convolutional neural network model and optimize the convolutional neural network model; Input the original image of the main transformer device into the trained convolutional neural network model to obtain a fused image of the main transformer device; The fused image of the main transformer equipment is compared with the original image of the main transformer equipment, and manual evaluation is performed to determine whether the fused image of the main transformer equipment includes visible light modalities and infrared light modalities.
5. The automatic disassembly method of the main transformer equipment according to claim 4, characterized in that: Inputting the main transformer device fusion image into the attention instance segmentation model for feature extraction to obtain the main transformer device fusion image features, including: The attention instance segmentation model uses hierarchical and progressive convolutional layers to obtain local features of the fused image of the main transformer device, and then further extracts low-dimensional features; The attention mechanism is used to obtain the target area of the main transformer device fusion image and obtain the main transformer device fusion image features.
6. The automatic disassembly method of the main transformer equipment according to claim 5, characterized in that: Based on the fusion image features of the main transformer device, the main transformer device is automatically disassembled, including: Extract features of the main transformer device fusion image through a feature pyramid to generate a prototype mask of the main transformer device; The main transformer device is identified and disassembled according to the main transformer device prototype mask.
7. The automatic disassembly method of the main transformer equipment according to claim 6, characterized in that: Generate the prototype mask of the main transformer device, which is expressed as: Among them, M instance is the instance mask, c i is the i-th coefficient, which represents the contribution weight of the prototype to the current instance, P i is the i-th prototype mask, and σ is the sigmoid function.
8. An automatic disassembly system for a main transformer device, applying the automatic disassembly method for a main transformer device according to any one of claims 1 to 7, characterized in that: include: A data acquisition module is used to obtain the original image of the main transformer equipment; A preprocessing module, configured to perform data alignment on the original image of the main transformer device using a chessboard method to obtain an image of the main transformer device; A feature extraction module is used to perform multimodal fusion on the main transformer device image through a convolutional neural network model to obtain a main transformer device fusion image; the main transformer device fusion image is input into an attention instance segmentation model for feature extraction to obtain main transformer device fusion image features; The identification and disassembly module is used to automatically disassemble the main transformer device based on the fusion image features of the main transformer device.
9. An electronic device comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the automatic disassembly method of the main transformer equipment according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the method for automatic disassembly of a main transformer device according to any one of claims 1 to 7.