High-altitude long-distance photovoltaic infrared image segmentation method based on Memory Visual Prompt full-automatic SAM model

By using a fully automated SAM model based on Memory Visual Prompt, the problem of infrared image segmentation for high-altitude, long-distance photovoltaic power stations was solved, achieving high-precision and real-time photovoltaic panel boundary recognition and improving the efficiency and accuracy of UAV inspections.

CN121904069APending Publication Date: 2026-04-21CGN WIND POWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CGN WIND POWER CO LTD
Filing Date
2025-12-25
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to perform accurate infrared image segmentation of photovoltaic power plants under high-altitude, long-distance conditions, especially when photovoltaic panels are arranged in complex patterns, there is background noise interference, and computing resources are limited. Segmentation models are difficult to achieve high accuracy and real-time performance.

Method used

A fully automated SAM model based on Memory Visual Prompt is adopted. Through data preprocessing, feature extraction and similarity retrieval, feature fusion and pseudomask generation, segmentation and boundary repair, combined with multi-scale feature decoding and specific loss function optimization, efficient segmentation of photovoltaic infrared images is achieved.

Benefits of technology

Stable and accurate infrared image segmentation of photovoltaic power plants under high-altitude, long-distance conditions by UAVs has been achieved, reducing the reliance on manual annotation and improving the automation level and segmentation accuracy of the segmentation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904069A_ABST
    Figure CN121904069A_ABST
Patent Text Reader

Abstract

The invention provides a high-altitude long-distance photovoltaic infrared image segmentation method based on a Memory Visual Prompt full-automatic SAM model, and belongs to the technical field of deep learning and unmanned aerial vehicle infrared inspection, and the method comprises the steps: firstly processing photovoltaic string data, carrying out the distortion correction, and constructing a corresponding training set, a test set and a verification set; the infrared picture is sent to a deep learning model to learn and simulate complex features; carrying out feature space matching on the learned features and good annotation data, and finding out a general area similar to the target object in the current infrared picture; and taking the areas as Prompt position areas in a mask encoder in an SAM model, and sending area information into the SAM to obtain a segmentation result. Calculating an error between the segmented photovoltaic panel and string segmentation effect and a true value, and training a model according to the error; and finally, testing the photovoltaic infrared segmentation model on the test set. According to the invention, a photovoltaic infrared image segmentation task of a long-distance small target can be stably and accurately completed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A High-Altitude Long-Range Photovoltaic Infrared Image Segmentation Method Based on the Memory Visual Prompt Fully Automatic SAM Model Technical Field

[0002] This invention belongs to the field of deep learning and UAV infrared inspection technology, specifically, it relates to a high-altitude long-range photovoltaic infrared image segmentation method based on the MemoryVisual Prompt fully automatic SAM model. Background Technology

[0003] Photovoltaic power plants typically consist of numerous photovoltaic panels, and manually inspecting each one is extremely time-consuming and labor-intensive. By using a positioning system, the inspection system can precisely pinpoint the faulty string and specific photovoltaic panel. Maintenance personnel can quickly locate the problem, eliminating the need to expend resources inspecting fault-free areas. This enables rapid and accurate troubleshooting and significantly improves inspection efficiency, reducing unnecessary comprehensive inspection time and maintenance personnel costs. This work not only affects the speed and cost of fault diagnosis but also directly relates to power generation efficiency and safety. Fault location enables early detection and accurate troubleshooting, preventing the fault from spreading to the entire photovoltaic system. Timely repair of faulty photovoltaic panels can prevent a drop in power across the entire string due to partial panel damage, significantly reducing power generation losses, improving power generation efficiency, and thus increasing the total power generation of the photovoltaic power plant. Improving the power generation efficiency of a photovoltaic power plant directly impacts its rate of return, enabling the plant to achieve a higher return on investment over its lifespan, enhancing the economic viability of the photovoltaic project, ensuring efficient and stable operation of the power plant, and extending equipment lifespan. In large-scale photovoltaic power plants, accurate positioning data helps maintenance personnel establish comprehensive maintenance records and databases, laying the foundation for subsequent intelligent operation and maintenance. By combining AI technology, predictive maintenance can be achieved, preventing potential problems. The accumulation and analysis of this data helps in developing more rational maintenance plans, promoting the intelligent development of power plant management.

[0004] Fault localization requires precision down to specific strings and individual photovoltaic panels. To achieve this, image segmentation models are typically introduced to delineate the boundaries of each photovoltaic panel and pinpoint its location. However, the application of segmentation models faces several challenges due to the characteristics and distribution of photovoltaic panels.

[0005] Photovoltaic panels in a photovoltaic (PV) field may be arranged at different angles, with varying row and column spacing and layouts, and the installation angles can differ depending on terrain and lighting conditions. This complexity of arrangement presents challenges to image segmentation models: the tilt angles and azimuths of the PV panels may vary, causing them to appear at different angles and shapes within the same image. These different perspectives can interfere with the model's segmentation accuracy. Furthermore, the spacing between different PV panels affects boundary segmentation; sometimes the spacing is very small or indistinct, making it difficult for the model to distinguish the boundaries between panels.

[0006] Photovoltaic panels are typically a single color, and reflective or shading effects can easily cause the panel to appear similar in color to the background (such as the ground, supports, or other objects) in the image. In such cases, the model struggles to accurately segment the edges of the photovoltaic panel: both the ground and the photovoltaic panel surface may be dark or gray, especially under shadow conditions, where the model may misidentify the background area as part of the photovoltaic panel. Due to sunlight reflection, aging of the photovoltaic panel, or dirt, the boundary lines may be unclear, which affects the model's segmentation accuracy and leads to unstable segmentation results.

[0007] In practical applications, segmentation models not only need to accurately identify the boundaries of photovoltaic panels, but also need to be real-time to meet the needs of large-scale inspections. This increases the complexity of the algorithm: segmentation calculations on high-resolution images require significant computational resources, especially during real-time inspections of large-area photovoltaic fields, where resource demands are even more pronounced. Photovoltaic panels are typically small in size, while inspection equipment (such as drones) may take pictures from a distance. To ensure the accuracy of the segmentation results, the model needs to be highly sensitive to image details.

[0008] In photovoltaic (PV) panel inspection tasks, using aerial images captured by drones to inspect large-scale PV panel arrays can effectively improve inspection efficiency. However, due to the resolution limitations of aerial photography, the size of individual PV panels in the images is small, which presents a challenge for segmentation models: because PV panels occupy only a few to tens of pixels in aerial images, the boundary areas are easily affected by various noises, such as the increased background information brought by aerial photography, ground, vegetation, supports, shadows, etc. When the boundaries of PV panels overlap or mix with the background, edge noise becomes more significant, increasing the difficulty for the model to distinguish the boundaries. With few pixels at the edges of PV panels, once noise intrudes or the edges become blurred, the model may struggle to distinguish the boundaries, leading to unclear boundaries or misclassification. Under different lighting conditions, the edge contours of PV panels may appear too bright or too dark, especially under conditions of shadows and reflections, making boundary noise more complex. Due to the small size of PV panels in aerial images, even small changes in shadows or lighting can exacerbate the boundary noise problem. Meanwhile, the small size of the photovoltaic panels makes it difficult for the model to accurately locate them, especially when the target is partially occluded or the background is too complex. Even if the model can detect the small target, localization errors may occur, affecting the final segmentation accuracy.

[0009] Segmentation models require a large amount of precisely labeled image data for training. However, in photovoltaic (PV) fields, each string typically contains multiple photovoltaic panels, which are densely packed and relatively small in area. To train the segmentation model, annotators need to individually label the boundaries of each PV panel, ensuring the accuracy of the labeling so that the model can distinguish each panel. Simultaneously, PV panel detection requires high-precision boundary segmentation so that subsequent algorithms can accurately determine the specific location and extent of defects. Excessive labeling errors can lead to poor model detection performance; therefore, labeling work must have pixel-level accuracy. To this end, annotators need to meticulously depict the edge regions of each PV panel, ensuring no omissions or errors. This places extremely high demands on the annotators' attention to detail and significantly increases labeling time. Especially since images acquired in actual PV field scenarios often contain complex backgrounds and various defect conditions, it is difficult to obtain a large amount of high-quality labeled data.

[0010] Precise fault location of photovoltaic panels is not only a core aspect of operation and maintenance management, but also plays a crucial role in cost reduction, efficiency improvement, increased revenue of photovoltaic power plants, and environmental protection. Through precise and intelligent fault location methods, the overall operation and maintenance level and long-term benefits of photovoltaic power plants will be significantly improved.

[0011] In summary, there are currently two bottlenecks in infrared image segmentation for UAV inspections: First, due to the presence of high-voltage power lines at some special sites, the safe cruising altitude of UAVs must be above 70m or even higher. Traditional segmentation models largely rely on large amounts of precisely labeled image data. However, photovoltaic panels in high-altitude images are densely packed and small in area, making labeling data extremely difficult and time-consuming, thus hindering the acquisition of a large amount of high-quality labeled data. Second, the currently popular visual segmentation SAM series models, which perform well in zero-shot generalization, require manual prompting (point, box, mask) engineering for their data in practical applications. This requirement undoubtedly limits the use of SAM series models in industrial fields. Summary of the Invention

[0012] To address the aforementioned issues, this invention proposes a high-altitude, long-range photovoltaic infrared image segmentation method based on the Memory Visual Prompt fully automatic SAM model. This method aims to provide all-weather unmanned inspection services for photovoltaic power plants and to complete stable and accurate photovoltaic infrared image segmentation tasks for small targets at long distances.

[0013] This invention is achieved through the following technical solution: A high-altitude long-range photovoltaic infrared image segmentation method based on the Memory Visual Prompt fully automatic SAM model: The method specifically includes the following steps: S1. Data preprocessing: Acquire raw high-altitude long-range photovoltaic infrared images, perform color mode conversion, optical distortion correction and coordinate projection transformation on the raw images to obtain standardized photovoltaic infrared images, and at the same time construct a historical image library that has undergone the same preprocessing operations. S2. Feature Extraction and Similarity Retrieval: A pre-trained ViT model is used to extract the current feature vector of the standardized photovoltaic infrared image and the historical feature vector of the historical image database; all historical feature vectors are stored in the Faiss vector database, and the top k historical feature vectors with the highest similarity to the current feature vector are selected by similarity calculation; S3. Feature Fusion and Pseudo-Mask Generation: The top k historical feature vectors selected in S2 are matched and enhanced with their corresponding labeled masks, and the features are refined through a self-attention mechanism; the current feature vector in S2 is refined in the same way, and the similarity between the refined historical feature vector and the current feature vector is calculated through cross-attention, and the softmax function is combined to generate pseudo-mask hints to guide segmentation. S4. Segmentation and Boundary Repair: The pseudo-mask hints generated in S3 are input into SAM. Multi-scale features are extracted by the SAM image encoder and obtained as a preliminary segmentation mask by the SAM mask decoder. The deep features output by the SAM image encoder and the features of the two intermediate layers are input into a multi-level decoder composed of UNet-Decoder. Multi-scale features are fused through Transformer and Patch Upsampling operations to gradually restore the image size and repair the segmentation boundary details to obtain the final segmentation result. S5. Model Training and Optimization: Calculate the error between the final segmentation result in S4 and the true label, adjust the model parameters using a specific loss function and optimizer, and improve the model's accuracy in distinguishing between photovoltaic panels and background boundaries through multi-type Prompt decoupling training.

[0014] Furthermore, in S1, the black-red color mode of the original photovoltaic infrared image is adjusted to the iron-red mode through color mode conversion; optical distortion correction is performed through distortion coefficients [k1, k2, p1, p2]; and coordinate projection from the physical imaging plane to the pixel coordinate system is completed using the camera intrinsic parameter matrix K.

[0015] Furthermore, in S3, The formula for calculating the annotation mask is: ; The historical feature vector is refined as follows: ; The current feature vector is refined as follows: ; in mF i These are the deep features of historical images previously selected as guidance through Query-Select. y i This corresponds to the segmentation result of the image. F This represents the deep features of the currently queried image. The corresponding annotation mask is used, MSA is the self-attention mechanism, and LN is the layer normalization. The formula for generating the pseudo-mask hint is:

[0016] d k It is a hidden dimension.

[0017] Furthermore, in S4, The three-layer features output by the image encoder of the SAM have scales corresponding to 1x, 1 / 2x, and 1 / 4x of the original photovoltaic infrared image, respectively, and are deep features and two intermediate layer features.

[0018] Furthermore, in S4, The multi-level decoder has three layers, which correspond one-to-one with the three layers of features output by the SAM image encoder; The decoder achieves bidirectional fusion of multi-scale features through the Cross-Attention mechanism, and gradually restores the feature map to the size of the original photovoltaic infrared image through the Patch Upsampling operation, thereby repairing the detail defects of the segmentation boundary.

[0019] Furthermore, in S5, The specific loss function is Focal Loss, and its calculation formula is as follows:

[0020] Where p is the model's predicted probability of the target region, and y is the true label of the region. This loss function strengthens the model's focus on the segmentation boundary region.

[0021] Furthermore, in S5, the multi-type Prompt decoupling training includes: The background segmentation result of the current image is obtained by using the background segmentation prompt as guidance information, the edge box segmentation prompt of the current photovoltaic panel is obtained by using the edge box segmentation prompt as guidance information, the internal segmentation result of the current photovoltaic panel is obtained by using the internal segmentation prompt of the photovoltaic panel as guidance information, and finally, the internal segmentation prompt of the photovoltaic panel is used as the final result in actual use.

[0022] A high-altitude long-range photovoltaic infrared image segmentation system based on the Memory Visual Prompt fully automatic SAM model; The system includes a data preprocessing module, a query-select module, a memory-attention module, a SAM segmentation and UNet-Decoder decoding module, and a training and optimization module; The data preprocessing module acquires the original high-altitude long-distance photovoltaic infrared image, performs color mode conversion, optical distortion correction and coordinate projection transformation on the original image to obtain a standardized photovoltaic infrared image, and at the same time constructs a historical image library that has undergone the same preprocessing operation. The Query-Select module is used for feature extraction and similarity retrieval: A pre-trained ViT model is used to extract the current feature vector of the standardized photovoltaic infrared image and the historical feature vector of the historical image database. All historical feature vectors are stored in the Faiss vector database, and the top k historical feature vectors with the highest similarity to the current feature vector are selected by similarity calculation. The Memory-Attention module is used for feature fusion and pseudo-mask generation: it matches and enhances the top k historical feature vectors selected in S2 with their corresponding labeled masks, and refines the features through a self-attention mechanism; it refines the current feature vector in S2 in the same way, and then calculates the similarity between the refined historical feature vector and the current feature vector through cross-attention, and generates pseudo-mask prompts for guiding segmentation by combining the softmax function; The SAM segmentation and UNet-Decoder decoding module is used for segmentation and boundary repair: the pseudo-mask prompts generated by S3 are input into SAM, multi-scale features are extracted by the SAM image encoder, and a preliminary segmentation mask is obtained by the SAM mask decoder; the deep features output by the SAM image encoder and the features of two intermediate layers are input into a multi-level decoder composed of UNet-Decoder, and the multi-scale features are fused by Transformer and Patch Upsampling operations to gradually restore the image size, repair the segmentation boundary details, and obtain the final segmentation result; The training and optimization module calculates the error between the final segmentation result and the true label in S4, adjusts the model parameters using a specific loss function and optimizer, and improves the model's accuracy in distinguishing between photovoltaic panels and background boundaries through multi-type Prompt decoupling training.

[0023] An electronic device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the above method.

[0024] A computer-readable storage medium for storing computer instructions that, when executed by a processor, implement the steps of the above-described method.

[0025] Beneficial effects of the invention This invention is based on Memory Visual Prompt and uses historical labeled data to generate pseudo-mask prompts, replacing the manual operation required by the traditional SAM model. Segmentation can be completed without human intervention, solving the core problem that the SAM model cannot be used for industrial-grade photovoltaic inspection due to its reliance on manual prompts, and meeting the automation requirements of large-scale unmanned inspection. Attached Figure Description

[0026] Figure 1 This is a flowchart illustrating the execution of the high-altitude photovoltaic instance segmentation model of the present invention. Figure 2 This is a structural diagram of the SAM automatic high-altitude long-range photovoltaic infrared imaging model based on Memory Visual Prompter. Figure 3 This is a design diagram of the Query-Select module of the present invention; Figure 4 This is a design diagram of the Memory-attention module of the present invention; Figure 5 This is a design diagram of the Multi-layel Decoder module of the present invention; Figure 6 Infrared images generated by using drones to inspect photovoltaic power stations from high altitude and long distance. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] Unless otherwise specified, the experimental methods used in the following embodiments are conventional methods. All materials, reagents, methods, and instruments used, unless otherwise specified, are conventional materials, reagents, methods, and instruments in the art, and can be obtained commercially by those skilled in the art. The SAM automatic high-altitude long-range photovoltaic infrared image model based on Memory Visual Prompter proposed in the embodiments is trained using the Adam optimizer with a learning rate set to 1e-4. The ViT encoder has a depth of 12 layers, an embedding length of 768, and a patch size of 16*16. All other configurations are consistent with the SAM_B model. The multi-scale Transformer decoder has 3 layers, representing 1, 1 / 2, and 1 / 4 different levels.

[0029] This invention proposes a high-altitude long-range photovoltaic infrared image segmentation method based on the Memory Visual Prompt SAM model. Its workflow is as follows: Figure 1As shown, the photovoltaic string data is first processed, distortion correction is performed, and corresponding training, testing, and validation sets are constructed. Then, infrared images are fed into a deep learning model to learn and simulate complex features. The learned features are matched with well-labeled data in the feature space to identify general regions in the current infrared image that are similar to the target object. These regions are then used as the Prompt locations in the mask encoder of the SAM model, and the region information is fed into the SAM to obtain the segmentation results. Next, the error between the segmented photovoltaic panels and strings and the true values ​​is calculated, and the model is trained based on this error. Finally, the photovoltaic infrared segmentation model is tested on the test set.

[0030] The raw infrared inspection data is stored in JPG format, using a black-and-red color mode by default. In practical applications, it needs to be adjusted to a reddish-brown mode. Furthermore, due to optical distortion in the infrared lens mounted on the UAV, radial and tangential distortions will appear at the edges of the raw image, requiring correction using distortion coefficients [k1, k2, p1, p2]. Simultaneously, high-altitude infrared imaging also involves coordinate projection calculations, using the camera intrinsic parameter matrix K to transform from the physical imaging plane to the pixel coordinate system, ultimately completing the data preparation for high-altitude long-range infrared imaging of the UAV.

[0031] After obtaining the photovoltaic infrared image, it is necessary to segment the components within the image to facilitate intuitive visualization and precise fault localization. Considering the lack of a large amount of high-quality labeled data, this invention employs the SAM series models as the backbone. The SAM series models aim to introduce zero-shot learning prompting techniques from the field of Natural Language Processing (NLP) into image segmentation tasks, enabling zero-shot and few-shot learning on new datasets and tasks with only a small amount of image cues. SAM uses three types of image cues (points, bounding boxes, and masks) to simply inform the model of the general objects of interest. By providing these initial cues, the model is guided to focus on the target region of interest, avoiding excessive attention to irrelevant backgrounds, allowing the segmentation model to adapt to various tasks or target categories without retraining. However, this operation requires manual prompting. Therefore, this invention introduces the Memory Visual Prompter method into SAM, aiming to use a small amount of labeled data to inform the model of the data information to be segmented, so that the segmentation model can obtain the segmented object results not only from the model structure, but also... It can also integrate Memory Visual Prompter visual guidance to enhance the segmentation effect, guiding the model to focus on specific regions or semantic content, thereby achieving more efficient and accurate segmentation. At the same time, it uses RAG technology and multi-scale decoding to improve the performance of segmenting small objects.

[0032] This high-altitude, long-range, automatic high-altitude, long-range photovoltaic infrared imaging model is as follows: Figure 2 As shown, the model consists of ViT (Vision Transformer), Query-Select, Memory-Attention, SAM (Segment Anything Model), and a multi-layer Decoder.

[0033] First, ViT extracts the representation tokens of the current photovoltaic infrared image to be processed and historical image representations. Then, the historical image representation tokens are stored in the vector database Faiss, enabling efficient similarity search. Using the representations in the same vector space obtained from the current image to be processed, the representations of the top k most similar images are selected as guidance. ViT is a pre-trained, fixed model that does not require fine-tuning for segmentation; its purpose is simply to select images similar to the current image to be processed in the high-dimensional feature space, such as similarity in illumination, terrain, and style. This selection of historical information facilitates subsequent guidance. The feature information from the selected historical images is then sent to the Memory-Attention module. In this module, a deep feature image mask is obtained using the annotation information of the corresponding image. This mask is used to manipulate the image features and select high-dimensional feature information of the object to be segmented, making the guidance more specific. Finally, a pseudo-mask indication of the current image to be processed is obtained through the interaction attention between the features extracted from the current object to be segmented and the deep features of the image to be processed. This pseudo-mask directs interaction with the SAM mask decoder to obtain the final image segmentation result.

[0034] Assumption This represents a photovoltaic infrared image after distortion correction. express The present invention uses a historical image that has already been processed to segment the current infrared image X, making its segmentation logic follow Y. The specific workflow is defined as follows:

[0035]

[0036]

[0037]

[0038]

[0039]

[0040]

[0041]

[0042]

[0043]

[0044] in This represents the Vision Transformer feature extractor. Specifically, it represents and tokenizes images by dividing them into patches. Operations representing similarity calculations. To select the best k features from these samples, This refers to the corresponding annotation information, i.e., the representation of the binary mask at this resolution. It is a cross-attention module. The image encoder representing SAM, A bidirectional attention fusion module for SAM, integrating image features and guidance information. This represents the initial mask obtained through the SAM pipeline. This represents the upsampling module of UNet. .

[0045] Next, we will introduce in detail Query-Select, Memory-Attention, SAM (Segment Anything Model), and UNet-Decoder, as well as the training process of the models.

[0046] Query-Select: To fully filter historical guidance images for UAV-based photovoltaic segmentation that better match the current image, this invention employs the Faiss vector database to quickly retrieve images similar to the current image in a high-dimensional feature space. Faiss is an open-source library developed by Meta, focusing on efficient similarity search and dense vector clustering. It supports various index structures, such as IVF (Inverted Index) and PQ (Product Quantization), enabling fast retrieval in large-scale vector data. Efficient similarity search can be achieved by adding feature vectors from historical images to the Faiss index.

[0047] The vectors used in this invention are derived from the ViT (Vision Transformer) architecture, which employs a general form for global attention mechanisms. Similar to traditional Transformer architectures, the core of ViT lies in segmenting the image into fixed-size patches and treating these patches as sequential inputs, thereby introducing 2D positional encoding to preserve the spatial information of the image. The multi-head attention mechanism is implemented by connecting M single-head attention modules in parallel, which are then integrated through a linear projection layer L. Residual connections, Dropout, and layer normalization are employed to enhance the model's stability and performance.

[0048]

[0049] in The , [] operation indicates splicing in the channel dimension.

[0050] At the same time, such as Figure 3 As shown, to improve the segmentation effect at high altitudes and long distances, this invention extracts not only single-scale features but also multi-scale features during the feature extraction stage to capture information from the image at different scales. Through multi-scale feature fusion, details and global information in the image can be better processed. Specifically, for a 12-layer ViT model structure, features from layers 4, 8, and 12 are extracted as information representations of the image, aiming to provide a more dimensional and comprehensive representation of information such as terrain and object shapes.

[0051] Memory-attention: like Figure 4 As shown, the Memory-attention architecture designed in this invention is an effective architecture for transferring historical object information representations. It can extract information about the object to be segmented from historical images, perform good selection learning, and transfer this information to the current image information to better capture object information in the input photovoltaic detection image. This can effectively locate the position of the target object to be segmented in the target image. Specifically, the model structure consists of two main modules: self-attention feature extraction and cross-attention modality learning. The process is defined as follows:

[0052]

[0053]

[0054]

[0055] in, mF i These are the deep features of historical images previously selected as guidance through Query-Select. y i This corresponds to the segmentation result of the image. F This represents the deep features of the currently queried image. d k This invention addresses the hidden layer dimension by first processing the corresponding ground truth segmentation value to make it identical in shape to the deep features. Then, it masks the deep features of historical images to extract information about the object to be segmented. A self-attention mechanism allows the features to capture the segmentation of other objects globally, thus refining their own object features. For example, by detecting that the masks of other parts are rectangular, the model constrains its own features to a rectangular pattern, helping it better handle background noise and improve target localization. Subsequently, by querying the similarity between image features and the features of the object to be segmented, the invention finds regions in the current image that may contain objects and obtains the corresponding mask information through these similar regions, thus obtaining the possible mask representation of the current image.

[0056] Multi-level decoding: While the multi-scale Query-Select and Memory-Attention methods described above, used in conjunction with SAM, can effectively model the spatial features of photovoltaic infrared images and obtain segmentation masks, the output shape of the SAM model is typically downsampled to four times the size of the original image. Therefore, it generally uses simple interpolation and upsampling as the final result. This approach usually doesn't present significant problems. However, this invention addresses the issue of small objects at high altitudes, where numerous photovoltaic panels are densely packed. Simply using interpolation and upsampling to obtain the final result would undoubtedly have disastrous effects on segmenting object boundaries. Therefore, this invention proposes using a multi-layered decoder to improve this operation, ensuring good boundary preservation even when the mask is restored to the original image size and shape, thus more accurately segmenting numerous small targets in photovoltaic infrared images.

[0057] Specifically, this invention not only uses the deepest features of the Image-Encoder like SAM, but also obtains and retains the features of the two intermediate layers of the Encoder, which represent different scale information. Then, when the size is restored step by step by the Transformer in conjunction with the patch Upsampling operation, the simple features are upsampled and repaired by the information of these corresponding scales, so as to perform target segmentation more accurately.

[0058] Model training: Since this invention utilizes drones to inspect photovoltaic power stations from high altitudes and long distances, the resulting infrared images contain numerous densely packed small targets, such as... Figure 6 As shown. In order to enforce constraints on these densely packed small targets during model training, this invention designs a training optimization for small target object segmentation and adopts a more suitable training function.

[0059] This invention argues that there is still room for improvement in the design of the model's prompts. Specifically, for a captured photovoltaic infrared image, assuming its mask is a mask, this invention believes that it provides three types of information about the image: edge information, internal information, and irrelevant background information. These three types of information are derived from the initial benchmark data and used as different prompts in the fine-tuning training of the model. This decouples the three types of information: the background class prompt is used as the guiding information to obtain the background class segmentation result of the current image; the edge box prompt is used as the guiding information to obtain the edge box of the current photovoltaic panel; and the interior of the photovoltaic panel is used as the guiding information to obtain the internal result of the current photovoltaic panel. Finally, in actual use, the internal prompt of the photovoltaic panel is used as the final result, thereby enabling the model to have a more accurate understanding of the background between photovoltaic panels.

[0060] Unlike traditional segmentation loss functions, this invention introduces Focal Loss, which makes the model focus more on edge loss, thereby improving the model's performance at edges in the segmentation results.

[0061] A high-altitude long-range photovoltaic infrared image segmentation system based on the Memory Visual Prompt fully automatic SAM model. The system includes a data preprocessing module, a query-select module, a memory-attention module, a SAM segmentation and UNet-Decoder decoding module, and a training and optimization module; The data preprocessing module acquires the original high-altitude long-distance photovoltaic infrared image, performs color mode conversion, optical distortion correction and coordinate projection transformation on the original image to obtain a standardized photovoltaic infrared image, and at the same time constructs a historical image library that has undergone the same preprocessing operation. The Query-Select module is used for feature extraction and similarity retrieval: A pre-trained ViT model is used to extract the current feature vector of the standardized photovoltaic infrared image and the historical feature vector of the historical image database. All historical feature vectors are stored in the Faiss vector database, and the top k historical feature vectors with the highest similarity to the current feature vector are selected by similarity calculation. The Memory-Attention module is used for feature fusion and pseudo-mask generation: it matches and enhances the top k historical feature vectors selected in S2 with their corresponding labeled masks, and refines the features through a self-attention mechanism; it refines the current feature vector in S2 in the same way, and then calculates the similarity between the refined historical feature vector and the current feature vector through cross-attention, and generates pseudo-mask prompts for guiding segmentation by combining the softmax function; The SAM segmentation and UNet-Decoder decoding module is used for segmentation and boundary repair: the pseudo-mask prompts generated by S3 are input into SAM, multi-scale features are extracted by the SAM image encoder, and a preliminary segmentation mask is obtained by the SAM mask decoder; the deep features output by the SAM image encoder and the features of two intermediate layers are input into a multi-level decoder composed of UNet-Decoder, and the multi-scale features are fused by Transformer and Patch Upsampling operations to gradually restore the image size, repair the segmentation boundary details, and obtain the final segmentation result; The training and optimization module calculates the error between the final segmentation result and the true label in S4, adjusts the model parameters using a specific loss function and optimizer, and improves the model's accuracy in distinguishing between photovoltaic panels and background boundaries through multi-type Prompt decoupling training.

[0062] An electronic device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the above method.

[0063] A computer-readable storage medium for storing computer instructions that, when executed by a processor, implement the steps of the above-described method.

[0064] The memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the methods described in this invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0065] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means such as coaxial cable, optical fiber, digital subscriber line, DSL, or wireless means such as infrared, wireless, microwave, etc. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium such as a floppy disk, hard disk, magnetic tape; an optical medium such as a high-density digital video disc, DVD; or a semiconductor medium such as a solid-state disk, SSD, etc.

[0066] In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software. The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or by a combination of hardware and software modules in the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are omitted here.

[0067] It should be noted that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiments can be completed by the integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied as execution by a hardware decoding processor, or as execution by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above methods.

[0068] The above provides a detailed description of the high-altitude long-range photovoltaic infrared image segmentation method based on the Memory Visual Prompt fully automatic SAM model proposed in this invention. The principles and implementation methods of this invention have been explained. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A high-altitude long-range photovoltaic infrared image segmentation method based on the Memory Visual Prompt fully automatic SAM model, characterized in that: The method specifically includes the following steps: S1. Data preprocessing: Acquire raw high-altitude long-range photovoltaic infrared images, perform color mode conversion, optical distortion correction and coordinate projection transformation on the raw images to obtain standardized photovoltaic infrared images, and at the same time construct a historical image library that has undergone the same preprocessing operations. S2. Feature Extraction and Similarity Retrieval: The ViT model is used to extract the current feature vector of the standardized photovoltaic infrared image and the historical feature vector of the historical image database; all historical feature vectors are stored in the Faiss vector database, and the top k historical feature vectors with the highest similarity to the current feature vector are selected by similarity calculation; S3. Feature Fusion and Pseudo-Mask Generation: The top k historical feature vectors selected in S2 are matched and enhanced with their corresponding labeled masks, and the features are refined through a self-attention mechanism; the current feature vector in S2 is refined in the same way, and the similarity between the refined historical feature vector and the current feature vector is calculated through cross-attention, and the softmax function is used to generate pseudo-mask prompts to guide segmentation. S4. Segmentation and Boundary Repair: The pseudo-mask hints generated in S3 are input into SAM. Multi-scale features are extracted by the SAM image encoder and obtained as a preliminary segmentation mask by the SAM mask decoder. The deep features output by the SAM image encoder and the features of the two intermediate layers are input into a multi-level decoder composed of UNet-Decoder. Multi-scale features are fused through Transformer and Patch Upsampling operations to gradually restore the image size and repair the segmentation boundary details to obtain the final segmentation result. S5. Model Training and Optimization: Calculate the error between the final segmentation result in S4 and the true label, adjust the model parameters using a specific loss function and optimizer, and improve the model's accuracy in distinguishing between photovoltaic panels and background boundaries through multi-type Prompt decoupling training.

2. The infrared image segmentation method according to claim 1, characterized in that: In S1, the black-red color mode of the original photovoltaic infrared image is adjusted to iron-red mode through color mode conversion; optical distortion correction is performed through distortion coefficients [k1, k2, p1, p2]; and coordinate projection from the physical imaging plane to the pixel coordinate system is completed using the camera intrinsic parameter matrix K.

3. The infrared image segmentation method according to claim 2, characterized in that: In S3, The formula for calculating the annotation mask is: ; The historical feature vector is refined as follows: ; The current feature vector is refined as follows: ; in mF i These are the deep features of historical images previously selected as guidance through Query-Select. y i This corresponds to the segmentation result of the image. F This represents the deep features of the currently queried image. For the corresponding annotation mask, MSA is the self-attention mechanism, and LN is the layer normalization; The formula for generating the pseudo-mask hint is: d k It is a hidden dimension.

4. The infrared image segmentation method according to claim 3, characterized in that: In S4, The three-layer features output by the image encoder of the SAM correspond to 1x, 1 / 2x, and 1 / 4x the scale of the original photovoltaic infrared image, respectively, and are the deep features and two intermediate layer features.

5. The infrared image segmentation method according to claim 4, characterized in that: In S4, The multi-level decoder has three layers, which correspond one-to-one with the three layers of features output by the SAM image encoder; The decoder achieves bidirectional fusion of multi-scale features through the Cross-Attention mechanism, and gradually restores the feature map to the size of the original photovoltaic infrared image through the PatchUpsampling operation, thereby repairing the detail defects of the segmentation boundary.

6. The infrared image segmentation method according to claim 5, characterized in that: In S5, The specific loss function is Focal Loss, and its calculation formula is as follows: Where p is the model's predicted probability of the target region, and y is the true label of the region. This specific loss function strengthens the model's focus on the segmentation boundary region.

7. The infrared image segmentation method according to claim 6, characterized in that: In S5, the multi-type Prompt decoupling training includes: The background segmentation result of the current image is obtained by using the background segmentation prompt as guidance information, the edge box segmentation prompt of the current photovoltaic panel is obtained by using the edge box segmentation prompt as guidance information, the internal segmentation result of the current photovoltaic panel is obtained by using the internal segmentation prompt of the photovoltaic panel as guidance information, and finally, the internal segmentation prompt of the photovoltaic panel is used as the final result in actual use.

8. A high-altitude long-range photovoltaic infrared image segmentation system based on the Memory Visual Prompt fully automatic SAM model, characterized in that: The system is used to perform the infrared image segmentation method according to any one of claims 1 to 7; The system includes a data preprocessing module, a query-select module, a memory-attention module, a SAM segmentation and UNet-Decoder decoding module, and a training and optimization module; The data preprocessing module acquires the original high-altitude long-distance photovoltaic infrared image, performs color mode conversion, optical distortion correction and coordinate projection transformation on the original image to obtain a standardized photovoltaic infrared image, and at the same time constructs a historical image library that has undergone the same preprocessing operation. The Query-Select module is used for feature extraction and similarity retrieval: A pre-trained ViT model is used to extract the current feature vector of the standardized photovoltaic infrared image and the historical feature vector of the historical image database. All historical feature vectors are stored in the Faiss vector database, and the top k historical feature vectors with the highest similarity to the current feature vector are selected by similarity calculation. The Memory-Attention module is used for feature fusion and pseudo-mask generation: it matches and enhances the top k historical feature vectors selected in S2 with their corresponding labeled masks, and refines the features through a self-attention mechanism; it refines the current feature vector in S2 in the same way, and then calculates the similarity between the refined historical feature vector and the current feature vector through cross-attention, and generates pseudo-mask prompts for guiding segmentation by combining the softmax function; The SAM segmentation and UNet-Decoder decoding module is used for segmentation and boundary repair: the pseudo-mask prompts generated by S3 are input into SAM, multi-scale features are extracted by the SAM image encoder, and a preliminary segmentation mask is obtained by the SAM mask decoder; the deep features output by the SAM image encoder and the features of two intermediate layers are input into a multi-level decoder composed of UNet-Decoder, and the multi-scale features are fused by Transformer and Patch Upsampling operations to gradually restore the image size, repair the segmentation boundary details, and obtain the final segmentation result; The training and optimization module calculates the error between the final segmentation result and the true label in S4, adjusts the model parameters using a specific loss function and optimizer, and improves the model's accuracy in distinguishing between photovoltaic panels and background boundaries through multi-type Prompt decoupling training.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method of claim 8.

10. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method of claim 8.