Visual identification method for solar photovoltaic panel in satellite image based on multi-task learning
Through the design of a multi-task learning framework and a shared encoder, the problem of neglecting background information in solar photovoltaic panel recognition is solved, efficient region segmentation and background recognition are achieved, and the recognition accuracy and adaptability of the model are improved.
Patent Information
- Application Number
- CN202510692778.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-05
AI Technical Summary
Existing technologies only focus on segmentation when identifying solar photovoltaic panels, while ignoring background information, resulting in limitations in solar power generation analysis and decision-making.
A multi-task learning-based method is adopted to extract multi-scale features through a shared encoder, and a segmentation decoder and a classification decoder are used to respectively realize the regional segmentation of solar photovoltaic panels and background type recognition. The MetaFormer framework and multiple training strategies are combined to optimize the model performance.
The model has improved its accuracy and efficiency in identifying solar photovoltaic panels, provided more comprehensive data support, enhanced its adaptability to complex scenarios, and improved segmentation and classification performance.
Smart Images

Figure CN120599487A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine vision technology, and in particular to a method for visually recognizing solar photovoltaic panels in satellite images based on multi-task learning. Background Art
[0002] Against the backdrop of an accelerating global energy transition, solar photovoltaic technology, as a key component of a clean energy system, is experiencing unprecedented growth opportunities. With the continuous advancement of photovoltaic technology, the conversion efficiency of photovoltaic cells has continued to increase, while costs have gradually decreased. Solar photovoltaic power generation has gained widespread adoption worldwide. However, in actual operation, the performance of photovoltaic panels (PVP) can be affected by various factors, such as surface contamination, hot spot effects, and mechanical damage, which can lead to reduced power generation efficiency. Therefore, timely monitoring and maintenance of PVP status is crucial for improving the efficiency and reliability of photovoltaic power generation systems, and has important practical implications for energy management, environmental monitoring, and urban planning.
[0003] With the rapid development of computer vision technology, detection methods based on image recognition have become a research hotspot. High-resolution image data acquired through drones and satellite remote sensing, combined with computer vision algorithms such as object detection and semantic segmentation, can enable automated analysis of equipment status and anomaly warnings. Based on this, it is possible to accurately identify and locate solar PVP in satellite images using PVP detection.
[0004] Most existing research focuses on accurately identifying PVP areas through semantic segmentation. After extracting PVP labels through semantic segmentation, researchers use the area of the labels to calculate the actual PVP area based on the ratio of the satellite image size to the actual size, thereby inferring the area's power generation capacity. However, in practical applications, simply identifying the area of the PVP is not enough. PVPs in different contexts, such as rooftops, ground, deserts, and water bodies, have significant differences in their power generation potential, layout planning, and environmental impact—all key factors affecting solar power generation estimation. The location, size, and distribution density of PVPs vary across diverse environments, including residential areas, mountainous areas, lakes, and deserts, and the methods for calculating potential solar energy also differ. Therefore, focusing solely on PVP segmentation while ignoring the extraction of contextual information will lead to limitations in subsequent analysis and decision-making. Summary of the Invention
[0005] The present invention provides a solar photovoltaic panel visual recognition method in satellite images based on multi-task learning to solve the technical problem that the existing technology only focuses on the segmentation of PVP but ignores the extraction of its background information, resulting in limitations in subsequent solar photovoltaic power generation analysis and related decision-making.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In one aspect, the present invention provides a method for visually recognizing solar photovoltaic panels in satellite images based on multi-task learning. The method comprises:
[0008] Obtain a solar photovoltaic panel satellite image dataset; wherein the solar photovoltaic panel satellite images in the dataset include not only solar photovoltaic panel semantic segmentation labels but also background information of the solar photovoltaic panels;
[0009] Constructing a multi-task learning model; wherein the input of the multi-task learning model is a satellite image of a solar photovoltaic panel, and the output is a segmentation result of the solar photovoltaic panel area and the background type of the solar photovoltaic panel;
[0010] Training the multi-task learning model using the dataset;
[0011] Use the trained multi-task learning model to achieve visual recognition of solar photovoltaic panels in satellite imagery.
[0012] Furthermore, the multi-task learning model includes a shared encoder, a segmentation decoder, and a classification encoder;
[0013] The shared encoder is used to extract multi-scale features from satellite images of solar photovoltaic panels input to the model;
[0014] The segmentation decoder is used to segment the area of the solar photovoltaic panel from the satellite image of the solar photovoltaic panel input to the model based on the multi-scale features extracted by the shared encoder;
[0015] The classification decoder is used to classify the background of the solar photovoltaic panel satellite image input to the model based on the multi-scale features extracted by the shared encoder, and identify the type of the background.
[0016] Furthermore, the shared encoder is a four-layer network structure designed using the MetaFormer framework; wherein the first layer of the shared encoder consists of a Stem layer and a basic block, and each of the second to fourth layers consists of a downsampling layer and a basic block respectively; the input of the first layer of the shared encoder is a satellite image of solar photovoltaic panels, and the input of each subsequent layer is the feature data output by the previous layer;
[0017] The Stem layer and the downsampling layer are used to adjust the feature resolution and realize channel dimension conversion; the basic block is used to further extract and aggregate context information of the features adjusted by the Stem layer or the downsampling layer.
[0018] Furthermore, the basic block includes a spatial feature extraction network and a multi-scale feature fusion network;
[0019] The spatial feature extraction network is used to extract spatial feature information of input features from multiple angles;
[0020] The multi-scale feature fusion network is used to capture multi-scale information of different channels of feature information extracted by the spatial feature extraction network, enhance information interaction between different channels, and fuse multi-scale information of feature channels.
[0021] Furthermore, the process of extracting spatial feature information of input features from multiple angles by the spatial feature extraction network includes:
[0022] First, the input feature X is extracted by 1×1 convolution and 3×3 depth-separable convolution DWConv to generate the core feature layer Y; then, the core feature layer Y is pooled by global average pooling GAP and γ s Spatial self-interaction and activation function GELU adjust the local perception ability of the feature to generate an adaptive local perception feature P. Then, a large-size deep convolution kernel is used to promote the interaction of long-distance features on the feature map to generate a wide-area static feature Q. Finally, the adaptive local perception feature P and the wide-area static feature Q are fused to generate the final feature representation Z. The size of the large-size deep convolution kernel is 13x13.
[0023] Furthermore, the multi-scale feature fusion network uses parallel depthwise convolutions with kernel sizes of [1, 3, 5, 7], and each depthwise convolution processes one quarter of the channels.
[0024] Furthermore, the segmentation decoder segments the region of the solar photovoltaic panels from the satellite image of the solar photovoltaic panels input to the model based on the multi-scale features extracted by the shared encoder, including:
[0025] The segmentation decoder first upsamples the multi-scale features extracted by the shared encoder through bilinear interpolation or transposed convolution to restore them to the input image size, then uses 1×1 convolution to compress the channel dimension, mapping the number of feature channels to the number of target categories. Finally, a probability distribution map is generated through pixel-by-pixel Softmax classification, and the maximum probability value is selected to determine the category of each pixel to segment the area of the solar photovoltaic panel.
[0026] Furthermore, the classification encoder classifies the background of the solar photovoltaic panel satellite image input to the model based on the multi-scale features extracted by the shared encoder. The process of identifying the type of the background includes:
[0027] The classification decoder compresses the multi-scale features extracted by the shared encoder into a fixed-length feature vector through global average pooling, and then predicts the background category of the satellite image of solar photovoltaic panels through a fully connected layer.
[0028] Furthermore, the training strategies for training the multi-task learning model using the data set include: a cyclic training strategy, a repeated sequence training strategy, a weighted random training strategy, and a parallel training strategy;
[0029] The cyclic training strategy alternately performs training for the segmentation task and the classification task. In each training cycle, the model first performs training for the segmentation task and then for the classification task, and so on.
[0030] The repeated sequence training strategy specifies a fixed training sequence before the training begins, first training one task several times continuously, and then training another task several times continuously;
[0031] The weighted random training strategy assigns a probability weight to each task, and in each training iteration, randomly selects a task for training based on these weights;
[0032] The parallel training strategy trains the segmentation task and the classification task simultaneously.
[0033] Furthermore, when training the model using the parallel training strategy, the loss function is:
[0034] L all =α·L class +(1-α)·L seg
[0035] Among them, α is the preset weight parameter; L all is the total loss; L class is the cross entropy loss for the background classification task; L seg Cross entropy loss for the solar photovoltaic panel segmentation task.
[0036] On the other hand, the present invention further provides an electronic device, comprising a processor and a memory; wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the above method.
[0037] In yet another aspect, the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, and the instruction is loaded and executed by a processor to implement the above method.
[0038] The beneficial effects brought about by the technical solution provided by the present invention include at least:
[0039] 1. Multi-task learning framework improves comprehensive recognition capabilities: The multi-task learning framework proposed in this paper simultaneously implements the semantic segmentation and background classification tasks of PVP by sharing an encoder, achieving 94.57% mIoU and 87.36% classification accuracy on the PV03 dataset. The shared feature extraction module reduces computational redundancy and improves model efficiency. The background classification results provide more comprehensive data support for subsequent energy management.
[0040] 2. PVFormer network enhances feature extraction capabilities: The PVFormer network designed in this paper based on the MetaFormer architecture improves the feature extraction effect. The spatial feature extraction network (SFEN) integrates local and global features, and the multi-scale feature fusion network (MSFFN) adopts [1, 3, 5, 7] parallel convolution kernels. Experiments show that the performance of the PVFormer designed in this paper is better than mainstream models such as Swin-T in both segmentation and classification tasks.
[0041] 3. Optimized training strategy: To address the training difficulties of multi-task learning, this paper proposes multiple training strategies. The parallel training strategy balances the classification and segmentation tasks through the weight parameter α (optimal value 0.2), and the cyclic training strategy achieves stable multi-task optimization. These strategies effectively alleviate the optimization conflicts between tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0043] Figure 1 Schematic diagram of the network structure of the multi-task learning model provided by an embodiment of the present invention;
[0044] Figure 2 2 is a schematic diagram of a network structure of a shared encoder provided by an embodiment of the present invention;
[0045] Figure 3 Schematic diagram of the structure of the spatial feature extraction network provided by an embodiment of the present invention;
[0046] Figure 4 Schematic diagram of the structure of a multi-scale feature fusion network provided by an embodiment of the present invention;
[0047] Figure 5 Schematic diagram of different training strategies for multi-task training provided by an embodiment of the present invention;
[0048] Figure 6This is a system block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0049] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0050] First, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "exemplarily" is intended to present concepts in a concrete manner. In addition, in the embodiments of the present invention, the meaning of "and / or" can be both or either of the two.
[0051] First embodiment
[0052] This embodiment provides a method for visual recognition of solar photovoltaic panels in satellite images based on multi-task learning, and proposes a PVP visual recognition framework based on multi-task learning (MTL). The framework uses a shared encoder to efficiently extract multi-scale features and uses two task-specific decoders to obtain different results. The MTL framework can simultaneously realize the semantic segmentation of PVP in satellite images and the classification and recognition of background information. By introducing the MTL framework, the model can not only accurately segment PVP, but also effectively extract its background information, thereby providing more comprehensive and accurate data support for subsequent energy management and planning. At the same time, this embodiment also proposes a deep learning network PVFormer based on the MetaFormer framework to improve the model's spatial feature extraction ability and the fusion effect of multi-scale information. In addition, in response to the difficulties in training in MTL, a variety of different MTL training strategies are proposed to promote the coordination between classification and segmentation tasks, further improving the model's feature extraction ability and multi-task classification effect when processing high-resolution satellite images, and can better cope with photovoltaic panel recognition tasks in complex scenarios.
[0053] The method can be implemented by an electronic device. Specifically, the execution process of the method includes the following steps:
[0054] S1, obtain a solar photovoltaic panel satellite image dataset; the solar photovoltaic panel satellite images in the dataset not only contain solar photovoltaic panel semantic segmentation labels, but also contain background information of the solar photovoltaic panels;
[0055] S2, constructing a multi-task learning model; wherein the input of the multi-task learning model is a satellite image of a solar photovoltaic panel, and the output is a segmentation result of the solar photovoltaic panel area and a background type of the solar photovoltaic panel;
[0056] S3, training the multi-task learning model using the dataset;
[0057] S4, uses the trained multi-task learning model to realize the visual recognition of solar photovoltaic panels in satellite images.
[0058] It should be noted that multi-task learning (MTL) expands supervision information by sharing and combining knowledge between different tasks, thereby avoiding the overfitting of each task by traditional methods. The MTL framework proposed in this embodiment is designed to simultaneously solve the semantic segmentation task of solar photovoltaic panels in satellite images and the background classification task. Its structure is as follows Figure 1 As shown in the figure, it mainly includes a shared image encoder and two task-specific decoders. First, the input image will undergo random data augmentation, then it will be input into the shared encoder to extract multi-scale features, and then the features will be decoded using the segmentation decoder and classification decoder respectively to obtain the results of the corresponding tasks.
[0059] The shared encoder extracts rich feature representations from the input satellite imagery, gradually downsampling the input image to generate feature maps of varying resolutions, providing multi-level feature information for subsequent segmentation and classification tasks. To this end, this example designs the PVFormer network based on the MetaFormer framework as a shared encoder to facilitate inter-task integration. The decoder's task is to gradually upsample the low-resolution feature maps output by the encoder to the same resolution as the input image, enabling image- and pixel-level learning.
[0060] The segmentation task classifies the input image at the pixel level, enabling the model to accurately identify whether each pixel belongs to a photovoltaic panel. The segmentation decoder is the key module in MTL responsible for semantic segmentation tasks. Its main goal is to accurately segment the PVP area from the input high-resolution satellite image. The four-stage feature map output by the encoder will be input into the segmentation decoder, which uses three-stage processing to complete pixel-level semantic segmentation. The specific implementation method is as follows: the segmentation decoder first upsamples the low-dimensional feature map through bilinear interpolation or transposed convolution to restore it to the input image size, and then uses 1×1 convolution to compress the channel dimension, mapping the number of feature channels to the number of target categories. Finally, a probability distribution map is generated through pixel-by-pixel Softmax classification, and the maximum probability value is selected to determine the category of each pixel. This process can effectively identify the precise outline of the photovoltaic panel and establish accurate spatial constraints for the background classification task.
[0061] In this embodiment, the segmentation decoder adopts the FPN and Mask2Former network structures, and generates a high-resolution feature map for segmentation through upsampling and feature fusion.
[0062] The classification task requires the model to be able to identify the environmental background of the photovoltaic panel from a global perspective, so global feature extraction of the feature map is required. The classification decoder is the key module in MTL responsible for the background classification task. Its main goal is to classify the background of the input high-resolution satellite image and identify the specific type of the background. The four-stage feature map output by the encoder will also be input into the classification decoder. The classification decoder compresses the feature map into a feature vector of fixed length through global average pooling, and then passes it through the fully connected layer to predict the category of the photovoltaic panel background. The classification decoder of this embodiment focuses on background environment analysis based on the results of the shared encoder, integrates the global semantic information of the feature map through global average pooling, and combines the fully connected layer to build a multi-classification model, breaking through the local field of view limitation. It can identify a variety of background types such as roofs, shrubs, and water bodies, thereby significantly improving the model's ability to understand complex geographical environments.
[0063] Based on the above, the MTL framework of this embodiment achieves simultaneous image segmentation and classification through a shared encoder. This has the advantage of leveraging the correlation between different tasks to improve model performance and efficiency. The features extracted by the shared encoder can be applied to both segmentation and classification tasks, thereby reducing the demand for computing resources and improving the generalization ability of the model. Compared with single-task methods, the MTL method proposed in this embodiment has cross-task versatility, provides a more comprehensive image representation, and reduces the overall cost across multiple tasks.
[0064] Among them, multi-scale feature maps play a vital role in the MTL framework. From experience, classification models tend to use global features as a summary of the image, and segmentation models need to parse high-resolution features of each pixel. Therefore, the MTL framework requires an efficient shared encoder to extract multi-scale features. To this end, this embodiment designs a four-layer network structure PVFormer as a shared encoder based on the MetaFormer framework; its structure is as follows Figure 2 As shown in Figure 2, the number of channels and basic blocks in the four stages of PVFormer are [96, 192, 384, 768] and [3, 3, 12, 3] respectively. In the i-th stage, the input image or feature is first fed into the Stem layer or downsampling layer to adjust the feature resolution and convert the channel into C i Assuming the resolution of the input image is H×W, the feature resolutions of the four stages are H / 4×W / 4, H / 8×W / 8, H / 16×W / 16, and H / 32×W / 32 respectively. The adjusted features are then fed into the L[Number] which consists of the spatial feature extraction and multi-scale feature fusion modules.i In order to further extract and aggregate context information, the encoder finally passes the retained four-level output to the classification decoder and segmentation decoder to obtain the final segmentation and classification results. It should be noted that each basic block in a stage is connected in series, because the size of the input and output feature maps of the basic blocks are the same, and the downsampling of the features is completed through the downsampling layer. This section explains that each basic block is composed of a spatial feature extraction network SFEN and a multi-scale feature fusion network MSFFN. The specific combination is as follows Figure 2 As shown in the figure, the input feature map is first residually connected with the output of batch normalization and SFEN, and then continues to be residually connected with the output of batch normalization and MSFFN (the detailed structure of SFEN and MSFFN will be described in detail later).
[0065] In order to more effectively integrate the segmentation and classification tasks, this embodiment proposes the following Figure 3 The Spatial Feature Extraction Network (SFEN) shown in the figure. The features extracted by the pure convolutional structure are generated by different branches in the figure, with adaptive local perception features and wide-area static features. The design goal of the SFEN module is to extract spatial information of features from multiple angles to better serve segmentation and classification tasks. The main steps of SFEN are as follows:
[0066] First, the input feature X is first extracted by 1×1 convolution and 3×3 depth-separable convolution DWConv to generate the core feature layer Y. Then, the core feature layer Y is pooled by global average pooling GAP and the learnable parameter γ. s Perform spatial self-interaction (see Equation 2) and activation function GELU to adjust the local perception ability of the feature and generate adaptive local perception feature P. Then, use large-size deep convolution kernels to promote the interaction of long-distance features on the feature map to generate wide-area static features Q. Finally, the adaptive local perception feature P and the wide-area static feature Q are fused to generate the final feature representation Z.
[0067] Formulas (1) to (4) illustrate the above process. Through this design, SFEN effectively integrates and interacts the features of the classification task with the features of the segmentation task, effectively balancing the needs of the segmentation and classification tasks while enhancing the model's adaptability to complex scenarios.
[0068] Y = DWConv 3×3 (Conv 1×1 (X))#(1)
[0069] P=GELU(Y+γ s ⊙(Y-GAP(Y)))#(2)
[0070] Q=Norm(DWConv 13×13 (Y))#(3)
[0071] Z=Conv 1×1 (P+Q)#(4)
[0072] In order to more effectively extract the feature information of high-resolution satellite images and fuse the multi-scale information of feature channels, this embodiment proposes the following Figure 4 As shown in the Multi-scale Feature Fusion Networks (MSFFN), MSFFN uses parallel depthwise convolutions with kernel sizes of [1, 3, 5, 7], each depthwise convolution processes a quarter of the channels, rather than using a single 3×3 depthwise convolution. Figure 4 As shown in the figure, in the MSFFN, the input features first undergo 1×1 convolution, BN normalization, and GLUE activation to adjust the feature channels. The features are then divided into four features along the channel dimension, with each feature occupying one-quarter of the number of channels. The divided features are then fed into depthwise convolutions with different kernel sizes. The outputs of the four depthwise convolutions are merged along the channel dimension to the pre-division channel dimension and then residually connected with the pre-division features. Finally, after GLUE activation and BN normalization, the feature channels are adjusted using 1×1 convolution, BN normalization, and GLUE activation. This approach not only effectively captures multi-scale information across different channels but also enhances information interaction between different channels, effectively capturing feature information at different scales and improving the model's processing capabilities for high-resolution satellite imagery.
[0073] Furthermore, it should be noted that MTL allows the model to learn multiple related tasks simultaneously. By sharing model parameters, the model can learn the correlation between multiple tasks, thereby improving the model's generalization ability and learning efficiency. However, in MTL, there are both similarities and correlations between different tasks, as well as differences and potential conflicts. Although the segmentation task and the classification task share the same input image and feature extraction module, the segmentation task focuses on accurate pixel-level segmentation, while the classification task focuses more on the recognition of global background information. This difference may make it difficult for the model to balance the needs of the two tasks during training, and may even cause the optimization of one task to have a negative impact on the other task. In addition, the simple process of learning all tasks in a single training iteration requires a lot of computing resources. Therefore, not only is it necessary to design an efficient, stable, and economical training method, but how to learn the correlation between different tasks in model design while dealing with the differences and conflicts between them is an important challenge in multi-task learning.
[0074] In this example, a training strategy is designed by determining the order and frequency of learning two tasks. The order of tasks can be fixed or arbitrary, depending on the trade-off between stabilizing training and prioritizing a particular task. Furthermore, the frequency of tasks can be the same or biased. A strategy with the same frequency helps alleviate data imbalance and can set the learning frequency of more important or weaker tasks higher, thereby increasing the model's performance on those tasks.
[0075] Based on the above, this embodiment is designed as follows Figure 5 The following four multi-task training strategies are shown. The cyclic training strategy alternates between training the segmentation and classification tasks. In each training cycle, the model first trains on the segmentation task, then on the classification task, and so on. This cyclic training strategy is very stable and treats both tasks fairly, preventing one task from dominating the training process. However, the training process is relatively slow due to the frequent switching between the two tasks. The repeated sequence training strategy specifies a fixed training sequence before training begins, for example, first training the segmentation task several times continuously, then training the classification task several times continuously. This strategy allows for manual adjustment of the network's training bias towards different tasks. By changing the order and number of repetitions of the training sequence, the model's training direction can be fine-tuned. However, this strategy requires pre-setting the training sequence, which reduces flexibility. The weighted random training strategy assigns a probabilistic weight to each task and randomly selects a task for training in each training iteration based on these weights. This strategy allows the network to automatically adjust its training bias towards different tasks, better adapting to the complex relationships between different tasks. However, the training process is more random, which may lead to fluctuations in model performance. The above three methods are all trained by alternating single tasks. After the input image enters the shared encoder, the output feature map will enter the classification decoder and segmentation decoder in sequence, and multiple tasks are trained in sequence.
[0076] The parallel training strategy trains the segmentation and classification tasks simultaneously, achieving both segmentation and classification losses. This results in high training efficiency, but requires balancing the weights of the classification and segmentation tasks. Otherwise, the loss of one task may dominate the entire training process. Therefore, the present invention adds a weight parameter α to balance the weights of the classification and segmentation tasks. At the same time, the losses obtained by the two decoders are combined into a joint loss. Formula (5) shows the specific design of the loss function, which dynamically adjusts the weight ratio of the classification loss and the segmentation loss through the weight parameter.
[0077] L all =α·L class +(1-α)·L seg #(5)
[0078] In the formula, α is the weight parameter, Lall is the total loss, L class is the cross entropy loss for the classification task, L seg is the cross entropy loss for the segmentation task.
[0079] The multi-task training strategy enables the model to learn features related to the photovoltaic panel background while performing the segmentation task. By designing an appropriate learning strategy, specific needs can be met. In this regard, the multi-task training process proposed in this embodiment is flexible and scalable. MTL makes the PVP visual recognition task more convenient and efficient, providing more comprehensive information for solar resource management and monitoring.
[0080] In summary, this embodiment provides a method for visual recognition of solar photovoltaic panels in satellite images based on multi-task learning, proposes a shared encoder + dual decoder (segmentation + classification) architecture, and introduces SFEN and MSFFN based on the MetaFormer architecture to design a shared encoder. This encoder combines local perception with wide-area static features to enhance the model's adaptability to complex scenes. Through parallel processing of multi-scale convolution kernels, the multi-scale information fusion capability is improved, and the intra-class diversity and background interference problems of PVP in high-resolution satellite images are solved. At the same time, this embodiment also designs four training strategies (loop, repeated sequence, weighted random, parallel), and optimizes task balance by dynamically adjusting the loss weight α parameter.
[0081] This framework can simultaneously achieve pixel-level segmentation of PVP and background environment classification. While traditional methods focus solely on PVP segmentation, this framework extracts background information (such as rooftops, water areas, and cultivated land) through multi-task collaborative optimization, providing more comprehensive data support for energy planning. Furthermore, by balancing the weights of segmentation and classification tasks to avoid task conflicts, an efficient shared feature extraction mechanism is designed. This significantly improves the segmentation accuracy and classification accuracy of the model in complex backgrounds. Extensive experiments have shown that this framework can accurately and quickly segment PVP from satellite imagery and distinguish its background, thereby promoting the development and utilization of renewable energy.
[0082] Second embodiment
[0083] This embodiment provides an electronic device, such as Figure 6 As shown, the electronic device includes: a processor and a memory; wherein the processor and the memory can be connected via a communication bus; the memory stores at least one instruction, which is loaded and executed by the processor to implement the method of the first embodiment described above. In addition, the electronic device may also include a transceiver; the processor and the transceiver can be connected via a communication bus, and the transceiver is used to communicate with other devices.
[0084] Next, combine Figure 6A detailed introduction to the various components of the electronic device is given below:
[0085] Among them, the processor is the control center of the electronic device, and the electronic device may include multiple processors, each of which may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here can be a processor or a general term for multiple processing elements. For example, the processor is one or more central processing units (CPUs), or other general-purpose processors, application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement an embodiment of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor can perform various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.
[0086] In a specific implementation, as an embodiment, the processor may include one or more CPUs, such as Figure 6 The CPU0 and CPU1 shown in FIG are, of course, only exemplary.
[0087] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment and will not be repeated here.
[0088] Optionally, the memory may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and accessed through the interface circuit ( Figure 6 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.
[0089] The transceiver may include a receiver and a transmitter ( Figure 6 The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function. The transceiver can be integrated with the processor or exist independently and communicate with the electronic device through the interface circuit ( Figure 6 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.
[0090] In addition, it should be noted that Figure 6 The structure of the electronic device shown in the figure does not constitute a limitation on the device. The actual device may include more or fewer components than shown, or may combine certain components, or arrange the components differently. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment can refer to the technical effects described in the first embodiment above, and therefore will not be repeated here.
[0091] Third embodiment
[0092] This embodiment provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the method of the first embodiment described above. The computer-readable storage medium may be a ROM, random access memory, CD-ROM, magnetic tape, floppy disk, or optical data storage device. The instructions stored therein can be loaded by a processor in a terminal to execute the method described above.
[0093] Furthermore, it should be noted that the present invention may be provided as a method, apparatus, or computer program product. Therefore, embodiments of the present invention may take the form of a fully or partially hardware embodiment, a fully or partially software embodiment, or an embodiment combining software and hardware aspects. Furthermore, when implemented using software, embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product comprises one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are fully or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired connection (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium. The semiconductor medium may be a solid state drive.
[0094] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0095] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0096] It should also be noted that, in this document, relational terms such as first and second are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between these entities or operations. The terms "include," "comprises," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. Without further limitation, an element defined by the phrase "comprising a..." does not preclude the presence of other identical elements in the process, method, article, or terminal device comprising the element. In addition, the term "and / or" is merely a description of an associative relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: the presence of A alone, the presence of A and B simultaneously, or the presence of B alone, where A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding. "At least one" means one or more, and "more" means two or more. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0097] In addition, it can be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0098] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0099] In the several embodiments provided herein, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of functional modules / units is merely a logical functional division. In actual implementation, other division methods may be used, such as multiple units or components being combined or integrated into another device, or some features being ignored or not implemented. Furthermore, the coupling or direct coupling or communication connection shown or discussed between each other may be through some interface, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms. Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, i.e., they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs. In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0100] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0101] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention. It should be noted that, although preferred embodiments of the present invention have been described, those skilled in the art, once understanding the basic inventive concepts of the present invention, may make various improvements and modifications without departing from the principles of the present invention. Such improvements and modifications should also be considered within the scope of protection of the present invention. Therefore, the appended claims are intended to be interpreted as covering the preferred embodiments and all variations and modifications that fall within the scope of the embodiments of the present invention.
Claims
1. A method for visual recognition of solar photovoltaic panels in satellite images based on multi-task learning, characterized in that: The multi-task learning-based solar photovoltaic panel visual recognition method in satellite images includes: Obtain a solar photovoltaic panel satellite image dataset; wherein the solar photovoltaic panel satellite images in the dataset include not only solar photovoltaic panel semantic segmentation labels but also background information of the solar photovoltaic panels; Constructing a multi-task learning model; wherein the input of the multi-task learning model is a satellite image of a solar photovoltaic panel, and the output is a segmentation result of the solar photovoltaic panel area and the background type of the solar photovoltaic panel; Training the multi-task learning model using the dataset; Use the trained multi-task learning model to achieve visual recognition of solar photovoltaic panels in satellite imagery.
2. The method for visually identifying solar photovoltaic panels in satellite images based on multi-task learning according to claim 1, wherein: The multi-task learning model includes a shared encoder, a segmentation decoder, and a classification encoder; The shared encoder is used to extract multi-scale features from satellite images of solar photovoltaic panels input to the model; The segmentation decoder is used to segment the area of the solar photovoltaic panel from the satellite image of the solar photovoltaic panel input to the model based on the multi-scale features extracted by the shared encoder; The classification decoder is used to classify the background of the solar photovoltaic panel satellite image input to the model based on the multi-scale features extracted by the shared encoder, and identify the type of the background.
3. The method for visually identifying solar photovoltaic panels in satellite images based on multi-task learning according to claim 2, wherein: The shared encoder is a four-layer network structure designed using the MetaFormer framework. The first layer of the shared encoder consists of a Stem layer and a basic block, and each of the second to fourth layers consists of a downsampling layer and a basic block. The first layer of the shared encoder takes as input a satellite image of solar photovoltaic panels, and the input of each subsequent layer is the feature data output by the previous layer. The Stem layer and the downsampling layer are used to adjust the feature resolution and realize channel dimension conversion; the basic block is used to further extract and aggregate context information of the features adjusted by the Stem layer or the downsampling layer.
4. The method for visually identifying solar photovoltaic panels in satellite images based on multi-task learning according to claim 3, wherein: The basic block includes a spatial feature extraction network and a multi-scale feature fusion network; The spatial feature extraction network is used to extract spatial feature information of input features from multiple angles; The multi-scale feature fusion network is used to capture multi-scale information of different channels of feature information extracted by the spatial feature extraction network, enhance information interaction between different channels, and fuse multi-scale information of feature channels.
5. The method for visually identifying solar photovoltaic panels in satellite images based on multi-task learning according to claim 4, wherein: The process of extracting spatial feature information of input features from multiple angles by the spatial feature extraction network includes: First, the input feature X is extracted by 1×1 convolution and 3×3 depth-separable convolution DWConv to generate the core feature layer Y; then, the core feature layer Y is pooled by global average pooling GAP and γ s Spatial self-interaction and activation function GELU adjust the local perception ability of the feature to generate an adaptive local perception feature P. Then, a large-size deep convolution kernel is used to promote the interaction of long-distance features on the feature map to generate a wide-area static feature Q. Finally, the adaptive local perception feature P and the wide-area static feature Q are fused to generate the final feature representation Z. The size of the large-size deep convolution kernel is 13x13.
6. The method for visually identifying solar photovoltaic panels in satellite images based on multi-task learning according to claim 4, wherein: The multi-scale feature fusion network uses parallel depthwise convolutions with kernel sizes of [1, 3, 5, 7], and each depthwise convolution processes a quarter of the channels.
7. The method for visually identifying solar photovoltaic panels in satellite images based on multi-task learning according to claim 2, wherein: The segmentation decoder segments the region of the solar photovoltaic panels from the satellite image of the solar photovoltaic panels input to the model based on the multi-scale features extracted by the shared encoder, including the following steps: The segmentation decoder first upsamples the multi-scale features extracted by the shared encoder through bilinear interpolation or transposed convolution to restore them to the input image size, then uses 1×1 convolution to compress the channel dimension, mapping the number of feature channels to the number of target categories. Finally, a probability distribution map is generated through pixel-by-pixel Softmax classification, and the maximum probability value is selected to determine the category of each pixel to segment the area of the solar photovoltaic panel.
8. The method for visually identifying solar photovoltaic panels in satellite images based on multi-task learning according to claim 2, wherein: The classification encoder classifies the background of the solar photovoltaic panel satellite image input to the model based on the multi-scale features extracted by the shared encoder. The process of identifying the type of background includes: The classification decoder compresses the multi-scale features extracted by the shared encoder into a fixed-length feature vector through global average pooling, and then predicts the background category of the satellite image of solar photovoltaic panels through a fully connected layer.
9. The method for visually identifying solar photovoltaic panels in satellite images based on multi-task learning according to claim 1, wherein: The training strategies for training the multi-task learning model using the data set include: a cyclic training strategy, a repeated sequence training strategy, a weighted random training strategy, and a parallel training strategy; The cyclic training strategy alternately performs training for the segmentation task and the classification task. In each training cycle, the model first performs training for the segmentation task and then for the classification task, and so on. The repeated sequence training strategy specifies a fixed training sequence before the training begins, first training one task several times continuously, and then training another task several times continuously; The weighted random training strategy assigns a probability weight to each task, and in each training iteration, randomly selects a task for training based on these weights; The parallel training strategy trains the segmentation task and the classification task simultaneously.
10. The method for visually identifying solar photovoltaic panels in satellite images based on multi-task learning according to claim 9, wherein: When training the model using the parallel training strategy, the loss function is: L all =α·L class +(1-a)·L seg Among them, α is the preset weight parameter; L all is the total loss; L class is the cross entropy loss for the background classification task; L seg Cross entropy loss for the solar photovoltaic panel segmentation task.