A Photovoltaic Recognition Method and Related Device Based on Visual Feature Constraints
By combining the adversarial training of Upernet segmentation network model and feature constraint module, the accuracy and reliability problems of traditional photovoltaic panel recognition methods are solved, and a more accurate photovoltaic panel recognition effect is achieved.
Patent Information
- Application Number
- CN202411533718.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-10-31
AI Technical Summary
The traditional photovoltaic panel recognition method has inaccurate and unreliable recognition results due to light changes and background complexity. The general deep learning segmentation model fails to make full use of the unique visual characteristics of the photovoltaic panel, which limits the improvement of segmentation effect.
The photovoltaic recognition method based on visual feature constraints is adopted, combined with the Upernet segmentation network model and the feature constraint module based on adversarial autoencoder, the Upernet segmentation network model is fully learned through adversarial training to improve the recognition accuracy.
It effectively improves the accuracy and reliability of photovoltaic panel recognition. Through adversarial training, the model is constantly approaching the real photovoltaic panel outline. The feature constraint module guides the model to pay attention to color, texture and shape characteristics to achieve more accurate segmentation results.
Smart Images

Figure CN119540751B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of photovoltaic panel recognition, and in particular to a photovoltaic recognition method and related device based on visual feature constraints. Background Art
[0002] With the continuous growth of the demand for renewable energy, photovoltaic power generation, as a clean and sustainable energy form, has received extensive attention and rapid development. Photovoltaic panels are the core components of a photovoltaic power generation system, and their performance and layout directly affect the efficiency of the photovoltaic power generation system. In order to improve the maintenance efficiency and operation reliability of photovoltaic panels, it is particularly important to accurately identify and segment photovoltaic panels in various application scenarios.
[0003] The segmentation of photovoltaic panels is an important issue in computer vision, aiming to identify and extract the photovoltaic panel area from a complex image background. When traditional image processing methods are used to process the segmentation of photovoltaic panels, the recognition results of photovoltaic panels are often inaccurate and unreliable due to reasons such as illumination changes and complex backgrounds.
[0004] With the rapid development of deep learning technology, segmentation models based on deep learning have demonstrated powerful performance in computer vision tasks, and photovoltaic panel segmentation has also benefited greatly from this. However, the visual features of photovoltaic panels in images have their uniqueness, such as uniform texture, significant color differences, and regular shapes. In view of the above visual features of photovoltaic panels, traditional general deep learning segmentation models often fail to fully utilize these characteristics when applied to the photovoltaic panel segmentation task, thus limiting the further improvement of the segmentation effect, and the final obtained recognition results of photovoltaic panels are also not very accurate. Summary of the Invention
[0005] The purpose of the present application is to provide a photovoltaic recognition method and related device based on visual feature constraints, which can effectively improve the accuracy of identifying photovoltaic panels.
[0006] To achieve the above purpose, the present application provides the following solutions:
[0007] In a first aspect, the present application provides a photovoltaic recognition method based on visual feature constraints, including the following steps:
[0008] Obtain a remote sensing image dataset; the remote sensing image dataset includes a number of remote sensing images and corresponding ground truth labels, each remote sensing image corresponding to one ground truth label, and the ground truth label being used to characterize whether the ground surface of the remote sensing image contains a photovoltaic panel and the location area of the photovoltaic panel.
[0009] Construct a Upernet segmentation network model; the Upernet segmentation network model takes the remote sensing image as input and the image segmentation result corresponding to the remote sensing image as output. The Upernet segmentation network model includes a feature pyramid network, a feature aggregation module, and a semantic segmentation detection head. The feature pyramid network is used to extract and fuse multi-scale features of the remote sensing image. The feature aggregation module is used to perform integration processing on the feature maps corresponding to the fused multi-scale features. The semantic segmentation detection head is used to output the image segmentation result according to the integrated feature maps. The image segmentation result is the result of predicting the area where the photovoltaic panels are segmented in the remote sensing image.
[0010] Construct a feature constraint module based on an adversarial autoencoder; the feature constraint module includes a color feature constraint module, a texture feature constraint module, and a shape feature constraint module. The color feature constraint module, the texture feature constraint module, and the shape feature constraint module all include an autoencoder, a decoder, and a discriminator. The autoencoder is used to encode the input feature vector into a low-dimensional manifold. The decoder is used to decode and reconstruct the input vector. The discriminator is used to distinguish the difference between the image segmentation result corresponding to the remote sensing image predicted by the Upernet segmentation network model and the ground truth label of the remote sensing image.
[0011] Use the remote sensing image dataset to perform adversarial training on the feature constraint module and the Upernet segmentation network model to obtain an optimal Upernet segmentation network model.
[0012] Obtain a to-be-tested remote sensing image of the target area and input the to-be-tested remote sensing image into the optimal Upernet segmentation network model to predict the image segmentation result corresponding to the to-be-tested remote sensing image of the target area.
[0013] Optionally, after the step of obtaining the remote sensing image dataset, the photovoltaic recognition method based on visual feature constraints further includes the following steps:
[0014] Preprocess each remote sensing image in the remote sensing image dataset respectively to obtain a preprocessed remote sensing image dataset. The preprocessed remote sensing image dataset is used to perform adversarial training on the feature constraint module and the Upernet segmentation network model to obtain the optimal Upernet segmentation network model.
[0015] Optionally, preprocessing each remote sensing image in the remote sensing image dataset respectively to obtain a preprocessed remote sensing image dataset specifically includes the following steps:
[0016] Perform clipping processing on each of the remote sensing images in the remote sensing image dataset to obtain a clipped remote sensing image dataset; the clipped remote sensing image dataset includes a number of clipped remote sensing images.
[0017] Perform data augmentation processing on each of the clipped remote sensing images in the clipped remote sensing image dataset to obtain the preprocessed remote sensing image dataset; the data augmentation processing includes image random rotation processing, and the image random rotation processing includes image random horizontal flipping processing and image random vertical flipping processing.
[0018] Optionally, constructing the color feature constraint module specifically includes the following steps:
[0019] For the color feature constraint module, respectively accept the predicted mask and the ground truth label of the Upernet segmentation network model, and mask the predicted mask and the ground truth label with the corresponding remote sensing image respectively to determine the target region image.
[0020] Calculate the color histogram of the target region according to the target region image.
[0021] Based on the color histogram of the target region, use the autoencoder of the color feature constraint module to respectively accept the color histogram feature vector of the positive region predicted by the Upernet segmentation network model and the color histogram feature vector of the real positive region, and map each color feature in the color histogram feature vector to a low-dimensional vector; the positive region is the region where the photovoltaic panel is located.
[0022] Use the decoder of the color feature constraint module to capture the color-related features in the low-dimensional vector while discarding the color-unrelated features, and reconstruct the input color features.
[0023] Use the discriminator of the color feature constraint module to distinguish whether the input color features come from the image segmentation result predicted by the Upernet segmentation network model or from the real remote sensing image sample.
[0024] Optionally, constructing the texture feature constraint module specifically includes the following steps:
[0025] For the texture feature constraint module, use the LBP algorithm to calculate the local binary pattern coding of the remote sensing image.
[0026] Determine the LBP coding map corresponding to the remote sensing image according to the local binary pattern coding of the remote sensing image.
[0027] Mask the LBP encoded image with the predicted mask of the corresponding Upernet segmentation network model and the ground truth label to obtain the LBP values of the target regions.
[0028] Use the LBP values of the target regions as texture feature vectors and feed them into the autoencoder of the texture feature constraint module respectively. Then, use the discriminator of the texture feature constraint module to distinguish the texture feature differences during the adversarial training process.
[0029] Optionally, constructing the shape feature constraint module specifically includes the following steps:
[0030] For the shape feature constraint module, input the image segmentation result predicted by the Upernet segmentation network model and the ground truth label into the autoencoder of the shape feature constraint module. Use the adversarial training method to train the autoencoder so that the autoencoder of the shape feature constraint module encodes two types of shapes and captures the differences between the two types of shapes. At the same time, use the Upernet segmentation network model to oppose the autoencoder of the shape feature constraint module so that the autoencoder of the shape feature constraint module cannot capture the differences between the two types of shapes; where the two types of shapes refer to the shapes of the binary maps of the image segmentation result predicted by the Upernet segmentation network model and the binary map of the ground truth label.
[0031] Optionally, the expression of the objective loss function used in the adversarial training process is:
[0032]
[0033] where, L all represents the total objective loss function of the adversarial training, L seg , L s , L t , L c represent the segmentation loss, shape loss, texture loss, and color loss respectively; λ 1 , λ 2 , λ 3 are hyperparameters; G and D represent the Upernet segmentation network model and the feature constraint module respectively.
[0034] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the photovoltaic recognition method based on visual feature constraints described in any one of the above.
[0035] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned photovoltaic identification methods based on visual feature constraints.
[0036] In a fourth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned photovoltaic identification methods based on visual feature constraints.
[0037] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0038] The present application provides a photovoltaic identification method and related device based on visual feature constraints, which combines an Upernet segmentation network model with a feature constraint module based on an adversarial autoencoder, and constructs an Upernet segmentation network model and a feature constraint module respectively, wherein the feature constraint module includes a color feature constraint module, a texture feature constraint module, and a shape feature constraint module, so that when adversarial training is performed between the feature constraint module and the Upernet segmentation network model, the Upernet segmentation network model fully learns the features of the color, texture, and shape of the photovoltaic panels in the remote sensing image. The Upernet segmentation network model is continuously approached to the real photovoltaic panel contour by adversarial training, and the feature constraint module attempts to distinguish the difference between the predicted results and the real results of the Upernet segmentation network model, and guides the Upernet segmentation network model to focus on the color features, texture features, and shape features of the photovoltaic panels. After the adversarial training is completed, the optimal Upernet segmentation network model can be obtained. In practical applications, the remote sensing image to be tested in the target area is obtained and input into the optimal Upernet segmentation network model, so that a more accurate and reliable image segmentation result can be predicted, which effectively improves the accuracy of identifying photovoltaic panels. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0040] Figure 1 A schematic flow chart of a photovoltaic identification method based on visual feature constraints provided in one embodiment of the present application.
[0041] Figure 2 A flowchart of an overall network architecture provided for an embodiment of the present application.
[0042] Figure 3 The structural schematic diagram of the UperNet segmentation network model provided by an embodiment of the present application.
[0043] Figure 4 The structural schematic diagram of the color feature constraint module provided by an embodiment of the present application.
[0044] Figure 5 The structural schematic diagram of the texture feature constraint module provided by an embodiment of the present application.
[0045] Figure 6 The structural schematic diagram of the shape feature constraint module provided by an embodiment of the present application. Detailed implementation manners
[0046] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0047] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0048] As Figure 1 and Figure 2 shown, this embodiment provides a photovoltaic recognition method based on visual feature constraints. The purpose of this method is to accurately identify the outline of photovoltaic panels from remote sensing images and segment the area where the photovoltaic panels are located. The method specifically includes the following steps:
[0049] Step S1, obtain a remote sensing image dataset.
[0050] In this embodiment, the remote sensing image dataset includes a number of remote sensing images and corresponding ground truth labels. Each remote sensing image corresponds to one ground truth label, and the ground truth label is used to characterize whether the ground surface of the remote sensing image contains photovoltaic panels and the area where the photovoltaic panels are located.
[0051] Step S2, construct a Upernet (Unified Perceptual Parsing Network) segmentation network model.
[0052] In this embodiment, the Upernet segmentation network model takes the remote sensing image as input and outputs the image segmentation result corresponding to the remote sensing image. The Upernet segmentation network model includes a Feature Pyramid Network (FPN), a feature aggregation module, and a semantic segmentation detection head. The Feature Pyramid Network is used to extract and fuse multi-scale features of the remote sensing image. The feature aggregation module is used to perform integration processing on the feature maps corresponding to the fused multi-scale features. The semantic segmentation detection head is used to output the image segmentation result according to the feature maps after the integration processing. The image segmentation result is the result of predicting the region where the photovoltaic panel is segmented in the remote sensing image.
[0053] Step S3: Construct a feature constraint module based on an adversarial autoencoder.
[0054] In this embodiment, the feature constraint module includes a color feature constraint module, a texture feature constraint module, and a shape feature constraint module. The color feature constraint module, the texture feature constraint module, and the shape feature constraint module all include an autoencoder, a decoder, and a discriminator. The autoencoder is used to encode the input feature vector into a low-dimensional manifold. The decoder is used to decode and reconstruct the input vector. The discriminator is used to distinguish the difference between the image segmentation result corresponding to the remote sensing image predicted by the Upernet segmentation network model and the ground truth label of the remote sensing image.
[0055] Step S4: Use the remote sensing image dataset to perform adversarial training on the feature constraint module and the Upernet segmentation network model to obtain an optimal Upernet segmentation network model.
[0056] Step S5: Obtain a to-be-detected remote sensing image of the target area, and input the to-be-detected remote sensing image into the optimal Upernet segmentation network model to predict the image segmentation result corresponding to the to-be-detected remote sensing image of the target area.
[0057] In this embodiment, after step S1 of obtaining the remote sensing image dataset, the following steps are further included:
[0058] Step S1.5: Preprocess each remote sensing image in the remote sensing image dataset to obtain a preprocessed remote sensing image dataset. The preprocessed remote sensing image dataset is used to perform adversarial training on the feature constraint module and the Upernet segmentation network model to obtain the optimal Upernet segmentation network model. Specifically, it includes the following steps:
[0059] Step S1.51: Perform cropping processing on each remote sensing image in the remote sensing image dataset to obtain a cropped remote sensing image dataset. The cropped remote sensing image dataset includes a number of cropped remote sensing images.
[0060] Step S1.52: Perform data augmentation processing on each of the cropped remote sensing image data sets to obtain the preprocessed remote sensing image data set; the data augmentation processing includes image random rotation processing, and the image random rotation processing includes image random horizontal flipping processing and image random vertical flipping processing.
[0061] In this embodiment, step S3 constructs a feature constraint module based on an adversarial autoencoder, which specifically includes the following steps:
[0062] Step S31: For the color feature constraint module, respectively receive the predicted mask and the true label of the Upernet segmentation network model, and mask the predicted mask and the true label with the corresponding remote sensing image to determine the target region image.
[0063] Step S32: Calculate the color histogram of the target region according to the target region image.
[0064] Step S33: Based on the color histogram of the target region, use the autoencoder of the color feature constraint module to respectively receive the color histogram feature vector of the positive region predicted by the Upernet segmentation network model and the color histogram feature vector of the true positive region, and map each color feature in the color histogram feature vector to a low-dimensional vector; the positive region is the region where the photovoltaic panel is located.
[0065] Step S34: Use the decoder of the color feature constraint module to capture the color-related features in the low-dimensional vector while discarding the color-unrelated features, and reconstruct the input color features.
[0066] Step S35: Use the discriminator of the color feature constraint module to distinguish whether the input color features come from the image segmentation result predicted by the Upernet segmentation network model or from the true remote sensing image sample.
[0067] Step S36: For the texture feature constraint module, use the LBP (Local Binary Patterns) algorithm to calculate the local binary pattern encoding of the remote sensing image.
[0068] Step S37: Determine the LBP encoding map corresponding to the remote sensing image according to the local binary pattern encoding of the remote sensing image.
[0069] Step S38: Mask the LBP encoding map with the predicted mask and the true label of the corresponding Upernet segmentation network model to obtain the LBP value of the target region.
[0070] Step S39: Use the LBP values of the target region as the texture feature vectors, and send them into the autoencoder of the texture feature constraint module respectively. Then, use the discriminator of the texture feature constraint module to distinguish the texture feature differences during the adversarial training process.
[0071] Step S40: For the shape feature constraint module, input the image segmentation result predicted by the Upernet segmentation network model and the true label into the autoencoder of the shape feature constraint module. Use the adversarial training method to train the autoencoder so that the autoencoder of the shape feature constraint module encodes two types of shapes and captures the differences between the two types of shapes. At the same time, use the Upernet segmentation network model to oppose the autoencoder of the shape feature constraint module so that the autoencoder of the shape feature constraint module cannot capture the differences between the two types of shapes; where the two types of shapes refer to the shapes of the binary maps of the image segmentation result predicted by the Upernet segmentation network model and the binary map of the true label.
[0072] To make the technical solution of this embodiment clearer, the following will take an example to illustrate the specific implementation process of the technical solution of this embodiment in detail. It specifically includes the following implementation steps:
[0073] Step1, Data collection and preprocessing.
[0074] S1-1: Collect the training data set, including the remote sensing image data set and the corresponding manually annotated sample images. In this embodiment, remote sensing images with a spatial resolution of 0.15m from Google Earth (Google Maps) and having three bands of red, green, and blue are used.
[0075] S1-2: Perform image cropping processing. In this embodiment, a sliding window is used to crop the remote sensing image into 256×256 pixel tiles.
[0076] S1-3: Perform data augmentation, including random rotation of the image, random horizontal flipping, and random vertical flipping, to enhance the generalization ability of the model.
[0077] Step2, Construct the Upernet segmentation network model. The structure of the Upernet segmentation network model is as Figure 3 shown. The Upernet segmentation network model uses a feature pyramid network for multi-scale feature extraction and fusion, integrates feature maps at different levels to obtain a richer feature representation. An additional feature aggregation module is used to further integrate and process the feature maps. Finally, the segmentation result is output through the semantic segmentation detection head.
[0078] Step 3: Construct a feature constraint module based on an adversarial autoencoder.
[0079] S3-1: The feature constraint module includes a color feature constraint module, a texture feature constraint module, and a shape feature constraint module. Each feature constraint module includes an autoencoder, a decoder, and a discriminator. The autoencoder encodes the input feature vector into a low-dimensional vector, the decoder reconstructs the input vector, and the discriminator is used to distinguish the difference between the real sample and the prediction result of the Upernet segmentation network model, that is, the difference between the image segmentation result corresponding to the remotely sensed image predicted by the Upernet segmentation network model and the real label of the remotely sensed image.
[0080] In this embodiment, considering the characteristics of different features, the autoencoders in the color feature constraint module and the texture feature constraint module are composed of multiple linear layers, and the autoencoder in the shape feature constraint module is composed of multiple convolutional layers. The steps to construct each feature constraint module are as follows:
[0081] S3-2: For the color feature constraint module, accept the prediction mask and the real label of the Upernet segmentation network model, and perform masking with the corresponding remotely sensed image to obtain the target region image, and calculate the target region color histogram. Construct an autoencoder, which is a trainable neural network that accepts the color histogram feature vectors of the predicted positive region and the real region of the Upernet segmentation network model respectively, maps each color feature to a low-dimensional vector, captures its significant features while discarding irrelevant features, and reconstructs the input color feature as much as possible. As Figure 4 shown, in the color feature constraint module, an autoencoder is constructed using multiple linear layers. The discriminator is used to distinguish whether the input feature comes from the prediction of the Upernet segmentation network model or a real sample.
[0082] S3-3: For the texture feature constraint module, as Figure 5 shown, first use the LBP algorithm to calculate the local binary pattern encoding of the remotely sensed image to construct the target region texture feature vector. The formula is as follows:
[0083]
[0084] where LBP(·) is the local binary pattern encoding, (x c , y c ) represents the central pixel of the neighborhood window, i p is the value of the central pixel of the neighborhood window, i c is the value of other pixels in the neighborhood, p is the number of sampling points in the neighborhood, and S(x) is the sign function, which is expressed as:
[0085]
[0086] Then, mask the LBP encoded map of the remote sensing image with the prediction mask and the ground truth label of the corresponding Upernet segmentation network model to obtain the LBP values of the target region, which are used as the texture feature vectors.
[0087] Send the LBP values of the target region into an autoencoder respectively. This autoencoder consists of multiple linear layers, learns the texture features of the photovoltaic panel images, and makes it difficult for the discriminator to distinguish the differences in texture features between them through adversarial training, so as to optimize the parameters of the Upernet segmentation network model.
[0088] S3-4: For the shape feature constraint module, as Figure 6 shown, directly send the prediction results and the ground truth labels of the Upernet segmentation network model into the autoencoder, try to reconstruct the input shape as much as possible, train the autoencoder through an adversarial training scheme, force the autoencoder to better encode the two types of shapes and capture their subtle differences, and at the same time encourage the Upernet segmentation network model to deceive the autoencoder into being unable to capture the subtle differences. Since the input is in the form of an image, this shape feature constraint module consists of multiple convolutional layers.
[0089] Step4, train the Upernet segmentation network model and the feature constraint module. Specifically, it includes the following steps:
[0090] S4-1: The training method adopts adversarial training, and the Upernet segmentation network model and the feature constraint module are alternately optimized. The Upernet segmentation network model tries to minimize the difference between the segmentation result and the ground truth label, while the feature constraint module tries to learn the encoding to maximize their distance. Therefore, the total objective loss function is defined as:
[0091]
[0092] where, L all represents the total objective loss function of adversarial training, L seg , L s , L t , L c represent the segmentation loss, shape loss, texture loss, and color loss respectively. λ 1 , λ 2 , λ 3 are hyperparameters, which are used to control the weights of the shape loss, texture loss, and color loss respectively. G and D represent the Upernet segmentation network model and the autoencoder of the feature constraint module respectively. During the training process of G, the results predicted by G will gradually approach the ground truth label, and the feature constraint module encourages the autoencoder to encode the subtle differences between them.
[0093] S4-2: The segmentation loss uses cross-entropy loss and is defined as:
[0094]
[0095] Among them, L seg represents the segmentation loss, which uses cross-entropy loss. q i is the probability of the model predicting the positive class, and p i is the true label of the i-th sample.
[0096] S4-3: The autoencoder has two loss terms, the reconstruction loss and the adversarial loss.
[0097] The reconstruction loss is defined as:
[0098]
[0099] Among them, L rec is the reconstruction loss, x represents the input feature vector, y represents its corresponding true label, G(x) represents the predicted mask of the segmentation module, and D(y) and D(G(x)) respectively represent the reconstruction results of the true label and the predicted mask.
[0100] The adversarial loss encourages the discriminator to distinguish different inputs, thereby guiding the Upernet segmentation network model to deceive the autoencoder. For the shape feature constraint module, texture feature constraint module, and color feature constraint module, they are respectively the adversarial shape loss, texture loss, and color loss. It is defined as:
[0101]
[0102] Among them, L d represents the feature constraint loss, which refers to the shape loss, texture loss, and color loss in the shape feature constraint module, texture feature constraint module, and color feature constraint module respectively. E(y) and E(G(x)) are the encodings of the true label and the predicted mask.
[0103] During training, first fix D and optimize G with the following loss:
[0104]
[0105] Among them, L G represents the loss in two aspects of the segmentation loss and the feature constraint loss of the Upernet segmentation network model.
[0106] Optimizing L G will encourage G to output a photovoltaic panel prediction result consistent with the ground truth label. Then, when G is fixed, optimize D with the following loss:
[0107]
[0108] Among them, λ is a hyperparameter used to control the weights of the shape loss, texture loss, and color loss, and ω is a hyperparameter used to make the training of the autoencoder more stable.
[0109] During the entire training process, the Upernet segmentation network model will continuously approach the true contour of the photovoltaic panel. When the feature constraint module cannot distinguish the difference between the prediction result and the true result, it no longer provides effective supervision. At this time, the Upernet segmentation network model can independently output a better segmentation result.
[0110] In this embodiment, the output result of the UperNet segmentation network model is input into the feature constraint module. The feature constraint module (including the color feature constraint module, texture feature constraint module, and shape feature constraint module) simultaneously receives the output result of the UperNet segmentation network model, the remote sensing image in the training set, and the true label. The feature constraint module is only used in the training stage. During training, the UperNet segmentation network model and the feature constraint module alternately update parameters. The UperNet segmentation network model first outputs the result, and then the result is sent to the feature constraint module. The feature constraint module extracts texture, color, and shape features respectively, and distinguishes the differences in texture, color, and shape between the output result of the UperNet segmentation network model and the corresponding region of the true label. The texture loss, color loss, and shape loss are used to measure the size of the difference, and the loss value is attached to the segmentation loss to control the optimization direction of the next iteration of the UperNet segmentation network model.
[0111] In this embodiment, the texture loss, color loss, and shape loss are weighted and added to the segmentation loss of the UperNet segmentation network model. That is, the overall objective function. Finally, the UperNet segmentation network model is jointly affected by the segmentation loss, texture loss, color loss, and shape loss. The segmentation loss is calculated by the output result of the UperNet segmentation network model and the true label. The texture loss is calculated in the texture feature constraint module, the color loss is calculated in the color feature constraint module, and the shape loss is calculated in the shape feature constraint module.
[0112] Step5, save the parameters of the optimal segmentation model and perform prediction.
[0113] Remove the feature constraint module, save the trained UperNet segmentation network model, and apply it to the recognition of large-scale remote sensing images. The purpose is to further improve the segmentation accuracy without increasing the model parameters of the UperNet segmentation network model during the prediction application stage, and to improve the accuracy and robustness of the UperNet segmentation network model for identifying photovoltaic panels on the basis of effectively controlling the calculation cost. This has high practicality for the recognition task of large-scale data.
[0114] The present application provides a photovoltaic recognition method based on visual feature constraints, including remote sensing image preprocessing; constructing an UperNet segmentation network model; constructing a color feature constraint module, a texture feature constraint module, and a shape feature constraint module based on an autoencoder; extracting the color features of photovoltaic panels in the remote sensing image using a color histogram, extracting the texture features of photovoltaic panels in the remote sensing image using local binary pattern coding, and extracting the shape features of photovoltaic panels in the remote sensing image using a binary image; through adversarial training, the UperNet segmentation network model continuously approaches the true contour of the photovoltaic panel, while the feature constraint module attempts to distinguish the difference between the prediction result and the true result, guiding the UperNet segmentation network model to focus on the color, texture, and shape features of the photovoltaic panel; finally, after the adversarial training ends, the feature constraint module is removed, and the optimized optimal UperNet segmentation network model is applied to the recognition of photovoltaic panels, so as to accurately recognize and segment the photovoltaic panel area, and finally accurately demarcate the area where the photovoltaic panel is located from the remote sensing image. It can achieve targeted optimization of the task of photovoltaic panel recognition without increasing the number of parameters of the UperNet segmentation network model, and effectively improve the accuracy of the UperNet segmentation network model in recognizing and segmenting photovoltaic panels.
[0115] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0116] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0117] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0118] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0119] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0120] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the methods and core ideas of the present application; at the same time, for those of ordinary skill in the art, according to the ideas of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A photovoltaic recognition method based on visual feature constraints, characterized in that: The photovoltaic recognition method based on visual feature constraints includes: Acquire a remote sensing image data set; the remote sensing image data set includes a plurality of remote sensing images and corresponding true labels, each remote sensing image corresponds to one true label, and the true label is used to characterize whether the surface of the remote sensing image contains a photovoltaic panel and the area where the photovoltaic panel is located; Constructing an Upernet segmentation network model; the Upernet segmentation network model takes the remote sensing image as input and takes the image segmentation result corresponding to the remote sensing image as output, the Upernet segmentation network model includes a feature pyramid network, a feature aggregation module and a semantic segmentation detection head, the feature pyramid network is used to extract and fuse the multi-scale features of the remote sensing image, the feature aggregation module is used to integrate the feature map corresponding to the fused multi-scale features, and the semantic segmentation detection head is used to output the image segmentation result according to the integrated feature map; the image segmentation result is the result of segmenting the area where the photovoltaic panel is located in the predicted remote sensing image; Constructing a feature constraint module based on an adversarial autoencoder; the feature constraint module includes a color feature constraint module, a texture feature constraint module and a shape feature constraint module, and the color feature constraint module, the texture feature constraint module and the shape feature constraint module all include an autoencoder, a decoder and a discriminator, the autoencoder is used to encode an input feature vector into a low-dimensional manifold, the decoder is used to decode and reconstruct the input vector, and the discriminator is used to distinguish the difference between the image segmentation result corresponding to the remote sensing image predicted by the Upernet segmentation network model and the true label of the remote sensing image; Using the remote sensing image data set, the feature constraint module and the Upernet segmentation network model are subjected to adversarial training to obtain an optimal Upernet segmentation network model; A remote sensing image to be measured of a target area is obtained, and the remote sensing image to be measured is input into the optimal Upernet segmentation network model to predict the image segmentation result corresponding to the remote sensing image to be measured of the target area.
2. The photovoltaic identification method based on visual feature constraints according to claim 1 is characterized in that: After the step of acquiring the remote sensing image data set, the photovoltaic recognition method based on visual feature constraints further includes: Preprocess each remote sensing image in the remote sensing image dataset to obtain a preprocessed remote sensing image dataset; the preprocessed remote sensing image dataset is used to perform adversarial training on the feature constraint module and the Upernet segmentation network model to obtain the optimal Upernet segmentation network model.
3. The photovoltaic identification method based on visual feature constraints according to claim 2 is characterized in that: Preprocessing each remote sensing image in the remote sensing image dataset to obtain a preprocessed remote sensing image dataset specifically includes: Performing cropping processing on each of the remote sensing images in the remote sensing image data set to obtain a cropped remote sensing image data set; the cropped remote sensing image data set includes a plurality of cropped remote sensing images; Data enhancement processing is performed on each of the cropped remote sensing images in the cropped remote sensing image dataset to obtain the preprocessed remote sensing image dataset; the data enhancement processing includes image random rotation processing, and the image random rotation processing includes image random horizontal flipping processing and image random vertical flipping processing.
4. The photovoltaic identification method based on visual feature constraints according to claim 1 is characterized in that: Constructing the color feature constraint module specifically includes: For the color feature constraint module, the predicted mask and the true label of the Upernet segmentation network model are respectively accepted, and the predicted mask and the true label are respectively masked with the corresponding remote sensing image to determine the target area image; Calculating a color histogram of the target area according to the target area image; Based on the color histogram of the target area, the automatic encoder of the color feature constraint module is used to respectively accept the color histogram feature vector of the positive area predicted by the Upernet segmentation network model and the color histogram feature vector of the real positive area, and map each color feature in the color histogram feature vector to a low-dimensional vector; the positive area is the area where the photovoltaic panel is located; Using a decoder of the color feature constraint module, the color-related features in the low-dimensional vector are captured while the color-irrelevant features are discarded, and the input color features are reconstructed; The discriminator of the color feature constraint module is used to distinguish whether the input color feature comes from the image segmentation result predicted by the Upernet segmentation network model or from a real remote sensing image sample.
5. The photovoltaic identification method based on visual feature constraints according to claim 1, characterized in that: Constructing the texture feature constraint module specifically includes: For the texture feature constraint module, the LBP algorithm is used to calculate the local binary pattern coding of the remote sensing image; Determine the LBP coding map corresponding to the remote sensing image according to the local binary pattern coding of the remote sensing image; Masking the LBP encoding map with the corresponding prediction mask of the Upernet segmentation network model and the true label to obtain the LBP value of the target area; The LBP values of the target area are used as texture feature vectors and sent to the automatic encoders of the texture feature constraint module respectively, and the discriminator of the texture feature constraint module is used to distinguish the texture feature differences in the adversarial training process.
6. The photovoltaic identification method based on visual feature constraints according to claim 1, characterized in that: Constructing the shape feature constraint module specifically includes: For the shape feature constraint module, the image segmentation result predicted by the Upernet segmentation network model and the true label are input into the autoencoder of the shape feature constraint module, and the autoencoder is trained using the adversarial training method so that the autoencoder of the shape feature constraint module encodes two types of shapes and captures the difference between the two types of shapes. At the same time, the Upernet segmentation network model is used to antagonize the autoencoder of the shape feature constraint module so that the autoencoder of the shape feature constraint module cannot capture the difference between the two types of shapes; wherein the two types of shapes refer to the shape of the binary image of the image segmentation result predicted by the Upernet segmentation network model and the shape of the binary image of the true label.
7. The photovoltaic identification method based on visual feature constraints according to claim 1, characterized in that: The expression of the objective loss function used in the adversarial training process is: Among them, L all Represents the total objective loss function of adversarial training, L seg , L s , L t , L c They represent segmentation loss, shape loss, texture loss, and color loss respectively; λ1, λ2, and λ3 are hyperparameters; G and D represent the Upernet segmentation network model and the feature constraint module respectively.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the photovoltaic identification method based on visual feature constraints described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the photovoltaic identification method based on visual feature constraints described in any one of claims 1 to 7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the photovoltaic identification method based on visual feature constraints described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Thyroid ultrasound image unsupervised domain adaptive semantic segmentation method based on feature decoupling
CN116630619A
Polarization image fusion method based on double attention mechanism generative adversarial network
CN117649349A