Graph Convolutional Network Model for Extracting Aerial Image Features and Method for Detecting Anomalies
The graph convolutional network model addresses classification challenges in aerial images by using a single-shot multibox detection model and reversible convolution to enhance feature extraction and detection accuracy, overcoming imbalanced sample quantities and image distortions.
Patent Information
- Application Number
- CN202210918508.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-01
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-08-01
AI Technical Summary
The prior art has the problem of low classification accuracy in the classification failure classification in the aerial image classification, especially in the absence of negative samples, it is difficult to effectively detect photovoltaic panel abnormalities.
The combination method of a single-shot multi-frame detection model, image feature conversion compression layer and deconvolution network is adopted. The aerial image features are extracted through the convolution neural network model and the Gate-Conv convolution model, and a hyperspherical surface is constructed for abnormal detection, avoid image correction, reduce the number of parameters, and retain image information.
It improves the accuracy of target abnormality detection in aerial images, can identify photovoltaic panel failures without negative samples, reduces the number of parameters and retains image information, and enhances semantic feature representation.
Smart Images

Figure CN115482473B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of convolutional network models, and in particular, to a graph convolutional network model for extracting features of aerial images and a method for detecting anomalies. Background Art
[0002] Traditional image classification methods require manual feature extraction based on a large amount of prior knowledge. This method is not only time-consuming but also the extracted features are not ideal. Compared with traditional methods, the greatest charm of CNN is that with sufficient computing power and sufficient training data, CNN can automatically learn the best features representing the original image according to the distribution of training samples.
[0003] In the process of implementing the present invention, the inventors found the following technical problems: When classifying faults for aerial images, first, currently CNN is generally based on supervised learning. However, in the scenario of damaged items, due to various situations of item damage, supervised learning will face the challenge of a large gap in the number of positive and negative samples. Usually, it is easy to obtain positive samples, so it is difficult to obtain enough negative samples, such as item breakage. And the image quality is affected by the attitude of the drone during aerial photography, which will cause deformation of the objects in the image. If it is corrected or other processed, it will cause further loss of information. On the premise of the lack of negative samples, it will be more difficult to classify faults. Summary of the Invention
[0004] The embodiments of the present invention provide a graph convolutional network model for extracting features of aerial images and a method for detecting anomalies, so as to solve the technical problem of low classification accuracy of traditional CNN for aerial images due to image object deformation in the prior art.
[0005] In a first aspect, the embodiments of the present invention provide a graph convolutional network model for extracting features of aerial images, including:
[0006] A single-shot multi-box detection model, configured to receive an aerial image obtained by a drone as an input image and output each feature sub-image in the aerial image;
[0007] The image feature transformation and compression layer includes: k stacked image feature transformation and compression modules. The first image feature transformation and compression module receives the input feature sub-image, and the k-th image feature transformation and compression module outputs the final encoded result. The output of each layer of the image feature transformation and compression module serves as the input to the next layer of the image feature compression module. The image feature transformation and compression module includes: a convolutional neural network model, a Gate-Conv convolutional model, and an output unit. The convolutional neural network model is used for convolutional filtering, the Gate-Conv convolutional model is used to reduce the number of parameters and extract information of different granularity sizes, and the image feature transformation and compression unit is used to fuse the operation result output by the convolutional neural network model in the previous layer and the Gate-Conv convolutional model to obtain an encoded feature map;
[0008] The deconvolution network is used to map the encoded feature map to the original image space through multiple steps to generate the feature representation of the feature sub-image.
[0009] Further, the Gate-Conv convolutional model includes: a series of 1*3 and 3*1 convolutional operations, which are in parallel with the 1*1 convolution, and the outputs of the two branches are spliced and fused, and after passing through the batch processing and pooling layer, a preliminary encoding is obtained.
[0010] Further, the Gate-Conv convolutional model is represented as follows:
[0011] ,
[0012] where , represents the weight parameters of the two branches, , are the bias terms of the two branches. () is the pooling function, is the activation function, is the preliminary encoding result.
[0013] Further, the output unit is used for:
[0014] Performing an activation operation on the output result of the convolutional neural network model , using the sigmoid activation function, performing a tanh activation on the output result of the Gate-Conv convolutional model , and fusing the two by short-circuiting to obtain an encoded feature map.
[0015] Further, the output unit operates as follows:
[0016] ,
[0017] where is the weight parameter of the filter, is the sigmoid activation function, , represents the Hadamard product, is the input, is the bias term.
[0018] Furthermore, the model further includes: an adjustment unit, which is configured to adjust the weight parameter and the bias term according to the loss between the output result of the image feature conversion compression layer and the output result of the deconvolution network.
[0019] Even further, the model further includes:
[0020] An anomaly detection model, which is configured to construct a hypersphere, calculate whether the feature representation falls within the hypersphere, and determine whether it is abnormal according to the calculation result.
[0021] In a second aspect, an embodiment of the present invention further provides a method for detecting photovoltaic panel anomalies based on aerial images, which is implemented based on any of the graph convolutional neural network models for extracting aerial image features provided in the above embodiments, and includes:
[0022] Obtain an aerial image including a photovoltaic panel image;
[0023] Input the aerial image into a single-shot multi-box detection model to obtain at least two photovoltaic panel sub-images;
[0024] Input the photovoltaic panel sub-images into an image feature conversion compression layer to perform multi-dimensional extraction and encoding on the information of the photovoltaic panel sub-images, and the encoding is used to reduce the number of parameters while ensuring that the information of the photovoltaic panel sub-images is not lost;
[0025] Input the encoding into a deconvolution network, and map the encoded feature map to the original image space through multiple steps to generate a feature representation of the feature sub-image;
[0026] Input the feature representation into the anomaly detection model, and determine whether the photovoltaic panel corresponding to the photovoltaic panel sub-image is abnormal according to the output anomaly result of the anomaly detection model.
[0027] Furthermore, the method further includes:
[0028] Calculate the loss function of the single-shot multi-box detection model according to the following method:
[0029] ,
[0030] where is the sample set, represents the th default box and the The matching degree of a true value box. If the th default box matches the th true value box, then , where is the localization loss of the overall objective loss function, is the confidence loss, is the weighted parameter that balances the localization loss and the confidence loss. Confidence refers to judging whether there is a photovoltaic panel inside the anchor box, and balancing the localization loss means whether the anchor box anchors the boundary of the photovoltaic panel correctly;
[0031] When the calculation result of the loss function is less than the set difference, at least two photovoltaic panel sub-images are obtained.
[0032] Furthermore, the method further includes:
[0033] Calculating the loss between the output result of the image feature transformation compression layer and the output result of the deconvolution network in the following manner:
[0034] ,
[0035] When the loss exceeds the preset loss threshold, adjust the weight parameters and bias terms of the image feature transformation compression layer until the loss does not exceed the preset loss threshold.
[0036] The graph convolutional network model for extracting aerial image features and the method for detecting anomalies provided by the embodiments of the present invention include: setting a single-shot multi-box detection model for receiving an aerial image obtained by a drone as an input image and outputting each feature sub-image in the aerial image; an image feature transformation compression layer, including: k stacked image feature transformation compression modules. The first layer of the image feature transformation compression module receives the input feature sub-image, and the kth layer of the image feature transformation compression module outputs the final encoded result. The output of each layer of the image feature transformation compression module is used as the input of the next layer of the image feature compression module. The image feature transformation compression module includes: a convolutional neural network model, a Gate-Conv convolutional model, and an output unit. The convolutional neural network model is used for convolutional filtering, the Gate-Conv convolutional model is used to reduce the number of parameters and extract information of different granularity sizes, and the image feature transformation compression unit is used to fuse the operation result output by the convolutional neural network model in the previous layer and the Gate-Conv convolutional model to obtain an encoded feature map;
[0037] The deconvolution network is used to map the encoded feature map to the original image space through multiple steps to generate the feature representation of the feature sub-image. The single-shot multi-box detection model can be used to accurately extract the target image in the aerial image. At the same time, the image feature transformation and compression layer is used to fully extract the features of the target image. On the premise of reducing the number of parameters, it is ensured that the obtained encoded feature map fully reflects the features of the target image. Through the encoding-decoding method, the semantic part in the image is increased, which is convenient for obtaining the image feature representation with rich data volume. Since there is no need to correct the target image, more image information can be retained. In addition, since a hypersphere is constructed to identify the target fault, the detection of the target photovoltaic panel fault can be realized without establishing negative samples. The accuracy of detecting target abnormal conditions in the case of aerial photography is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0039] Figure 1 It is a structural diagram of the graph convolutional neural network model for extracting the features of aerial images provided in Embodiment 1 of the present invention;
[0040] Figure 2 It is a flowchart of the photovoltaic panel anomaly detection method based on aerial images provided in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings.
[0042] Embodiment 1
[0043] Figure 1 It is a structural diagram of the graph convolutional neural network model for extracting the features of aerial images provided in Embodiment 1 of the present invention. Refer to Figure 1 The graph convolutional neural network model for extracting the features of aerial images may include: a single-shot multi-box detection model, which is used to receive the aerial image obtained by the drone as the input image and output each feature sub-image in the aerial image;
[0044] The image feature transformation and compression layer includes: k stacked image feature transformation and compression modules. The first image feature transformation and compression module receives the input feature sub-image, and the k-th image feature transformation and compression module outputs the final encoding result. The output of each layer of the image feature transformation and compression module serves as the input to the next layer of the image feature compression module. The image feature transformation and compression module includes: a convolutional neural network model, a Gate-Conv convolutional model, and an output unit. The convolutional neural network model is used for convolutional filtering, the Gate-Conv convolutional model is used to reduce the number of parameters and extract information of different granularity sizes, and the image feature transformation and compression unit is used to fuse the operation result output by the convolutional neural network model of the previous layer and the Gate-Conv convolutional model to obtain an encoded feature map;
[0045] The deconvolution network is used to map the encoded feature map to the original image space through multiple steps to generate the feature representation of the feature sub-image.
[0046] In this embodiment, taking the aerial image as an aerial photovoltaic panel image as an example, the graph convolutional neural network model for extracting the features of the aerial image is described.
[0047] The single-shot multi-box detection model is mainly composed of a basic network, followed by multiple multi-scale feature blocks. The basic network structure is used to extract features from the input image. It mainly makes predictions by directly regressing the target category and location. In addition, because different scales of feature layers are used for prediction, the target can be well detected even when the image has a low resolution, ensuring its accuracy. During the training process, an end-to-end training method is adopted. The basic network module can use a typical convolutional neural network structure, such as the VGG16 network structure, or for simplicity, network structures such as ResNet and MobileNet can be used instead. It can be simplified as: input image → CNN → prediction network → generate anchor boxes. The extraction process of the anchor boxes is also included in the prediction network. The extraction process is to take each unit of the feature map as the center, find its position in the original image by a geometric ratio method, and extract bounding boxes of different scales centered on this point. Each anchor box will separately predict the probability of the corresponding target category (such as whether it is a photovoltaic panel) and the coordinate values. Each point in the feature map corresponds to different anchor boxes. Then the photovoltaic panels calibrated by the anchor boxes can be obtained, and these anchor boxes are cropped to obtain all the target image information.
[0048] To generate a relatively large number of anchor boxes from the feature map, which can be used to detect targets with smaller sizes. Each multi-scale feature block in the SSD model can be used to reduce (e.g., halve) the height and width of the feature map provided by the previous layer, making the receptive field of each unit in the feature map on the input image broader. Through the multi-scale feature block, single-shot multi-box detection generates anchor boxes of different sizes and detects targets of different sizes by predicting the class and offset of the bounding box.
[0049] When the drone collects video information of the photovoltaic panels, due to the principle of perspective, the photovoltaic panels at different positions are all different inclined planes in the video rather than rectangles. In the traditional processing method, each photovoltaic panel needs to be perspective-corrected to obtain the front-view image of each photovoltaic panel, and the pictures of photovoltaic panels of different sizes are compressed or expanded to a unified size. However, this method will remove some image features or introduce new image features during the processing. Since negative samples cannot be established, it is obvious that this method will increase the error of the detection result. Therefore, in this embodiment, the encoder-decoder structure is introduced.
[0050] The decoder consists of k stacked image feature transformation and compression modules, and each image feature transformation and compression module is composed of a convolutional neural network module and a Gate-Conv convolutional model. The data operations in the image feature transformation and compression layer are respectively passed through the filter module and the Gate-Conv convolutional model, and then the operations of the two branches are fused together to perform operations such as activation pooling to obtain the encoded feature data. , and at the same time is decoded to reconstruct the input data, and the network is trained to make it fit the reconstructed data and the input data.
[0051] The Gate-Conv convolutional model is composed of serial 1*3 and 3*1 convolutional operations, which are parallel to the 1*1 convolutional operation. The outputs of the two branches are fused, and the fusion method is splicing fusion. Then, through the batch processing and pooling layer, the preliminarily encoded result is obtained. . The formula here can be expressed as:
[0052] ,
[0053] where , represents the weight parameters of the two branches, , are the bias terms of the two branches. For more efficient feature compression, the Gate-Conv convolutional model effectively reduces the number of parameters, and through the two branches, it can extract information of different granularity sizes, so as to more efficiently perform multi-dimensional feature extraction and data feature space compression on the sample data, pictures After the calculation of the Gate-Conv module, we get .
[0054] For the convolutional neural network model, the filter is a 3*3 or 5*5 convolution. After passing through the filter layer, we get .
[0055] The output unit can perform an activation operation on . The activation function uses the sigmoid activation. For , it performs a tanh activation. Finally, the two are fused, and drawing on the idea of the residual network structure, they are fused in a shortcut manner to incorporate the information of . The data can be compressed into a hypersphere with the smallest volume, facilitating subsequent anomaly detection. This module can be represented by the following formula:
[0056] ,
[0057] where are the weight parameters of the filter, is the sigmoid activation function, , represents the Hadamard product, that is, element-wise multiplication calculation.
[0058] The decoder consists of a transposed convolution network, aiming to map the encoded features Figure 1 step by step into the original image space. After the feature map is restored and mapped, useful information can be obtained. After the low-level features are mapped into the original image space, some strong semantic information such as colors, corners, textures, and contours can be obtained. These highly semantic high-level information can well characterize the anomaly categories to which the objects belong. The high-level feature map has the ability to reconstruct the contours of the objects to be recognized to a certain extent. That is to say, through certain operations, the high-level feature map can restore the position information of the objects. This is for the first purpose of discovering useful information from the hidden feature space, facilitating the training of the encoder, and greatly improving the accuracy of the encoder in discovering anomaly information. As can be seen from the previous text, the process of the encoder can be expressed as . Then the transposed convolution part can be described as .
[0059] In this embodiment, since both the encoding and decoding involve weight parameters and bias terms, the settings of the weight parameters and bias terms determine the accuracy of the encoding result. Therefore, the model further includes: an adjustment unit, which is used to adjust the weight parameters and bias terms according to the loss between the output results of the image feature conversion compression layer and the transposed convolution network.
[0060] Exemplarily, the process of the encoder can be expressed as . Then the transposed convolution part can be described as Define the reconstruction error Loss as the Euclidean distance between the original input and the reconstructed information. Then the error of this part can be expressed as:
[0061] .
[0062] Optimize the weight parameters and bias terms according to the result of the loss function to obtain more suitable weight parameters and bias terms, further improving the accuracy of the encoding and decoding processes.
[0063] Optionally, the graph convolutional neural network model for extracting aerial image features may further include: an anomaly detection model. The anomaly detection model is used to construct a hypersphere, calculate whether the feature representation falls within the hypersphere, and judge whether it is abnormal according to the calculation result. In this embodiment, classification can be performed through the following idea. Map the original training sample x to a high-dimensional inner product space (or feature space) through a non-linear mapping; then, find a hypersphere (optimal hypersphere) with the smallest volume that contains all or most of the training samples mapped to the feature space in the feature space; finally, through the non-linear mapping, if the image of the new sample point in the feature space falls within the optimal hypersphere, the sample is regarded as a normal point; otherwise, if the image of the new sample in the feature space falls outside the optimal hypersphere, the new sample is regarded as an abnormal point. The optimal hypersphere is determined by its center and radius.
[0064] Optionally, it is necessary to find the center and the radius of the smallest hypersphere that contains most of the data in the feature space . It can be expressed by the following formula:
[0065] ,
[0066] ,
[0067] Slack variable , allowing a trade-off between the slack boundary and the hyperparameter to control the penalty and the volume of the sphere. Points that fall outside the sphere, i.e., , are regarded as anomalies.
[0068] Encode an image into a hidden state through the encoder. For an input color image , 3 is the number of channels, H is the height, and W is the width. After the transformation of the neural network , map its input to the output space , that is , is a A neural network with one hidden layer, and this neural network has parameters , . That is to say, is the feature representation given by the network with parameters . The purpose of this framework is to jointly learn the network parameters and minimize the amount of data, enclosing the hypersphere in the output space , whose characteristics are the radius and the center . Assuming that the given training data is on , the training objective of the network is defined as
[0069] .,
[0070] This network simply uses the quadratic loss to penalize the distance of each network to . The second term is the network weight decay regularizer of the hyperparameter . It is also possible to find a hypersphere with the smallest volume centered at c, and the hypersphere is shrunk by directly penalizing the radius and the data falling outside the sphere, and the sphere is shrunk by minimizing the average distance of all data representations to the center. Similarly, in order to map the overall data as close as possible to the center , the neural network must extract the common factors of the anomalies. Penalize the average distance of all data points, rather than allowing some points to fall outside the hypersphere.
[0071] For a given test picture, the distance from the encoded point of the picture to the center of the hypersphere can be defined as the anomaly score calculated by the framework , that is
[0072] .
[0073] where are the network parameters after training the model, and this score can be adjusted by subtracting the final radius of the trained model, so that anomalies (points with representations outside the sphere) have positive scores, while points inside the hypersphere have negative scores. In this way, the anomaly score of the abnormal photo can be directly obtained. Through the anomaly score, it can be judged whether the target image is abnormal.
[0074] In this embodiment, a single-shot multi-box detection model is set up to receive an aerial photo image captured by a drone as an input image and output each feature sub-image in the aerial photo image; an image feature transformation and compression layer, including: k stacked image feature transformation and compression modules. The first image feature transformation and compression module receives the input feature sub-image, and the k-th image feature transformation and compression module outputs the final encoded result. The output of each layer of the image feature transformation and compression module serves as the input to the next layer of the image feature compression module. The image feature transformation and compression module includes: a convolutional neural network model, a Gate-Conv convolutional model, and an output unit. The convolutional neural network model is used for convolutional filtering, the Gate-Conv convolutional model is used to reduce the number of parameters and extract information with different granularity sizes, and the image feature transformation and compression unit is used to fuse the operation result output by the convolutional neural network model in the previous layer and the Gate-Conv convolutional model to obtain an encoded feature map;
[0075] A deconvolution network is used to map the encoded feature map to the original image space through multiple steps to generate a feature representation of the feature sub-image. The single-shot multi-box detection model can be used to accurately extract the target image in the aerial photo image. At the same time, the image feature transformation and compression layer is used to fully extract the features of the target image. On the premise of reducing the number of parameters, it is ensured that the obtained encoded feature map fully reflects the features of the target image. Through the encoding-decoding method, the semantic part in the image is increased, which is convenient for obtaining a rich image feature representation with a large amount of data. Since there is no need to correct the target image, more image information can be retained. In addition, since a hypersphere is constructed to identify the target fault, the detection of the target photovoltaic panel fault can be realized without establishing negative samples. The accuracy of detecting target abnormal conditions in the case of aerial photography is improved.
[0076] Embodiment 2
[0077] Figure 2 The figure shows a flowchart of a photovoltaic panel anomaly detection method based on an aerial photo image provided in this second embodiment. This embodiment relies on the graph convolutional neural network model for extracting the features of the aerial photo image provided in the above embodiment. Specifically, it includes the following steps:
[0078] Step 210, obtain an aerial photo image including a photovoltaic panel image.
[0079] Most of the existing visible-light-based base station battery panel anomaly detections are based on supervised learning. However, in actual scenarios, supervised learning faces the challenge of a large gap in the number of positive and negative samples. It is easy for us to obtain positive samples, while it is difficult for us to obtain enough samples of negative samples such as damage. On the other hand, the existing methods all collect battery panel photos manually, which makes it difficult to achieve automatic recognition in actual application scenarios.
[0080] Photovoltaic equipment requires regular inspections to maintain the normal operation of the system. At present, the environment around photovoltaic equipment is mixed with people, the site is lush with greenery, and the plants grow and block the photovoltaic panels. In view of the practical problems such as the difficulty of infrastructure maintenance, high risk factor, and strong hidden faults, and focusing on the major needs of urban operation and management for intelligent maintenance of urban infrastructure, an intelligent inspection prototype system based on drone perception is developed to improve the accuracy, safety and timeliness of the urban infrastructure maintenance process. It is difficult for the existing technology to target a variety of abnormal situations, and it is impossible to determine the degree of abnormality of each abnormal situation, such as the degree of dust accumulation, the degree of obstruction by foreign objects, damage to photovoltaic panels and other abnormal situations. There is no indicator to measure the severity of these abnormalities. Therefore, this embodiment obtains images by aerial photography, and the graph convolutional network model that extracts aerial image features can quickly and accurately judge a variety of abnormal situations of photovoltaic panels and give scores for the degree of abnormality.
[0081] The image acquisition device can be configured on the drone and perform 3D map route planning. The drone collects videos of photovoltaic panels along the planned route and extracts frame images from the video according to the rules.
[0082] Specifically, for the detection problem, ,in Indicates pictures, , for a color image, then The whole detection method can be expressed as .in The output image after anomaly detection. Anomaly detection model , is the model weight parameter.
[0083] Step 220: input the aerial image into a single-shot multi-frame detection model to obtain at least two photovoltaic panel images.
[0084] First, collect a large number of photovoltaic panel videos, use marking tools to annotate the photovoltaic panel data, and mark a large amount of photovoltaic panel data.
[0085] The data is input into the single-shot multi-box detection model, which generates anchor boxes of different sizes and detects objects of different sizes by predicting the category and offset of the bounding box.
[0086] Optionally, the method may also calculate the loss function of the single-shot multi-frame detection model in the following manner:
[0087] ,
[0088] in is a sample set, Indicates The matching degree between the th default box and the th ground truth box. If the th default box matches the th ground truth box, then where is the localization loss of the overall objective loss function, is the confidence loss, and
[0089] is the weighted parameter that balances the localization loss and the confidence loss. The confidence indicates whether the object in the anchor box is a photovoltaic panel, and the balanced localization loss indicates whether the anchor box anchors the boundary of the photovoltaic panel correctly; when the calculation result of the loss function is less than the set difference, at least two photovoltaic panel sub-images are obtained.
[0090]
[0091] where is the sample set, represents the th default box and the th ground truth box. If the th default box matches the th ground truth box, then where is the localization loss of the overall objective loss function, is the confidence loss, and is the weighted parameter that balances the localization loss and the confidence loss. The confidence indicates whether the object in the anchor box is a photovoltaic panel, and the balanced localization loss indicates whether the anchor box anchors the boundary of the photovoltaic panel correctly.
[0092] Step 230: Input the photovoltaic panel sub-image into the image feature transformation and compression layer, perform multi-dimensional extraction on the information of the photovoltaic panel sub-image, and perform encoding, where the encoding is used to reduce the number of parameters while ensuring that the information of the photovoltaic panel sub-image is not lost.
[0093] Encode the photovoltaic panel sub-image through the image feature transformation and compression layer. It can fully ensure the integrity of information and reduce the number of parameters.
[0094] Step 240: Input the encoding into the deconvolution network, map the encoded feature map to the original image space through multiple steps, and generate the feature representation of the feature sub-image.
[0095] In addition, the following method can be used to calculate the loss between the output result of the image feature transformation compression layer and the output result of the deconvolution network:
[0096] , when the loss exceeds a preset loss threshold, adjust the weight parameters and bias terms of the image feature transformation compression layer until the loss does not exceed the preset loss threshold.
[0097] Step 250: Input the feature representation into the anomaly detection model, and determine whether the photovoltaic panel corresponding to the photovoltaic panel sub-image is abnormal according to the output anomaly result of the anomaly detection model.
[0098] In this embodiment, an aerial image including a photovoltaic panel image is obtained; the aerial image is input into a single-shot multi-box detection model to obtain at least two photovoltaic panel sub-images; the photovoltaic panel sub-images are input into an image feature transformation compression layer to perform multi-dimensional extraction and encoding on the information of the photovoltaic panel sub-images, and the encoding is used to reduce the number of parameters while ensuring no loss of the information of the photovoltaic panel sub-images; the encoded data is input into a deconvolution network, and the encoded feature map is mapped to the original image space through multiple steps to generate a feature representation of the feature sub-image; the feature representation is input into the anomaly detection model, and it is determined whether the photovoltaic panel corresponding to the photovoltaic panel sub-image is abnormal according to the anomaly result output by the anomaly detection model. By optimizing the computational network architecture of the model, reducing the parameters and computational amount of the model, the accuracy of anomaly detection can be maintained. There is no need to use a large number of abnormally labeled pictures to model normal image features, thereby realizing an unsupervised learning anomaly detection method. In terms of feature extraction, efficient and discriminative features are achieved with very few parameters. The accuracy of detecting target anomalies in aerial photography is improved.
[0099] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0100] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A graph convolutional neural network model for extracting features of aerial images, characterized in that, Including: A single-shot multi-box detection model for receiving an aerial image captured by a drone as an input image and outputting each feature sub-image in the aerial image; An image feature transformation and compression layer, including: k stacked image feature transformation and compression modules. The first image feature transformation and compression module receives the input feature sub-image, and the k-th image feature transformation and compression module outputs the final encoding result. The output of each image feature transformation and compression module serves as the input to the next image feature compression module. The image feature transformation and compression module includes: a convolutional neural network model, a Gate-Conv convolutional model, and an output unit. The convolutional neural network model is used for convolutional filtering, and the Gate-Conv convolutional model is used to reduce the number of parameters and extract information of different granularity sizes. The output unit is used for: For the output result of the convolutional neural network model perform an activation operation, using the sigmoid activation function for activation, and for the output result of the Gate-Conv convolutional model perform tanh activation, and fuse the two in a short-circuit manner to obtain an encoded feature map; the image feature transformation and compression module is used to fuse the operation result output by the previous layer of the convolutional neural network model and the Gate-Conv convolutional model to obtain an encoded feature map; A deconvolution network for mapping the encoded feature map to the original image space through multiple steps to generate a feature representation of the feature sub-image; An anomaly detection model that is used to construct a hypersphere, calculate whether the feature representation falls within the hypersphere, and determine whether it is abnormal based on the calculation result; The training objective of the network model is defined as:
2. The model according to claim 1, wherein The Gate-Conv convolutional model includes: a series of 1*3 and 3*1 convolutional operations that are parallel to the 1*1 convolution operation, and the outputs of the two branches are concatenated and fused, and after passing through a batch normalization layer and a pooling layer, a preliminary encoding is obtained.
3. The model according to claim 2, characterized in that The Gate-Conv convolutional model is represented as follows: Among them, W G1 , W G2 represent the weight parameters of the two branches, b1 and b2 are the bias terms of the two branches, Pooling() is the pooling function, σ is the activation function, is the preliminary encoding result.
4. According to the model described in claim 1, the output unit operates in the following manner: where W1 is the weight parameter of the filter, σ is the sigmoid activation function, ⊙ represents the Hadamard product, and x i is the input, and b is the bias term.
5. The model according to claim 4, characterized in that, The model further includes: an adjustment unit that is used to adjust the weight parameters and bias terms according to the loss between the output result of the image feature transformation and compression layer and the output result of the deconvolution network.
6. A method for detecting anomalies in photovoltaic panels based on aerial images, characterized in that, The method is implemented based on the graph convolutional neural network model for extracting aerial image features described in any one of claims 1-5, and includes: Obtaining an aerial image including a photovoltaic panel image; Inputting the aerial image into the single-shot multi-box detection model to obtain at least two photovoltaic panel sub-images; Inputting the photovoltaic panel sub-images into the image feature transformation and compression layer to perform multi-dimensional extraction and encoding of the information of the photovoltaic panel sub-images. The encoding is used to reduce the number of parameters while ensuring that the information of the photovoltaic panel sub-images is not lost; Inputting the encoding into the deconvolution network to map the encoded feature map to the original image space through multiple steps to generate a feature representation of the feature sub-image; Inputting the feature representation into the anomaly detection model, and determining whether the photovoltaic panel corresponding to the photovoltaic panel sub-image is abnormal according to the output anomaly result of the anomaly detection model.
7. The method according to claim 6, characterized in that, The method further includes: Calculating the loss function of the single-shot multi-box detection model according to the following manner: where x is the sample set, x ij = {0, 1} represents the matching degree between the i-th default box and the j-th ground truth box. If the i-th default box matches the j-th ground truth box, then x ij = 1, where L loc (x, l, g) is the localization loss of the overall objective loss function, and L conf (x, c) is the confidence loss. α is the weighting parameter that balances the localization loss and the confidence loss. The confidence indicates whether there is a photovoltaic panel inside the anchor box, and the balanced localization loss indicates whether the boundary of the anchor box for the photovoltaic panel is correctly anchored; When the calculation result of the loss function is less than a set difference, at least two photovoltaic panel sub-images are obtained.
8. The method according to claim 6, characterized in that, The method further includes: Calculating the loss between the output result of the image feature transformation and compression layer and the output result of the deconvolution network by using the following manner: When the loss exceeds a preset loss threshold, adjust the weight parameters and bias terms of the image feature transformation and compression layer until the loss does not exceed the preset loss threshold.
Citation Information
Patent Citations
Pulmonary nodule detection method and device and computer storage medium
CN110838114A
Image semantic segmentation method based on coding and decoding structure
CN113807355A