Method for identifying human activities in ecological protection areas based on high-resolution remote sensing data transfer learning
Patent Information
- Application Number
- CN202510016120.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-01-06
AI Technical Summary
然而,在实际应用中,遥感影像的质量和数据分析能力仍面临诸多挑战,限制了其在生态保护中的有效性和实时性
[0066]1、本发明解决了遥感影像去云、小样本分析和人类异常活动监测的难题,并构建一体化的监测系统,实现对生态保护区域人类异常活动的监测与预警,为打击非法活动、保护生态环境提供强有力的技术支持。
Smart Images

Figure CN120071171B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human identification technology, and in particular to a method for identifying human activities in ecological protection areas based on transfer learning of high-resolution remote sensing data. Background Technology
[0002] With the rapid development of remote sensing technology (electromagnetic wave detection), its application in ecological protection is becoming increasingly widespread. Remote sensing technology acquires surface information through platforms such as satellites and drones, enabling large-scale and efficient monitoring of human activities in ecological protection areas, especially abnormal behaviors such as illegal construction. However, in practical applications, the quality of remote sensing imagery and data analysis capabilities still face many challenges, limiting its effectiveness and real-time performance in ecological protection.
[0003] First, cloud cover is one of the main problems affecting the quality of remote sensing images. Ecological protection areas are mostly located in remote or cloudy regions, where clouds severely interfere with the acquisition of surface information, making the image data unusable for monitoring abnormal human activities. Traditional de-clouding methods (such as threshold-based cloud detection and multi-temporal image synthesis) can reduce the impact of clouds to some extent, but their effectiveness is limited when dealing with complex cloud distributions and high-resolution images, making it difficult to meet the needs of accurate monitoring.
[0004] Secondly, the insufficient sample size of remote sensing images in ecological protection areas is another major challenge. Due to high data acquisition costs, limited historical data, and difficulties in annotation, the number of samples available for training is small, leading to a significant decline in the performance of traditional machine learning models. Although deep learning methods perform well in remote sensing image analysis, they rely on a large amount of labeled data and are difficult to apply directly to small-sample scenarios, especially in the special environment of ecological protection areas.
[0005] Furthermore, the complexity and concealment of abnormal human activities increase the difficulty of monitoring. Traditional monitoring methods often rely on single features or simple models, making it difficult to comprehensively capture the dynamic changes of abnormal activities, resulting in insufficient monitoring accuracy. The complex terrain and vegetation cover of ecological protection areas further increase the complexity of monitoring. Summary of the Invention
[0006] This invention discloses a method for identifying human activities in ecological protection areas based on transfer learning of high-resolution remote sensing data. The specific method is as follows:
[0007] Collect remote sensing image data of ecological protection areas to serve as a dataset of real cloud-covered images;
[0008] Based on a real set of cloud-covered images, a generative adversarial network model incorporating an adaptive spatial attention algorithm is used to generate a cloudless image dataset.
[0009] A feature extraction model trained on a general dataset is used to extract data features from a cloudless image dataset.
[0010] The extracted data features were used to train the YOLOv5 model;
[0011] The YOLOv5 model was used to identify human activities in ecological protection areas.
[0012] Furthermore, the YOLOv5 model is used to identify human activities in ecological protection areas. The specific method is as follows:
[0013] Collect remote sensing image data of ecological protection areas to be identified;
[0014] An adaptive spatial attention algorithm is used to perform cloud removal on the remote sensing image data to be identified.
[0015] The cloudless image after cloud removal is processed using a feature extraction model;
[0016] The extracted features are input into the YOLOv5 model, and the YOLOv5 model outputs the recognition results.
[0017] Furthermore, a cloudless image dataset is generated, using the following method:
[0018] A densely dilated convolutional merging network is used to extract local features from a real cloud-covered image dataset.
[0019] Expanding the receptive field of view through dilated convolution to capture global features;
[0020] An attention weight map is generated using an adaptive spatial attention algorithm;
[0021] Based on the attention weight map, dynamically adjust the attention allocation for cloud-covered areas;
[0022] Obtain a set of real, cloudless images;
[0023] Generative adversarial network models generate cloudless image datasets based on real cloudless image sets.
[0024] Furthermore, the generative adversarial network model generates a cloudless image dataset based on a set of real cloudless images, using the following method:
[0025] Generate a set of simulated cloudless images using a generator;
[0026] A discriminator is used to distinguish between simulated cloudless image sets and real cloudless image sets;
[0027] Once the discriminator's discrimination result reaches the preset target, the iteration stops, and the simulated cloudless image generated by the generator is mixed with the real cloudless image as the generated cloudless image dataset.
[0028] Furthermore, during the iteration process, the total loss function formula is as follows:
[0029] L total =L GAN +λ1L L1 +λ2L att
[0030] In the formula, L GAN L1 is the original loss function of the generative adversarial network model, and L2 is the loss function that constrains the pixel-level differences between the generated image and the real image. art The loss function constrains the difference between the attention weight map and the cloud mask, where λ1 and λ2 are weight coefficients.
[0031] Furthermore, the feature extraction model trained on a general dataset is used to extract data features from the cloudless image dataset. The specific method is as follows:
[0032] The feature extraction models include the VGG16 model, the ViT model, and the MobileViT model, which are used to extract low-level, mid-level, and high-level features, respectively.
[0033] VGG16, ViT, and MobileViT models were trained using a general dataset to obtain preliminary weights.
[0034] Data features of cloudless image datasets were extracted using the VGG16 model, the ViT model, and the MobileViT model, respectively.
[0035] Based on the reinforcement learning model, dynamic fusion weights are generated, and the data features extracted by the VGG16 model, ViT model and MobileViT model are dynamically fused to extract the data features of the cloudless image dataset.
[0036] Furthermore, the feature fusion formula for the cloudless image dataset is as follows:
[0037] T Fused =w1·T VGG16 +w2·T ViT +w3·T MobileViT
[0038] In the formula, T VGG16 T ViT T MobileViT The data features extracted from the VGG16 model, ViT model, and MobileViT model are respectively, and w1, w2, and w3 are fusion weights generated by the reinforcement learning model.
[0039] Furthermore, the reward function formula for the reinforcement learning model to generate fusion weights is as follows:
[0040]
[0041] In the formula, r t γ is the reward at step t, and γ is the discount factor;
[0042] The input to the reinforcement learning model is state s t The output is action a. t The probability distribution of the reinforcement learning model is expressed mathematically as follows:
[0043] π(a t |s t ;θ)=Softmax(W3·ReLU(W2·ReLU(W1s t +b1)+b2)+b3)
[0044] Where s t =[T VGG16 ,T ViT ,T MobileViT ] represents the state, θ = {W1, W2, W3, b1, b2, b3} represents the parameters of the policy network, W is the weight matrix of the policy network, and b is the bias vector of the policy network.
[0045] Furthermore, the YOLOv5 model predicts the bounding box, class, and confidence score through convolutional layers, as shown in the following formula:
[0046] P bbox ,P cls ,P obj =f YOLOv5 (T Fused ;θ YOLOv5 )
[0047] Among them, P bbox It is the predicted value of the target bounding box; P cls It is the predicted value of the category; P obj It is the predicted value of the target confidence level; θ YOLOv5 These are the parameters of the YOLOv5 detector head;
[0048] The classification loss function is as follows:
[0049]
[0050] Where, N target y is the number of samples in the target domain. i p is the true class label of the i-th sample. i The class probability distribution of the i-th sample predicted by the model is obtained from the annotation of the dataset and the prediction output of the model;
[0051] The regression loss function is as follows:
[0052]
[0053] Among them, t i Let be the coordinates of the i-th bounding box predicted by the model. Let i be the true coordinates of the i-th bounding box, and smooth L1 The loss function is Smooth L1.
[0054] The target confidence loss is as follows:
[0055]
[0056] Among them, y i The true target confidence level (0 or 1) of the i-th sample, p i It is the target confidence score predicted by the model for the i-th sample.
[0057] The total loss is:
[0058] L target =L cls +L reg +L obj
[0059] The reward is based on the metric mAP, and the specific method is as follows:
[0060] Calculate Precision and Recall for each category, then plot the Precision-Recall curve, and calculate the area under the curve to obtain the AP for that category. Average the APs of all categories to obtain mAP.
[0061] r t =mAP t
[0062] Update the parameters θ of the policy network using the policy gradient method:
[0063]
[0064] Where η is the learning rate.
[0065] Due to the adoption of the above technical solutions, the present invention has the following beneficial effects:
[0066] 1. This invention solves the problems of cloud removal from remote sensing images, small sample analysis, and monitoring of abnormal human activities, and constructs an integrated monitoring system to realize the monitoring and early warning of abnormal human activities in ecological protection areas, providing strong technical support for combating illegal activities and protecting the ecological environment.
[0067] 2. This invention provides a remote sensing image cloud removal method that combines an adaptive spatial attention mechanism with a generative adversarial network, which effectively removes cloud interference and improves image quality.
[0068] 3. To address the problem of insufficient remote sensing image sample size in ecological protection areas, deep transfer learning technology is introduced to transfer rich remote sensing data knowledge from the source domain to the target domain, solving the model training problem under small sample conditions. At the same time, spatial, spectral and texture features of remote sensing images are extracted by combining multi-feature networks, and a fusion method is selected based on dynamic feature selection of reinforcement learning to achieve high-precision monitoring of abnormal human activities.
[0069] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0070] The accompanying drawings of this invention are described below.
[0071] Figure 1 This is a schematic diagram of the cloud removal algorithm process.
[0072] Figure 2 This is a schematic diagram of the workflow of a generative adversarial network model.
[0073] Figure 3 A schematic diagram of the framework structure of a human activity monitoring and decision-making platform.
[0074] Figure 4 This is a schematic diagram of the overall process. Detailed Implementation
[0075] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0076] A method for identifying human activities in ecological protection areas based on transfer learning of high-resolution remote sensing data, such as Figure 4 As shown, the specific steps are as follows:
[0077] S1. Collect remote sensing image data of ecological protection areas to serve as a dataset of real cloud-covered images.
[0078] S2. Based on a real set of cloud-covered images, a generative adversarial network model incorporating an adaptive spatial attention algorithm is used to generate a dataset of cloudless images.
[0079] In step S2, such as Figure 1 As shown, the specific steps for removing clouds are as follows:
[0080] First, input a satellite remote sensing image with cloud cover. cloudy Then, DDCM-Net is used for feature extraction, extracting local features through dense convolutional layers:
[0081] x0=Conv(I cloudy )
[0082] x0 is the output feature map of the dense convolutional layer.
[0083] Extracting multi-scale global features using dilated convolutional layers:
[0084] x l =H l ([x0,x1,...,x l-1 ])
[0085] Where, x l H is the output feature map of the l-th layer. l This refers to the operations at layer l, including dilated convolution, activation functions, and batch normalization, [x0, x1, ..., x]. l-1 The ] indicates that the outputs of the first l layers are concatenated together, and finally, all intermediate layer features are fused through the merging layer:
[0086] x * =Merge([x0,x1,...,x L ])
[0087] x * This is the merged output feature map, with a shape of C. feat ×H in ×W in .
[0088] Then, the attention weight map A is generated using the ASAM adaptive spatial attention mechanism, as follows:
[0089]
[0090] A = σ(W*g + b)
[0091] Here, g is the global feature vector obtained through global average pooling. Its function is to summarize the global information of each channel of the input feature map and generate a vector that can represent the global information of the entire feature map. W and b are the weight matrix and bias vector of the convolutional layer, and σ is the sigmoid activation function.
[0092] Attention weight map A is used to adjust the weights of the input feature map x, enhancing features in important regions and suppressing unimportant regions.
[0093] x att [c,i,j]=x * [c,i,j]·A[i,j]
[0094] x att The adjusted feature map has a shape of C. feat ×H in×W in A[i,j] is the value of the attention weight map at position (i,j).
[0095] Finally, output a cloudless image:
[0096] I cloudless =Conv(x att )
[0097] Right now:
[0098] I cloudless =G(I cloudy A)
[0099] In step S2, such as Figure 2 As shown, the specific workflow of the generative adversarial network model is as follows:
[0100] The discriminator is a standard convolutional neural network used to determine the authenticity of generated images. Its structure is described below:
[0101] Input layer: The input is a cloudless image I generated by the generator. cloudless Or a true cloudless image I ground truth .
[0102] Convolutional layer: Extracts image features f1 = Conv(I).
[0103] Batch normalization layer: accelerates training and stabilizes the model f2 = BatchNorm(f1).
[0104] LeakyReLU: An activation function that introduces nonlinearity f3 = LeakyReLU(f2).
[0105] Output layer: The probability of the output image being true is D(I) = Conv(f3).
[0106] The total loss function consists of three parts:
[0107] GAN loss is used to optimize the adversarial training of the generator and discriminator:
[0108]
[0109] x represents a true cloudless image, derived from the true data distribution P. data Sampled from (x); z represents random noise or the input clouded image, from the noise distribution P. zSampling in (z); G(z) represents the cloudless image generated by the generator based on the input z; D(x) represents the probability of the discriminator D being true to the real image x (the closer to 1, the more real the discriminator considers the image); D(G(z)) represents the probability of the discriminator D being true to the generated image G(z); E represents the expected value, which means taking the average of the data distribution.
[0110] L1 loss is used to constrain the pixel-level differences between the generated image and the real image:
[0111]
[0112] C, H, and W represent the number of channels, the height of the feature map, and the width of the feature map, respectively.
[0113] Attention loss is used to constrain the difference between the attention weight map and the cloud mask:
[0114]
[0115] Where M is a binary mask for the cloud region, where the cloud region is 1 and the non-cloud region is 0.
[0116] The total loss is:
[0117] L total =L GAN +λ1L L1 +λ2L att
[0118] λ1 and λ2 are weighting coefficients used to control the importance of L1 loss and attention loss in the total loss. The weighting coefficients can be dynamically adjusted according to the specific objectives of the task. If the task focuses more on the quality of the generated image, the weight of L1 loss can be increased; if the task focuses more on the accuracy of the attention mechanism, the weight of attention loss can be increased.
[0119] S3. Use a feature extraction model trained on a general dataset to extract data features from a cloudless image dataset.
[0120] In step S3, the general dataset can be remote sensing image data of roads, industrial and mining land, residential and construction land, farmland and other man-made facilities. The data features of the cloudless image dataset are extracted using the following method:
[0121] First, pre-trained VGG16, ViT, and MobileViT are used to extract features. VGG16 is a classic convolutional neural network capable of extracting low-level and mid-level features. The input image x∈R H×W×C Where H is the image height, W is the image width, and C is the number of channels, features are then extracted through multiple convolutional layers:
[0122] FVGG16 =f VGG16 (x;θ VGG16 )
[0123] Where θ VGG16 These are the parameters of VGG16, determined through training.
[0124] ViT segments the image into multiple patches and extracts global features using a Transformer encoder. The input image x∈R is first... H×W×C If the image is divided into N patches, and each patch is P×P, then
[0125] Flatten each patch into a vector and obtain the patch embedding through linear projection:
[0126]
[0127] x class It is a category token; It is the i-th patch; It is the projection matrix, where D is the embedding dimension; E pos ∈R (N+1)×D It is a positional encoding.
[0128] Features are extracted using an L-layer Transformer encoder.
[0129] z l ′=MHA(LN(z l-1 ))+z l-1 ,l=1,...,L
[0130] z l =MLP(LN(z) l ′))+z l ′,l=1,...,L
[0131] Among them, MHA is multi-head self-attention mechanism, LN is layer normalization, and MLP is multilayer perceptron.
[0132] Take the classification token from the last layer as the global feature:
[0133]
[0134] MobileViT combines CNN and Transformer to capture global features while maintaining high efficiency. First, the input image x∈R... H×W×C Local features are extracted using a lightweight CNN:
[0135] F CNN =fCNN (x;θ CNN )
[0136] Where θ CNN These are the parameters of the CNN, determined through training.
[0137] Inputting CNN features into a Transformer encoder to extract global features:
[0138] T MobileViT =f Transformer (F CNN ;θ Transformer )
[0139] After feature extraction, dynamic feature selection and fusion based on reinforcement learning are performed. The reinforcement learning model is designed as follows:
[0140] State: Input feature T VGG16 T ViT T MobileViT ;
[0141] Action: Generate the weights ω1, ω2, ω3 for the weighted summation;
[0142] Reward: mAP of the object detection task;
[0143] The goal of reinforcement learning is to maximize cumulative reward:
[0144]
[0145] Where, r t γ is the reward (mAP) at step t, and γ is the discount factor (usually between 0.9 and 0.99).
[0146] The policy network input is state s t The output is action a. t The probability distribution, policy network π(a) t |s t The mathematical expression for θ is as follows:
[0147] π(a t |s t ;θ)=Softmax(W3·ReLU(W2·ReLU(W1s t +b1)+b2)+b3)
[0148] Where s t =[T VGG16 ,T ViT ,T MobileViTLet θ represent the state, and θ = {W1, W2, W3, b1, b2, b3} represent the parameters of the policy network. W is the weight matrix of the policy network, and b is the bias vector of the policy network.
[0149] Weights are generated using the Softmax function:
[0150] w = Softmax(W3h2 + b3)
[0151] Where w = [w1, w2, w3] are the generated weights, satisfying w1 + w2 + w3 = 1, and h2 is the output of the second hidden layer.
[0152] In other words, reinforcement learning models generate weights based on input features:
[0153] w1,w2,w3=f RL (T VGG16 ,T ViT ,T MobileViT )
[0154] Where f RL It is a reinforcement learning model.
[0155] The features are then weighted and summed based on the generated weights, i.e., feature fusion is performed.
[0156] T Fused =w1·T VGG16 +w2·T ViT +w3·T MobileViT
[0157] S4. Train the YOLOv5 model using the extracted data features.
[0158] In step S4, the YOLOv5 algorithm is used for target detection to detect human activities in the remote sensing image. The fused feature T is then input. Fused The YOLOv5 detection head predicts bounding boxes, categories, and confidence scores through convolutional layers:
[0159] P bbox ,P cls ,P obj =f YOLOv5 (T Fused ;θ YOLOv5 )
[0160] Among them, P bbox It is the predicted value of the target bounding box; P cls It is the predicted value of the category; P obj It is the predicted value of the target confidence level; θ YOLOv5 These are the parameters of the YOLOv5 detector head.
[0161] The classification loss function is as follows:
[0162]
[0163] Where, N target y is the number of samples in the target domain. i p represents the true class label (one-hot encoded) of the i-th sample. i The probability distribution of the class of the i-th sample predicted by the model is obtained from the annotations of the dataset and the model's prediction output.
[0164] The regression loss function is as follows:
[0165]
[0166] Among them, t i Let be the coordinates of the i-th bounding box predicted by the model. Let i be the true coordinates of the i-th bounding box, and smooth L1 This is the Smooth L1 loss function.
[0167] The target confidence loss is as follows:
[0168]
[0169] Among them, y i The true target confidence level (0 or 1) of the i-th sample, p i It is the target confidence level predicted by the model for the i-th sample.
[0170] The total loss is:
[0171] L target =L cls +L reg +L obj
[0172] By continuously optimizing the total loss function, the model gradually learns better feature representations and object detection capabilities.
[0173] Then, the mAP metric is used as a reward to guide the reinforcement learning model in optimizing its feature fusion strategy. Calculating mAP first requires calculating Precision and Recall for each class, then plotting the Precision-Recall curve, and finally calculating the area under the curve to obtain the AP for that class. The AP of all classes is then averaged to obtain mAP. mAP is used as the reward for reinforcement learning.
[0174] r t =mAP t
[0175] Update the parameters θ of the policy network using the policy gradient method.
[0176]
[0177] Where η is the learning rate.
[0178] After completing the dynamic feature selection and fusion step of reinforcement learning, the transfer learning step is performed. The core idea of transfer learning is to transfer the knowledge learned in the source domain ImageNet to the target domain. VGG16, ViT, and MobileViT are pre-trained on ImageNet to learn general feature representations. The goal of pre-training is to minimize the classification loss function.
[0179]
[0180] Where N source It is the number of samples in the source domain, y i It is the true label of the source domain sample, p i It is the probability distribution predicted by the model.
[0181] Then, the VGG16, ViT, and MobileViT weights pre-trained in the source domain are transferred to the target domain model. The shallow convolutional layers of VGG16, ViT, and MobileViT are frozen to retain their general feature extraction capabilities. The model is then fine-tuned in the target domain, and the object detection loss function is minimized to complete the transfer learning, thus solving the model training problem under the given conditions.
[0182] S5. Identify human activities in ecological protection areas using the YOLOv5 model.
[0183] In this embodiment, by integrating these technologies into a single software platform, users can easily upload remote sensing images, automatically complete complex tasks such as cloud removal, feature extraction, and human activity detection, and intuitively view the analysis results and decision-making suggestions. Figure 3 As shown, the platform not only lowers the technical threshold but also improves processing efficiency, making the analysis of remote sensing data and the monitoring of human activities more convenient and practical. Furthermore, through its monitoring and early warning functions, the platform helps users respond quickly to changes in the Earth's surface, providing scientific evidence and decision support for fields such as urban planning, environmental monitoring, and disaster assessment, thereby maximizing the application value of the technology.
[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A method for identifying human activities in ecological protection areas based on transfer learning of high-resolution remote sensing data, characterized in that, The specific method is as follows: Collect remote sensing image data of ecological protection areas to serve as a dataset of real cloud-covered images; Based on a real set of cloud-covered images, a generative adversarial network model incorporating an adaptive spatial attention algorithm is used to generate a cloudless image dataset. A feature extraction model trained on a general dataset is used to extract data features from a cloudless image dataset. The extracted data features were used to train the YOLOv5 model; Identifying human activities in ecological protection areas using the YOLOv5 model; The feature extraction model trained on a general dataset is used to extract data features from a cloudless image dataset. The specific method is as follows: The feature extraction models include the VGG16 model, the ViT model, and the MobileViT model. The VGG16 model extracts low-level and mid-level features, while the ViT model and the MobileViT model extract global features. VGG16, ViT, and MobileViT models were trained using a general dataset to obtain preliminary weights. Data features of cloudless image datasets were extracted using the VGG16 model, the ViT model, and the MobileViT model, respectively. Based on the reinforcement learning model, dynamic fusion weights are generated, and the data features extracted by the VGG16 model, ViT model and MobileViT model are dynamically fused to extract the data features of the cloudless image dataset.
2. The method for identifying human activities in ecological protection areas based on high-resolution remote sensing data transfer learning as described in claim 1, characterized in that, The YOLOv5 model is used to identify human activities in ecological protection areas. The specific method is as follows: Collect remote sensing image data of ecological protection areas to be identified; An adaptive spatial attention algorithm is used to perform cloud removal on the remote sensing image data to be identified. The cloudless image after cloud removal is processed using a feature extraction model; The extracted features are input into the YOLOv5 model, and the YOLOv5 model outputs the recognition results.
3. The method for identifying human activities in ecological protection areas based on high-resolution remote sensing data transfer learning as described in claim 1, characterized in that, The specific method for generating a cloudless image dataset is as follows: A densely dilated convolutional merging network is used to extract local features from a real cloud-covered image dataset. Expanding the receptive field of view through dilated convolution to capture global features; An attention weight map is generated using an adaptive spatial attention algorithm; Based on the attention weight map, dynamically adjust the attention allocation for cloud-covered areas; Obtain a set of real, cloudless images; Generative adversarial network models generate cloudless image datasets based on real cloudless image sets.
4. The method for identifying human activities in ecological protection areas based on high-resolution remote sensing data transfer learning as described in claim 3, characterized in that, Generative adversarial network models generate cloudless image datasets from real cloudless image sets, using the following method: Generate a set of simulated cloudless images using a generator; A discriminator is used to distinguish between simulated cloudless image sets and real cloudless image sets; Once the discriminator's discrimination result reaches the preset target, the iteration stops, and the simulated cloudless image generated by the generator is mixed with the real cloudless image as the generated cloudless image dataset.
5. The method for identifying human activities in ecological protection areas based on high-resolution remote sensing data transfer learning as described in claim 4, characterized in that, During the iteration process, the total loss function is formulated as follows: In the formula, The original loss function of the generative adversarial network model is... To constrain the pixel-level differences between the generated image and the real image, To constrain the loss function based on the difference between the attention weight map and the cloud mask, and These are the weighting coefficients.
6. The method for identifying human activities in ecological protection areas based on high-resolution remote sensing data transfer learning as described in claim 1, characterized in that, The feature fusion formula for cloudless image datasets is as follows: In the formula, , , The data features extracted from the VGG16 model, ViT model, and MobileViT model are respectively. , , All are fusion weights generated by reinforcement learning models.
7. The method for identifying human activities in ecological protection areas based on high-resolution remote sensing data transfer learning as described in claim 6, characterized in that, The reward function formula for generating fusion weights using a reinforcement learning model is as follows: In the formula, This is the reward for step t. It is a discount factor; The input to a reinforcement learning model is a state. The output is an action. The probability distribution of the reinforcement learning model is expressed mathematically as follows: in It's a state. These are the parameters of the policy network. It is the weight matrix of the policy network. It is the bias vector of the policy network.
8. The method for identifying human activities in ecological protection areas based on high-resolution remote sensing data transfer learning as described in claim 7, characterized in that, The YOLOv5 model predicts bounding boxes, categories, and confidence scores through convolutional layers, as shown in the following formula: in, It is the predicted value of the target bounding box; It is the predicted value for the category; It is the predicted value of the target confidence level; These are the parameters of the YOLOv5 detector head; The classification loss function is as follows: in, The number of samples in the target domain. Let i be the true class label of the i-th sample. The class probability distribution of the i-th sample predicted by the model is obtained from the annotation of the dataset and the prediction output of the model; The regression loss function is as follows: in, Let be the coordinates of the i-th bounding box predicted by the model. Let i be the true coordinates of the i-th bounding box. The loss function is Smooth L1. The target confidence loss is as follows: in, The true target confidence level of the i-th sample. It is the target confidence score predicted by the model for the i-th sample; The total loss is: The reward is based on the metric mAP, and the specific method is as follows: Calculate Precision and Recall for each category, then plot the Precision-Recall curve, and calculate the area under the curve to obtain the AP for that category. Average the APs of all categories to obtain mAP. Update the parameters of the policy network using the policy gradient method. : in It is the learning rate.
Citation Information
Patent Citations
Remote sensing image target extraction processing method and device based on deep learning
CN117437555A