Remote sensing image authenticity identification method and system based on feature joint learning
Through feature joint learning method, the color, clarity and texture features of remote sensing images are extracted, and the CFANet model is constructed to make multi-scale feature fusion decisions, solving the accuracy problem of remote sensing image forgery detection and achieving efficient optical remote sensing image identification.
Patent Information
- Application Number
- CN202510566704.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing natural image authenticity and false identification methods cannot effectively adapt to remote sensing image forgery detection, and lack identification methods for high-quality forgery images generated by generative models.
Using a joint feature learning method, a CFANet model is constructed by extracting manual features such as color richness, clarity and image texture of remote sensing images, combining deep learning multi-scale features, a CFANet model is constructed, and a first-order and second-order statistical feature fusion decision is made, and a support vector machine is used for classification.
The discrimination accuracy of remote sensing images is improved, the discrimination ability of optical remote sensing images is enhanced, and the robustness and scalability of the model are improved.
Smart Images

Figure CN120495738A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a remote sensing image authenticity identification method and system based on feature joint learning. Background Art
[0002] In recent years, significant progress has been made in image forgery tasks, which are mainly divided into forgery localization and forgery detection. Forgery localization mainly determines the precise location of the altered original image, while forgery detection aims to verify the overall authenticity of the image. By combining forgery localization and detection techniques based on deep learning, excellent authenticity identification performance has been achieved in many forgery detection tasks. Currently, convolutional neural networks, generative adversarial networks, and diffusion models are widely used in deep forgery image detection. Models based on convolutional neural networks mainly use convolutional layers for effective feature extraction and integrate attention modules into the neural network architecture to enhance global feature extraction capabilities. Models based on generative adversarial networks mainly use generator and discriminator training strategies to enhance their ability to generate and detect forged images.
[0003] In the field of computer vision, some general models have been developed to identify complex forgery operations. For example, they use tampered boundary artifacts and noisy views to learn semantically independent features that can adapt to various types of operations. Due to differences in image acquisition methods and perception technologies, models designed for natural images may not fully meet the needs of forgery localization in remote sensing images. Furthermore, the performance of existing natural image authenticity verification methods may vary depending on the content and operations of the forged image, limiting their ability to comprehensively address all types of images and forgery detection. Therefore, in the field of remote sensing, there is currently a lack of specific methods to identify high-quality forged images produced by generative models. Summary of the Invention
[0004] In view of the defects in the prior art, the purpose of the present invention is to provide a remote sensing image authenticity identification method and system based on feature joint learning.
[0005] According to the present invention, a remote sensing image authenticity identification method based on feature joint learning is provided, comprising:
[0006] Step S1: input remote sensing image data and set the training cycle;
[0007] The remote sensing image data includes real remote sensing image data and forged remote sensing image data;
[0008] Step S2: extracting manual features from each remote sensing image data, including color richness, clarity and image texture;
[0009] Step S3: Construct CFANet, unify the feature map sizes of different convolutional layers through pooling, and extract deep learning multi-scale features through concatenation operations;
[0010] Step S4: Fuse the hand-crafted features and deep learning multi-scale features to generate a new fully connected layer and train the model;
[0011] Step S5: Use the trained model to extract first-order statistical features and second-order statistical features, perform multi-order information fusion decision-making, and output the authenticity detection results of the remote sensing image.
[0012] Preferably, the step S1 includes:
[0013] The input data set of N data Divide into training set and validation set
[0014] Among them, D L 、D tr 、D va Respectively represent the input data set, training set, and validation set; x r Represents the image features of the rth data input to the initialization module; y r Indicates the image label of the rth data input to the initialization module; Represents the image features of the jth data in the training set; Represents the image label of the jth data in the training set; Represents the image features of the p-th data in the validation set; Represents the image label of the p-th data in the validation set; N tr Indicates the number of training samples, N va Indicates the number of validation set samples.
[0015] Preferably, step S2 includes the following sub-steps:
[0016] Step S2.1: Extract image data spatial domain features SDF i ;
[0017] Spatial domain feature SDF i =[CFI i ASM i CON i ENT i IDM i ];
[0018]
[0019]
[0020]
[0021]
[0022]
[0023] Among them, CFI i represents the image color richness index of the i-th image; Indicates the number of different color types of the i-th image; Indicates the total number of pixels of the i-th image; ASM i represents the angular second-order moment of the i-th image based on the gray-level co-occurrence matrix GLCM; P(j,k,d,θ) represents the gray-level co-occurrence matrix; is the number of gray levels of the i-th image; CON i Indicates the contrast of the i-th image based on the gray-level co-occurrence matrix GLCM; ENT i represents the entropy of the i-th image based on the gray-level co-occurrence matrix GLCM; IDM i Represents the inverse difference moment of the i-th image based on the gray-level co-occurrence matrix GLCM;
[0024] Step S2.2: Extract image data color histogram feature CHF i ;
[0025] Step S2.3: Extract image frequency domain features FDF i ;
[0026] Step S2.4: Splicing to obtain manual feature vector HF i ;
[0027] HF i =[SDF i CHF i FDF i ];
[0028] Among them, SDF i Obtain spatial domain features for the i-th image S1; CHF i Obtain the color histogram features for the i-th image S2; FDF i Obtain spatial domain features for the i-th image S1.
[0029] Preferably, the color histogram feature CHF i include:
[0030]
[0031] Let c represent the color channel (R, G, B)
[0032]
[0033]
[0034]
[0035]
[0036]
[0037]
[0038]
[0039]
[0040]
[0041]
[0042]
[0043]
[0044] TIQ i =∑ j,k |L i (j,k)|;
[0045] L i =Laplacian(I i );
[0046] in, Represents the mean color histogram of the i-th image c; Represents the standard deviation of the color histogram of the i-th image; Represents the color histogram of the i-th image Figure 3 Order central moment; N i represents the total number of pixels in the i-th image; h k,c Indicates the value of the kth color channel histogram corresponding to the color channel c; p k Represents the probability of gray level k appearing; Indicates the number of gray levels of the i-th image; MEAN i Indicates the mean of the grayscale histogram of the i-th image; STD i Indicates the standard deviation of the grayscale histogram of the i-th image; SKEW i Indicates the skewness of the grayscale histogram of the i-th image; KURT i Represents the kurtosis of the grayscale histogram of the i-th image; IET i represents the information entropy of the grayscale histogram of the i-th image; BIQ i Represents the i-th image blockiness index; Ii represents the i-th image; H i Indicates the height of the i-th image; W i Indicates the width of the i-th image; TIQ i represents the i-th image texture index; Represents the Sobel gradient in the x direction of the i-th image; Represents the Sobel gradient of the y direction of the i-th image; LIQ i represents the Laplace index of the i-th image; L i Represents the Laplacian operator result of the i-th image.
[0047] Preferably, the image frequency domain feature FDF i include:
[0048] FDF i =[FASM i FCON i FENT i FIDM i ];
[0049]
[0050]
[0051]
[0052]
[0053] Among them, FASM i represents the frequency domain angular second-order moment of the i-th image; P f (j, k, d, θ) represents the co-occurrence matrix based on the frequency domain; is the number of gray levels of the i-th image; FCON i represents the frequency domain contrast of the i-th image based on the frequency domain co-occurrence matrix; FENT i represents the frequency domain entropy of the i-th image based on the frequency domain co-occurrence matrix; FIDM i Represents the frequency domain inverse difference moment of the i-th image based on the frequency domain co-occurrence matrix.
[0054] Preferably, step S3 includes the following sub-steps:
[0055] Step S3.1: Build a new CFANet model based on A-ConvNet, use pooling operations to unify the sizes of different feature maps, and use concatenation operations to aggregate convolutional feature maps in different scale feature spaces;
[0056] Step S3.2: Expand the feature map tensors of different convolutional layers into vectors to represent the deep learning features of the i-th image.
[0057] Preferably, step S4 includes the following sub-steps:
[0058] Step S4.1: Stacking manual features and deep features to obtain joint learning features F i ;
[0059]
[0060] Among them, HF i represents the manual feature vector of the i-th image obtained in step S2;
[0061] represents the depth feature vector obtained by expanding the FM4 tensor of the i-th image obtained in step S3;
[0062] Step S4.2: Perform model training;
[0063] Construct a classification task, the loss function is
[0064] Where l(·) represents the cross entropy loss function; f(·) represents the model output probability distribution; Represents the true label The one-hot encoding vector of θ2 represents the parameters that need to be learned for model training; N tr Indicates the number of training set data; represents the jth image in the training set; Represents the j-th image label in the training set.
[0065] Preferably, step S5 includes:
[0066] Deep learning features and And the resulting joint learning feature F i It is called the first-order statistical feature;
[0067] Input training set, validation set and test set sample D tr 、D va , obtain their first-order statistical feature vectors in FM1, FM2, FM3, FM4 and joint learning features, use the obtained training set data to learn five SVM classifiers, and obtain five soft classification results of the validation set and test set through these classifiers.
[0068] Preferably, assuming Represents the soft classification results of the validation set samples using SVM under the four convolutional feature map representations; constructs an objective function, that is, minimizes the mean square error between the fusion results and the true value on the validation set to learn the weights η = [η1, η2, η3, η4, η5];
[0069]
[0070] Solve the constrained optimization problem and find the optimal value
[0071] After obtaining the weights, the weighted arithmetic mean rule is used to fuse the soft classification results
[0072]
[0073] Among them, D tr 、D va Represent the input training set and validation set respectively; and Respectively represent the vectors obtained by expanding the feature map tensors of the i-th image FM1, FM2, FM3, and FM4; Represents the true label The one-hot encoding vector of ; ‖·‖ represents the Euclidean norm; Indicates the membership degree of the i-th image as forged in the p-th SVM classifier of the first-order statistical feature; Indicates the true membership of the i-th image in the p-th SVM classifier of the first-order statistical feature; N va represents the number of validation set data; η k Represents the weight of the k-th classification result of the first-order statistical feature.
[0074] The covariance matrix of FM is usually used to reflect the second-order statistical characteristics. In practice, it can also be used to build classification models and has achieved good results. Where W, H, and D are width, height, and depth, respectively. We can obtain N0 = W·H vectors, which are obtained by vectorizing χ along the third dimension, and use these vectors to construct a feature matrix X∈RD×N. Therefore, its covariance matrix It can be calculated as:
[0075] The matrix logarithm operation is used to transform the covariance matrix from the manifold space to the Euclidean space, thereby obtaining the second-order statistical features for the downstream classification task. Let U and Σ be the eigenvector matrix and eigenvalue matrix of C respectively, then C=UΣU T At this point, the logarithm of the covariance matrix can be obtained by Calculated. It can be seen that is a symmetric matrix, so here we vectorize the elements of the upper triangular matrix to get a vector
[0076] Assume that the covariance matrices of the feature maps of different convolutional layers are expressed as:
[0077] and
[0078] For the data input to the CFANet model, the second-order statistical feature vectors of FM1, FM2, FM3 and FM4 will be obtained:
[0079] and
[0080] Four SVM classifiers are learned using the training set data represented by second-order statistical features. These classifiers provide four soft classification results for the validation set. The second-order statistical features of different convolutional feature maps are also complementary to each other, and fusing them at the decision layer enhances the reliability of the soft classification results.
[0081] assumed Represents the soft classification results of the validation set obtained by these classifiers. Their weights γ = [γ1, γ2, γ3, γ4] can also be obtained by minimizing the mean square error between the fusion result and the true value on the validation set, that is,
[0082]
[0083] After obtaining the weights, the weighted arithmetic mean rule is used to fuse the soft classification results
[0084]
[0085] in, Represents the true label The one-hot encoded vector of ;
[0086] ||·|| represents the Euclidean norm;
[0087] represents the membership of the i-th image to category 0 (forged) in the p-th SVM classifier of the second-order statistical features;
[0088] Indicates the membership of the i-th image to category 1 (true) in the p-th SVM classifier of the second-order statistical features;
[0089] N va Indicates the number of validation set data;
[0090] γ kRepresents the weight of the k-th classification result of the second-order statistical feature;
[0091] S3, soft classification result fusion;
[0092] The soft classification results obtained using the first-order and second-order statistical features and the soft classification results obtained using the CFANet model are complementary to a certain extent. By effectively integrating these soft classification results, the classification accuracy is expected to be further improved.
[0093] assumed It represents the soft classification results of the validation set obtained by the softmax layer of the joint learning model of manual features and deep features.
[0094] For the validation sample Can be combined and Make category decisions. Soft classification results generated by SVM classifier based on first-order and second-order statistical features and And the soft classification results obtained by the softmax layer of the joint learning model of manual features and deep features Reliability is different, so fusion needs to consider weights. Here we still use the weighted arithmetic average rule to fuse them, where the weight ω = [ω F ,ω S ,ω A ]Use the following formula to learn:
[0095]
[0096] st ω F +ω S +ω A =1,0≤{ω F ,ω S ,ω A}≤1
[0097] Solve the constrained optimization problem and find the optimal value The weighted arithmetic mean fusion result can be obtained as
[0098]
[0099] For the target x to be identified q ,according to The element with the largest calculated probability value is the identification result of the sample.
[0100] According to the present invention, a remote sensing image authenticity identification system based on feature joint learning is provided, comprising:
[0101] Initialization module, used to input real and fake remote sensing image data and set the training cycle;
[0102] Manual feature extraction module, which extracts manual features of color richness, clarity and image texture from each remote sensing image data;
[0103] The deep feature extraction module builds CFANet, unifies the feature map sizes of different convolutional layers through pooling, and extracts multi-scale features of deep learning through serial operations to obtain soft classification results of the first-order and second-order statistical features of each different convolutional feature layer;
[0104] The evidence fusion module estimates the relative weights of each convolutional feature map and softmax layer classification results based on accuracy, estimates the final weights of each classification result, takes the weighted arithmetic average of the results, performs multi-order information fusion decision-making, and finally obtains the sample category based on the fusion results.
[0105] Compared with the prior art, the present invention has the following beneficial effects:
[0106] 1. This invention adopts an innovative feature processing strategy to fuse artificial features with features extracted by deep networks and conduct joint learning. It retains the accurate depiction of local details by artificial features and also leverages the ability of deep features to capture global semantics, thereby expanding the receptive field of the model and obtaining richer image information. This makes the parameter update more complete during the back-propagation process and the gradient transfer more efficient, thereby significantly improving the optimization effect.
[0107] 2. This invention innovatively optimizes the deep feature extraction module and makes targeted improvements to the A-CovNet structure. When processing the feature maps output by different convolutional layers, it introduces a pooling operation to unify the feature map sizes of different convolutional layers, and organically integrates the features of multiple convolutional layers through a series operation, fully extracting the multi-scale features of deep learning, enhancing feature discriminability, and making it more robust.
[0108] 3. This paper employs multi-level information decision-making fusion. First, we systematically extract first- and second-order statistical feature vectors to construct a rich feature representation system. Based on these features, we train a new classifier using a support vector machine (SVM) and use it to make category decisions. The probability outputs of the joint learning model of manual and deep features are arithmetically weighted with the probability outputs of these new classifiers, further improving the model's performance.
[0109] 4. Strong field targeting: For the forgery identification of optical remote sensing images, manual features directly reflect the traces of remote sensing image tampering and improve the judgment accuracy.
[0110] 5. Dynamic weight optimization mechanism: In multi-source information fusion, weights are dynamically learned by minimizing the mean square error of the validation set, rather than preset fixed values.
[0111] 6. Scalability and flexibility: CFANet's modular design allows for the replacement or addition of convolutional layers to adapt to remote sensing images of different resolutions or complexities; the multi-stage fusion framework can easily integrate new features or classifiers, facilitating subsequent improvements.
[0112] Other beneficial effects of the present invention will be explained through the introduction of specific technical features and technical solutions in the specific implementation methods. Those skilled in the art should be able to understand the beneficial technical effects brought about by the introduction of these technical features and technical solutions. BRIEF DESCRIPTION OF THE DRAWINGS
[0113] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0114] Figure 1 The figure is a schematic diagram of the remote sensing image authenticity identification method and system flow based on feature joint learning of the present invention.
[0115] Figure 2 This is a flow chart of the remote sensing image authenticity identification method and system model based on feature joint learning of the present invention.
[0116] Figure 3 Schematic diagram of a sample remote sensing image dataset of the present invention. DETAILED DESCRIPTION
[0117] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.
[0118] Reference Figure 1 As shown, a remote sensing image authenticity identification method based on feature joint learning includes:
[0119] Step 1: Initialize the module to input real and fake remote sensing image data and set the training cycle;
[0120] The input data of the initialization module includes the image and the label corresponding to the image (i.e., true or false);
[0121] If the initialization module inputs N data to form a data set It can be divided into training set and validation set
[0122] D L 、D tr 、D vaRepresent the input data set, training set, and validation set respectively;
[0123] x r Represents the image features of the rth data input to the initialization module; y r Indicates the image label of the rth data input to the initialization module; Represents the image features of the jth data in the training set; Represents the image label of the jth data in the training set; Represents the image features of the p-th data in the validation set; Represents the image label of the p-th data in the validation set; N tr Indicates the number of training samples, N va Indicates the number of validation set samples.
[0124] To achieve optimal identification results for optical remote sensing images, the most direct strategy is to improve quality by learning more features and a wider range of classifiers. Obtaining richer feature vectors to train classifiers and combining the soft outputs of multiple classifiers can yield more reliable outputs and more robust learned features.
[0125] Step 2: Extracting manual features such as color richness, clarity, and image texture of remote sensing image data;
[0126] We select unique and common indicators in remote sensing images for manual feature extraction (blocking effect BIQ, texture TIQ, Laplace index LIQ, etc.), directly targeting common signs of remote sensing image tampering (such as compression artifacts, edge discontinuities, etc.). Specifically, it includes:
[0127] S1. Extract the spatial domain features (SDF) of image data i ;
[0128] The spatial domain SDF i The calculation is as follows:
[0129] SDF i =[CFI i ASM i CON i ENT i IDM i ]
[0130]
[0131]
[0132]
[0133]
[0134]
[0135] Where, CFI i Represents the image color richness index of the i-th image; ASM i represents the angular second-order moment of the i-th image based on the gray-level co-occurrence matrix GLCM; CON i Indicates the contrast of the i-th image based on the gray-level co-occurrence matrix GLCM; ENT i represents the entropy of the i-th image based on the gray-level co-occurrence matrix GLCM; IDM i Represents the inverse difference moment of the i-th image based on the gray-level co-occurrence matrix GLCM;
[0136] Indicates the number of different color types of the i-th image; represents the total number of pixels of the i-th image; is the number of gray levels of the i-th image;
[0137] P(j,k,d,θ) represents the gray-level co-occurrence matrix;
[0138] S2. Extract the color histogram feature of the image data (ColorHistogramFeature, CHF) CHF i ;
[0139] Color histogram feature CHF i The calculation is as follows:
[0140]
[0141] Let c represent the color channel (R, G, B)
[0142]
[0143]
[0144]
[0145]
[0146]
[0147]
[0148]
[0149]
[0150]
[0151]
[0152]
[0153]
[0154] TIQ i =∑ J,k |L i (j,k|
[0155] L i =Laplacian(I i )
[0156] in, Represents the mean color histogram of the i-th image c; Represents the standard deviation of the color histogram of the i-th image; Represents the color histogram of the i-th image Figure 3 order central moment; MEAN i Indicates the mean of the grayscale histogram of the i-th image; STD i Indicates the standard deviation of the grayscale histogram of the i-th image; SKEW i Indicates the skewness of the grayscale histogram of the i-th image; KURT i Represents the kurtosis of the grayscale histogram of the i-th image; IET i represents the information entropy of the grayscale histogram of the i-th image; BIQ i Indicates the i-th image blockiness index; TIQ i Represents the i-th image texture index; LIQ i represents the Laplace index of the i-th image;
[0157] N i represents the total number of pixels in the i-th image; Represents the number of gray levels of the i-th image;
[0158] h k,c Indicates the value of the kth color channel histogram corresponding to the color channel c; p k Represents the probability of gray level k appearing;
[0159] I i represents the i-th image; H i Indicates the height of the i-th image; W i Indicates the width of the i-th image;
[0160] Represents the Sobel gradient in the x direction of the i-th image; Represents the Sobel gradient in the y direction of the i-th image;
[0161] Li Represents the Laplacian operator result of the i-th image;
[0162] S3. Extract frequency domain features (FDF) of the image i ;
[0163] Frequency Domain Features (FDF) i The calculation method is:
[0164] FDF i =[FASM i FCON i FENT i FIDM i ]
[0165]
[0166]
[0167]
[0168]
[0169] Among them, FASM i FCON represents the frequency domain angular second-order moment of the i-th image; i represents the frequency domain contrast of the i-th image based on the frequency domain co-occurrence matrix; FENT i represents the frequency domain entropy of the i-th image based on the frequency domain co-occurrence matrix; FIDM i represents the frequency domain inverse difference moment of the i-th image based on the frequency domain co-occurrence matrix;
[0170] P f (j, k, f, θ) represents the co-occurrence matrix based on the frequency domain;
[0171] is the number of gray levels of the i-th image;
[0172] S4. Splicing to obtain handcrafted feature vector (HF) i ;
[0173] Spatial domain feature SDF obtained based on S1, S2, and S3 i , color histogram feature CHF i , frequency domain feature FDF i In order to obtain more comprehensive artificial features for training and judgment, a manual feature vector HF is constructed i , which is composed of the above three manual features, and has a wider receptive field.
[0174] Handcrafted feature vector HF i The calculation method is:
[0175] HF i =[SDF i CHF i FDF i ]
[0176] Among them, SDF i Obtain spatial domain features for the i-th image S1; CHF i Obtain the color histogram features for the i-th image S2; FDF i Obtain spatial domain features for the i-th image S1;
[0177] Step 3: Construct a convolutional feature aggregation network (CFANet), use pooling operations to unify the feature map sizes of different convolutional layers, and then extract multi-scale features through concatenation operations;
[0178] The discriminative information in different convolutional layers is complementary. In order to effectively utilize the complementary information in shallow and deep convolutional feature maps, a convolutional feature aggregation module is proposed, the A-ConvNet model is optimized, and deep feature extraction is performed to obtain global information of the image and prevent image tampering from a global perspective. Specifically, it includes:
[0179] S1. CFANet construction;
[0180] The A-ConvNet model focuses on extracting high-level semantic information of the target. When using A-ConvNet, the utilization of shallow convolution features is relatively limited. Based on A-ConvNet, a new convolutional feature aggregation network (CFANet) model is constructed. Pooling operations are used to unify the sizes of different feature maps, and concatenation operations are used to aggregate convolutional feature maps in different scale feature spaces. The proposed CFANet model structure is as follows: Figure 2 The Deep Feature Extraction Module module is shown in the figure.
[0181] The CFANet model input is set to 88×88, and five convolutional layers are added. Since the pooling process does not increase the parameters that need to be learned and has a larger receptive field, average pooling and maximum pooling operations are used to unify the width and height of different convolutional feature maps (FMs).
[0182] In order to effectively aggregate feature maps of different sizes produced by multiple convolutional layers, the width and height of FM are uniformly set to 3×3.
[0183] For CFANet, the size of FM1 is 42×42×16, and the maximum pooling and average pooling operations are used without padding, and the width and height of the new feature map will become 3×3.
[0184] The size of FM2 is 19×19×32, and using the maximum pooling and average pooling operations with padding, the width and height of the new feature map will become 3×3.
[0185] The size of FM3 is 7×7×64, and using the average pooling operation with padding, the width and height of the new feature map will become 3×3.
[0186] This results in three new feature maps of sizes 3×3×16, 3×3×32, and 3×3×64.
[0187] The stacking operation is then used to connect the convolutional feature maps of different lengths in series. At this time, the size of FM4 will become 3×3×16+3×3×32+3×3×64+3×3×112=3×3×240.
[0188] S2, deep feature extraction;
[0189] The feature maps of different convolutional layers can usually also be regarded as features of downstream classification tasks. The feature maps generated by the CFANet model can prepare for subsequent training of new classifiers and category decisions.
[0190] FM1, FM2, FM3 and FM4 are used to represent the feature maps of different convolutional layers, and the corresponding feature map tensors are expressed as and For the i-th picture, these tensors can be expanded into vectors, represented as and Their dimensions are
[0191] 42×42×16=28224, 19×19×32=11552, 7×7×64=3136, and 3×3×240=2160. These vectors can represent the deep learning features of the i-th image.
[0192] Step 4: Fuse manual features with deep features to obtain new discriminative features;
[0193] The manual feature vector obtained in step 2 is combined with the deep feature vector obtained in step 3. The obtained vectors form a new fully connected layer, which can enrich the information carried by the original deep feature vector, have a wider receptive field, and better optimize the parameters during model training.
[0194] S1. Stack manual features and deep features to obtain joint learning features F i ;
[0195]
[0196] Among them, HF i represents the manual feature vector of the i-th image obtained in step 2;
[0197] represents the depth feature vector obtained by expanding the FM4 tensor of the i-th image obtained in step 3;
[0198] S2, model training;
[0199] For this classification task, the joint learning feature F obtained by S1 is used i As a fully connected layer, since the manual features are locally targeted, F i Obtaining expert-specific information in addition to global information is more conducive to learning the classification task, and is expected to improve the generalization ability of the model and solve the overfitting problem.
[0200] Construct a classification task, the loss function is
[0201] l(·) represents the cross entropy loss function; f(·) represents the model output probability distribution;
[0202] Represents the true label The one-hot encoded vector of ; represents the jth image in the training set; Represents the j-th image label of the training set;
[0203] θ2 represents the parameters that need to be learned for model training;
[0204] N tr Indicates the number of training set data;
[0205] Step 5: Extract the first-order and second-order statistical features of the deep convolution feature map to achieve multi-order information fusion and decision-making;
[0206] A single feature vector can only reflect the basic distribution of the data, but cannot capture the high-order correlation between features, which may lead to overfitting. Therefore, the multi-order information fusion decision-making method is used to combine the first-order and second-order statistical features to complement each other and reflect the comprehensive image information, which can improve the ability to identify complex tampering patterns and enhance robustness. Figure 2 The model shown extracts first-order and second-order statistical features and then makes multi-order information fusion decisions. Specifically, it includes:
[0207] S1, first-order statistical feature information mining;
[0208] Step 3: Deep Learning Features and And the joint learning feature F obtained in step 4 i These are called first-order statistical features and can be used directly to train classifiers (e.g., SVMs). The first-order statistical features obtained using different convolutional feature maps are usually complementary, so the soft classification results of the target to be identified using these features are also complementary. They can then be combined from the perspective of decision-level fusion.
[0209] Input training set, validation set and test set sample D tr 、D va , we can obtain their first-order statistical feature vectors in FM1, FM2, FM3, FM4 and joint learning features. Then, we use the obtained training set data to learn five SVM classifiers, and use these classifiers to obtain five soft classification results of the validation set and test set.
[0210] assumed Represents the soft classification results obtained by SVM using the four convolutional feature maps for the validation set samples. Since these first-order statistical features are extracted from the same object and are highly correlated, the soft classification results obtained are also highly correlated. In addition, the importance of feature maps and joint learning features of different convolutional layers for downstream classification tasks is usually different. These soft classification results should not be treated equally during fusion. In order to obtain high-quality fusion results, their weights must be considered. Here, an objective function is constructed to minimize the mean square error between the fusion result and the true value on the validation set to learn the weights η = [η1, η2, η3, η4, η5];
[0211]
[0212]
[0213] Solve the constrained optimization problem and find the optimal value After obtaining the weights, the weighted arithmetic average rule is used to fuse the soft classification results
[0214]
[0215] Among them, D tr 、D va Represent the input training set and validation set respectively;
[0216] and They represent the vectors obtained by expanding the feature map tensors of the i-th image FM1, FM2, FM3, and FM4 in step 3 respectively;
[0217] Represents the true label The one-hot encoded vector of ;
[0218] ||·|| represents the Euclidean norm;
[0219] represents the membership degree of the i-th image to category 0 (forged) in the p-th SVM classifier of the first-order statistical features; represents the membership of the i-th image to category 1 (true) in the p-th SVM classifier of the first-order statistical feature;
[0220] N va represents the number of validation set data; η k Represents the weight of the k-th classification result of the first-order statistical feature;
[0221] S2. Second-order statistical feature mining
[0222] The covariance matrix of FM is usually used to reflect the second-order statistical characteristics. In practice, it can also be used to build classification models and has achieved good results. Where W, H, and D are width, height, and depth, respectively. We can obtain N0 = W·H vectors, which are obtained by vectorizing χ along the third dimension, and use these vectors to construct a feature matrix X∈RD×N. Therefore, its covariance matrix It can be calculated as:
[0223] The matrix logarithm operation is used to transform the covariance matrix from the manifold space to the Euclidean space, thereby obtaining the second-order statistical features for the downstream classification task. Let U and Σ be the eigenvector matrix and eigenvalue matrix of C respectively, then C=UΣU T At this point, the logarithm of the covariance matrix can be obtained by Calculated. It can be seen that is a symmetric matrix, so here we vectorize the elements of the upper triangular matrix to get a vector
[0224] Assume that the covariance matrices of the feature maps of different convolutional layers are expressed as:
[0225] and
[0226] For the data input to the CFANet model, the second-order statistical feature vectors of FM1, FM2, FM3 and FM4 will be obtained:
[0227] and
[0228] Four SVM classifiers are learned using the training set data represented by second-order statistical features. These classifiers provide four soft classification results for the validation set. The second-order statistical features of different convolutional feature maps are also complementary to each other, and fusing them at the decision layer enhances the reliability of the soft classification results.
[0229] The soft classification results of them. Their weights γ = [γ1, γ2, γ3, γ4] can also be obtained by minimizing the mean square error between the fusion results and the true value on the validation set, that is,
[0230]
[0231]
[0232] After obtaining the weights, the weighted arithmetic average rule is used to fuse the soft classification results
[0233]
[0234] in, Represents the true label The one-hot encoded vector of ;
[0235] ||·|| represents the Euclidean norm;
[0236] represents the membership of the i-th image to category 0 (forged) in the p-th SVM classifier of the second-order statistical features;
[0237] Indicates the membership of the i-th image to category 1 (true) in the p-th SVM classifier of the second-order statistical features;
[0238] N va Indicates the number of validation set data;
[0239] γ k Represents the weight of the k-th classification result of the second-order statistical feature;
[0240] S3, soft classification result fusion;
[0241] The soft classification results obtained using the first-order and second-order statistical features and the soft classification results obtained using the CFANet model are complementary to a certain extent. By effectively integrating these soft classification results, the classification accuracy is expected to be further improved.
[0242] assumed It represents the soft classification results of the validation set obtained by the softmax layer of the joint learning model of manual features and deep features.
[0243] For the validation sample Can be combined and Make category decisions. Soft classification results generated by SVM classifier based on first-order and second-order statistical features and And the soft classification results obtained by the softmax layer of the joint learning model of manual features and deep features Reliability is different, so fusion needs to consider weights. Here we still use the weighted arithmetic average rule to fuse them, where the weight ω = [ω F ,ω S ,ω A ]Use the following formula to learn:
[0244]
[0245] st ω F +ω S +ω A =1,0≤{ω F ,ω S ,ω A}≤1
[0246] Solve the constrained optimization problem and find the optimal value The weighted arithmetic mean fusion result can be obtained as
[0247]
[0248] For the target x to be identified q ,according to The element with the largest calculated probability value is the identification result of the sample.
[0249] like Figure 3 As a preferred example, the remote sensing image dataset of Beijing and Seattle contains real and fake images, and the size of each sample is a slice image of 88×88 pixels. The basic information of the sample is as follows Figure 3 shown.
[0250] A support vector machine (SVM) with a linear kernel was used as the base classifier to obtain soft classification results and pseudo labels for the samples. The training period was set to 10. This example used classification accuracy, recall, and F1 score as algorithm evaluation metrics. The test results are shown in Tables 1 and 2.
[0251] For optical remote sensing image data, this method simultaneously extracts both handcrafted and deep features and utilizes multi-order statistical feature fusion for decision-making, improving both target recognition performance and stability. Compared to other single-feature extraction and decision-making algorithms, it achieves higher accuracy, achieving recognition accuracy rates exceeding 95%. The results demonstrate that this remote sensing image authenticity verification method and system based on joint feature learning is an effective means for accurate recognition of optical remote sensing images and possesses significant practical application value.
[0252]
[0253] Table 1
[0254]
[0255] Table 2
[0256] The present invention also provides a remote sensing image authenticity identification system based on feature joint learning. The remote sensing image authenticity identification system based on feature joint learning can be implemented by executing the process steps of the remote sensing image authenticity identification method based on feature joint learning, that is, those skilled in the art can understand the remote sensing image authenticity identification method based on feature joint learning as a preferred implementation of the remote sensing image authenticity identification system based on feature joint learning.
[0257] According to the present invention, a remote sensing image authenticity identification system based on feature joint learning is provided. Figure 2 For example, including:
[0258] Initialization module, used to input real and fake remote sensing image data and set the training cycle.
[0259] Manual feature extraction module, which extracts manual features such as color richness, clarity, and image texture from each remote sensing image data;
[0260] The deep feature extraction module builds CFANet, unifies the feature map sizes of different convolutional layers through pooling, and extracts multi-scale features of deep learning through serial operations to obtain soft classification results of the first-order and second-order statistical features of each different convolutional feature layer;
[0261] The evidence fusion module estimates the relative weights of each convolutional feature map and softmax layer classification results based on accuracy, estimates the final weights of each classification result, takes a weighted arithmetic average of the results, performs multi-order information fusion decision-making, and then obtains the sample category based on the fusion results.
[0262] The input real and forged remote sensing image data include images and labels corresponding to the images;
[0263] The manual feature extraction module includes:
[0264] A spatial domain feature extraction unit, used for extracting image spatial domain features;
[0265] A color histogram feature extraction unit, used for extracting image color histogram features;
[0266] A frequency domain feature extraction unit, used for extracting frequency domain features of an image;
[0267] The depth feature extraction module includes:
[0268] Construct a loss unit to find the optimal solution for the loss term and obtain the convolutional feature map and the fully connected layer;
[0269] Statistical feature extraction unit, used to extract first-order and second-order statistical features of FM1-FM4;
[0270] The evidence fusion module includes:
[0271] The weighted fusion unit is used to classify images according to various types of classifiers trained based on first-order and second-order statistical features to obtain soft classification results, and to fuse these soft classification results based on the weighted arithmetic average rule to obtain a fused soft classification result.
[0272] In more preferred embodiments, the initialization module inputs real and forged remote sensing image data and sets a training cycle;
[0273] The manual feature extraction module extracts manual features of spatial features, color histogram features and frequency domain features from remote sensing image data;
[0274] The deep feature extraction module builds CFANet, unifies the size of the convolutional layer feature map through pooling, and extracts deep learning multi-scale features through serial operations to obtain soft classification results of the first-order and second-order statistical features of each different convolutional feature layer;
[0275] The evidence fusion module estimates the relative weights of the convolutional feature map and the softmax layer classification results based on the accuracy, estimates the final weights of each classification result, takes the weighted arithmetic average of the results, performs multi-order information fusion decision-making, and then obtains the category of the sample based on the fusion result.
[0276] Traverse all remote sensing image data and obtain the authenticity identification results of all remote sensing images.
[0277] Those skilled in the art will appreciate that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various devices, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; the devices, modules, and units for implementing various functions can also be considered as both software modules implementing the method and structures within the hardware component.
[0278] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.
Claims
1. A remote sensing image authenticity identification method based on feature joint learning, characterized in that: include: Step S1: input remote sensing image data and set the training cycle; The remote sensing image data includes real remote sensing image data and forged remote sensing image data; Step S2: extracting manual features from each remote sensing image data, including color richness, clarity and image texture; Step S3: Construct CFANet, unify the feature map sizes of different convolutional layers through pooling, and extract deep learning multi-scale features through concatenation operations; Step S4: Fuse the hand-crafted features and deep learning multi-scale features to generate a new fully connected layer and train the model; Step S5: Use the trained model to extract first-order statistical features and second-order statistical features, perform multi-order information fusion decision-making, and output the authenticity detection results of the remote sensing image.
2. The remote sensing image authenticity identification method based on feature joint learning according to claim 1 is characterized in that: The step S1 comprises: The input data set of N data Divide into training set and validation set Among them, D L 、D tr 、D va Respectively represent the input data set, training set, and validation set; x r Represents the image features of the rth data input to the initialization module; y r Represents the image label of the rth data input to the initialization module; Represents the image features of the jth data in the training set; Represents the image label of the jth data in the training set; Represents the image features of the p-th data in the validation set; Represents the image label of the p-th data in the validation set; N tr Indicates the number of training samples, N va Indicates the number of validation set samples.
3. The remote sensing image authenticity identification method based on feature joint learning according to claim 2 is characterized in that: The step S2 includes the following sub-steps: Step S2.1: Extracting spatial features SF of image data i ; Spatial domain feature SDF i =[CFI i ASM i CON i ENT i IDMI i ]; Among them, CFI i represents the image color richness index of the i-th image; Indicates the number of different color types of the i-th image; Indicates the total number of pixels of the i-th image; ASM i represents the angular second-order moment of the i-th image based on the gray-level co-occurrence matrix GLCM; P(j,k,d,θ) represents the gray-level co-occurrence matrix; is the number of gray levels of the i-th image; CON i Indicates the contrast of the i-th image based on the gray-level co-occurrence matrix GLCM; ENT i represents the entropy of the i-th image based on the gray-level co-occurrence matrix GLCM; IDM i Represents the inverse difference moment of the i-th image based on the gray-level co-occurrence matrix GLCM; Step S2.2: Extract image data color histogram feature CHF i ; Step S2.3: Extract image frequency domain features FDF i ; Step S2.4: Splicing to obtain manual feature vector HF i ; <h2 style=";text-align:left;direction:ltr">HF<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> =[SDF<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> CHF<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> FDF<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> ]; Among them, SDF i Obtain spatial domain features for the i-th image S1; CHF i Obtain the color histogram features for the i-th image S2; FDF i Obtain spatial domain features for the i-th image S1.
4. The remote sensing image authenticity identification method based on feature joint learning according to claim 3 is characterized in that: The color histogram feature CHF i include: Let c represent the color channel (R, G, B) WARM i =∑ j,k |L i (j,k)|; L i =Laplacian(Ii); in, Represents the mean color histogram of the i-th image c; Represents the standard deviation of the color histogram of the i-th image; Represents the third-order central moment of the color histogram of the i-th image; N i represents the total number of pixels in the i-th image; h k,c Indicates the value of the kth color channel histogram corresponding to the color channel c; p k Represents the probability of gray level k appearing; Indicates the number of gray levels of the i-th image; MEAN i Indicates the mean of the grayscale histogram of the i-th image; STD i Indicates the standard deviation of the grayscale histogram of the i-th image; SKEW i Indicates the skewness of the grayscale histogram of the i-th image; KURT i Represents the kurtosis of the grayscale histogram of the i-th image; IET i represents the information entropy of the grayscale histogram of the i-th image; BIQ i Represents the i-th image blockiness index; I i represents the i-th image; H i Indicates the height of the i-th image; W i Indicates the width of the i-th image; TIQ i Represents the i-th image texture index; Represents the Sobel gradient in the x direction of the i-th image; Represents the Sobel gradient of the y direction of the i-th image; LIQ i represents the Laplace index of the i-th image; L i Represents the Laplacian operator result of the i-th image.
5. The remote sensing image authenticity identification method based on feature joint learning according to claim 4 is characterized in that: The image frequency domain feature FDF i include: FDF i =[FASM i FCON i FOUND i FIDM i ]; Among them, FASM i represents the frequency domain angular second-order moment of the i-th image; P f (j, k, d, θ) represents the co-occurrence matrix based on the frequency domain; is the number of gray levels of the i-th image; FCON i represents the frequency domain contrast of the i-th image based on the frequency domain co-occurrence matrix; FENT i represents the frequency domain entropy of the i-th image based on the frequency domain co-occurrence matrix; FIDM i Represents the frequency domain inverse difference moment of the i-th image based on the frequency domain co-occurrence matrix.
6. The remote sensing image authenticity identification method based on feature joint learning according to claim 5 is characterized in that: The step S3 includes the following sub-steps: Step S3.1: Build a new CFANet model based on A-ConvNet, use pooling operations to unify the sizes of different feature maps, and use concatenation operations to aggregate convolutional feature maps in different scale feature spaces; Step S3.2: Expand the feature map tensors of different convolutional layers into vectors to represent the deep learning features of the i-th image.
7. The remote sensing image authenticity identification method based on feature joint learning according to claim 6 is characterized in that: The step S4 includes the following sub-steps: Step S4.1: Stacking manual features and deep features to obtain joint learning features F i ; Among them, HF i represents the manual feature vector of the i-th image obtained in step S2; represents the depth feature vector obtained by expanding the FM4 tensor of the i-th image obtained in step S3; Step S4.2: Perform model training; Construct a classification task, the loss function is Where l(·) represents the cross entropy loss function; f(·) represents the model output probability distribution; Represents the true label The unique heat encoding vector; θ2 represents the parameters that need to be learned for model training; N tr Indicates the number of training set data; represents the jth image in the training set; Represents the j-th image label in the training set.
8. The remote sensing image authenticity identification method based on feature joint learning according to claim 7 is characterized in that: The step S5 comprises: Deep learning features and And the resulting joint learning feature F i It is called the first-order statistical feature; Input training set, validation set and test set sample D tr 、D va , obtain their first-order statistical feature vectors in FM1, FM2, FM3, FM4 and joint learning features, use the obtained training set data to learn five SVM classifiers, and obtain five soft classification results of the validation set and test set through these classifiers.
9. The remote sensing image authenticity identification method based on feature joint learning according to claim 8, characterized in that: assumed Represents the soft classification results of the validation set samples using SVM under the four convolutional feature map representations; constructs an objective function, that is, minimizes the mean square error between the fusion results and the true value on the validation set to learn the weights η = [η1, η2, η3, η4, η5]; Solve the constrained optimization problem and find the optimal value After obtaining the weights, the weighted arithmetic average rule is used to fuse the soft classification results Among them, D tr 、D va Represent the input training set and validation set respectively; and Respectively represent the vectors obtained by expanding the feature map tensors of the i-th image FM1, FM2, FM3, and FM4; Represents the true label The one-hot encoding vector of ; ||·|| represents the Euclidean norm; Indicates the membership degree of the i-th image as forged in the p-th SVM classifier of the first-order statistical feature; Indicates the true membership of the i-th image in the p-th SVM classifier of the first-order statistical feature; N va represents the number of validation set data; η k Represents the weight of the k-th classification result of the first-order statistical feature.
10. A remote sensing image authenticity identification system based on feature joint learning, characterized in that: include: Initialization module, used to input real and fake remote sensing image data and set the training cycle; Manual feature extraction module, which extracts manual features of color richness, clarity and image texture from each remote sensing image data; The deep feature extraction module builds CFANet, unifies the feature map sizes of different convolutional layers through pooling, and extracts multi-scale features of deep learning through serial operations to obtain soft classification results of the first-order and second-order statistical features of each different convolutional feature layer; The evidence fusion module estimates the relative weights of each convolutional feature map and softmax layer classification result based on the accuracy, estimates the final weight of each classification result, takes the weighted arithmetic average of the results, performs multi-order information fusion decision-making, and finally obtains the sample category based on the fusion result.
Citation Information
Patent Citations
Image forgery detection method and device and computer storage medium
CN114444566A
Deep pseudo video evidence obtaining method based on semi-supervised learning
CN116824430A
Early fire detection method based on ESNN
CN117409347A