Visual question answering method, system, device and medium based on remote sensing tampered images
By introducing a visual question-and-answer network based on edge prior guidance in remote sensing image processing, the problems of inaccurate edge information recovery and insufficient feature fusion in replica-mobile tampered images are solved, and higher image processing accuracy and sensitivity are achieved.
Patent Information
- Application Number
- CN202510104910.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-23
AI Technical Summary
In the prior art, when processing replica-mobile tampering remote sensing images, especially marine remote sensing images, there are problems such as inaccurate edge information recovery and insufficient fusion of global and local features, resulting in low accuracy of tampering detection.
The remote sensing replication-mobile tampered image visual question-and-answer network is adopted based on edge prior guidance. Through the combination of the main branch network and the prior branch network, the multi-layer encoder-decoder architecture and edge prior guidance blocks are used to perform visual feature extraction and edge detection, achieving effective fusion of global and local features.
It significantly improves the edge artifact recovery effect of copy-moving tampered images, enhances the accuracy and sensitivity of image processing, and improves the analysis and understanding accuracy of remote sensing tampered images.
Smart Images

Figure CN119539093B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing and computer vision technology, and in particular to a visual question answering method, system, device and medium based on remote sensing tampered images. Background Art
[0002] Remote Sensing Visual Question Answering (RSVQA) is an application of visual question answering in the field of remote sensing. It combines remote sensing image processing, natural language processing, and machine learning techniques to answer questions related to image content by analyzing remote sensing images.
[0003] However, remote sensing image data may be forged and tampered with by means of copy-move forgery, which can create a seemingly reasonable but actually completely false scene by copying and moving a certain area in the image. For copy-move tampered remote sensing images, especially copy-move tampered marine remote sensing images, current visual question answering methods have problems such as low accuracy in analyzing and understanding tampered remote sensing images. The following will discuss them one by one:
[0004] First, there is a lack of effective use of edge prior knowledge. Existing methods (including the SHRNet model) generally lack effective use of edge prior knowledge, which affects the accuracy of image analysis and understanding. For example, the edge information of marine tampered images usually contains rich semantic content, such as the outline of ships, the shape of coastlines, the boundaries of buoys, etc. Due to the lack of use of edge prior knowledge, existing methods may produce blurred or erroneous edges when restoring the edge information of tampered images.
[0005] Second, there is a lack of effective fusion of global and local features. Existing methods (including the SHRNet model) often find it difficult to balance global and local features, resulting in the inability to simultaneously utilize detailed information of the global background and local targets during tampering detection. For example, in marine tampering images, global features can provide overall information about the marine environment, such as weather conditions, tidal changes, etc., while local features focus on edge information of specific targets such as ships and buoys. Due to the lack of effective fusion strategies in existing methods, local edge information is easily lost during the processing process, especially under the interference of natural factors such as waves and fog, the edge information becomes more blurred and less obvious. The insufficient fusion of global and local features makes the tampering detection system not sensitive enough to small changes in the image, and it is difficult to accurately identify the tampered edge area. Summary of the invention
[0006] The technical problem to be solved by the present application is to overcome the deficiencies of the prior art and provide a visual question answering method, system, device and medium based on remote sensing tampered images, which can effectively improve the accuracy of question answering for copy-move tampered remote sensing images.
[0007] To achieve the above-mentioned purpose, the first aspect of the present application provides a visual question-answering method based on remote sensing tampered images, constructing and training a remote sensing copy-movement tampered image visual question-answering network based on edge prior guidance, wherein the question-answering network includes a main branch network and a prior branch network, and the main branch network adopts a multi-layer encoder-decoder architecture. Constructing the question-answering network includes the following steps:
[0008] Visual feature extraction, inputting the copy-move tampered image, performing visual feature extraction on the copy-move tampered image through the main branch network, each layer of the encoder is composed of a plurality of edge prior guide blocks connected, and the edge prior guide blocks perform multi-scale feature extraction on the input features of each layer of the encoder;
[0009] Performing edge detection on the copy-move tampered image through a priori branches to obtain edge prior features, each of the priori branches corresponding to each edge prior guide block performs feature fusion, the edge prior features are fused with the input features in each edge prior guide block and then output as input features of the next edge prior guide block for visual feature extraction;
[0010] Cross-modal fusion features, guided by edge prior knowledge, extract features from the input text to obtain text features, and cross-modally fuse the extracted visual features and the extracted text features to obtain fused multi-modal features;
[0011] Based on the fused multimodal features, multimodal reasoning is performed to output question and answer results.
[0012] Furthermore, the edge prior guidance block performs feature extraction on the input image, including global feature extraction and local feature extraction, and the feature extraction of the edge prior guidance block is output after passing through a self-attention mechanism and an edge prior gated feedforward neural network;
[0013] In the process of extracting global features by the edge prior guidance block, the edge prior guidance block uses a self-attention mechanism to achieve global modeling;
[0014] In the process of extracting local features by the edge prior guidance block, the edge prior features are used as gated features of an edge prior gated feedforward neural network. In the edge prior guidance block, features of an input image are extracted and outputted through the edge prior gated feedforward neural network to obtain visual features.
[0015] Furthermore, given the input feature , as the input of the xth edge prior guidance block, the given feature After batch normalization, the output is processed by the self-attention mechanism and combined with the given features Output features by adding ;
[0016] The output characteristics After batch normalization again, the features are input into the edge prior gated feedforward neural network, and the features processed after batch normalization are processed using edge prior features. As a gated feature Output through addition operation;
[0017] The edge prior guided block feature extraction process is expressed as:
[0018] ;
[0019] ;
[0020] in, represents batch normalization, represents the self-attention mechanism, represents the output of the self-attention mechanism in the xth edge prior guided block, represents the output of the edge prior gated feedforward neural network in the xth edge prior guided block, represents an edge prior gated feedforward neural network, represents the x-1th marginal prior feature, Represents the feature map of the x-1th marginal prior.
[0021] Further, in the edge prior gated feedforward neural network feature extraction process, a first branch and a second branch are included, and the first branch is used to provide edge prior features to the second branch;
[0022] In the first branch, the edge prior features The high-dimensional features are obtained by pixel-by-pixel convolution operation, and then the inter-group gating weights are generated by group convolution with different convolution kernel sizes. , and then the inter-group gating weight Integrate into the second branch for feature extraction, and also pass the prior features of group convolution Through element-wise convolution operation and component gating weights After fusion, it is used as edge prior features Output, the edge prior feature As the input of the first branch in the next edge prior gated feedforward neural network;
[0023] The first branch extraction is expressed as:
[0024] ;
[0025] ;
[0026] in, represents group convolution with different kernel sizes, represents the gating weight between groups, represents the pixel-by-pixel convolution operation, express Activation function, represents the prior features after group convolution, represents the x-1th marginal prior feature.
[0027] Furthermore, in the second branch, for a given feature , firstly widen the channel dimension through pixel-by-pixel convolution operation, and then refine it through depth-wise convolution operation, using the inter-group gating weights generated in the first branch The split output features are weighted and then combined with the prior features after group convolution. After the Hadamard matrix product operation, the output is then pixel-by-pixel convolution operation. The second branch extraction process is expressed as:
[0028] ;
[0029] in, is the Hadamard matrix product, represents the depthwise convolution operation, Represents a given feature, represents the output of the edge prior gated feedforward neural network in the xth edge prior guided block, Represents the prior features after group convolution.
[0030] Furthermore, through edge prior features For input text Guided, marginal prior features After passing through a multi-layer perceptron, it is combined with the input text After matrix addition operation, we get text features , expressed as:
[0031] ;
[0032] in, Represents input text, represents a multi-layer perceptron, Represents edge prior features; represents the matrix addition operation, Represented by edge prior features Text features after induction;
[0033] Through the edge prior features Visual features obtained after guidance and text features After fusion, multimodal features are obtained for visual question answering prediction :
[0034] ;
[0035] in, Represents multimodal features, represents the visual feature representation, Represents text features, Indicates feature fusion.
[0036] Furthermore, the loss function consists of tampering detection loss and visual question answering loss. The tampering detection loss is calculated based on the root mean square error, and the visual question answering loss is determined by the cross entropy loss. The tampering detection loss is calculated as follows:
[0037] ;
[0038] in, represents the number of samples, represents the real mask, represents the prediction mask;
[0039] The cross entropy loss of the visual question answering is expressed as:
[0040] ;
[0041] in, Indicates the true answer, Represented by multimodal features The predicted probability of
[0042] The final loss function is defined as follows:
[0043] ;
[0044] in, Represents the trade-off coefficient.
[0045] To achieve the above-mentioned purpose, the second aspect of the present application provides a visual question answering system based on remote sensing tampered images, the system comprising:
[0046] A remote sensing copy-move tampered image visual question answering network based on edge prior guidance, the question answering network includes a main branch network and a prior branch network, the main branch network adopts a multi-layer encoder-decoder architecture, and also includes:
[0047] A visual feature extraction network, wherein the visual feature extraction network is used to input a copy-move tampered image, and the main branch network is used to extract visual features of the copy-move tampered image, wherein each layer of the encoder is composed of a plurality of edge prior guide blocks connected, and the edge prior guide blocks perform multi-scale feature extraction on the input features of each layer of the encoder;
[0048] Performing edge detection on the copy-move tampered image through a priori branches to obtain edge prior features, each of the priori branches corresponding to each edge prior guide block performs feature fusion, the edge prior features are fused with the input features in each edge prior guide block and then output as input features of the next edge prior guide block for visual feature extraction;
[0049] A cross-modal fusion feature network, wherein the cross-modal fusion feature network is guided by edge prior knowledge to extract features from input text to obtain text features, and cross-modally fuses the visual features extracted by the visual feature extraction network and the extracted text features to obtain fused multi-modal features;
[0050] The question and answer result output network performs multimodal reasoning based on the fused multimodal features and outputs the question and answer result.
[0051] To achieve the above-mentioned purpose, the third aspect of the present application provides a visual question-answering device based on remote sensing copy-move tampered images, including a processor and a memory, wherein a computer program is stored on the memory, and when the computer program is executed by the processor, the visual question-answering method based on remote sensing tampered images as described above is implemented.
[0052] To achieve the above-mentioned purpose, the third aspect of the present application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the visual question answering method based on remote sensing tampered images as described above is implemented.
[0053] After adopting the above technical solution, the present application has the following beneficial effects compared with the prior art:
[0054] In this application, edge information, as one of the most basic features of an image, is of great significance for image understanding and analysis. By introducing a priori branches to provide edge prior features to the main branch, the edge information in the image can be effectively restored and enhanced. The introduction of prior branches significantly restores and strengthens the edge artifacts of the copy-move tampered image, effectively improves the image processing effect, and further improves the accuracy of analysis and understanding of remote sensing tampered images.
[0055] In this application, the Roberts edge operator is used to detect the edge of the image, and the prior branch is introduced to provide the edge prior features to the main branch. The edge prior features are processed by pixel-by-pixel convolution to extract high-dimensional features. Then, group convolutions with different convolution kernel sizes are applied to generate inter-group gating weights, which are used by the deep convolutional feedforward neural network. The use of group convolution can reduce the number of parameters and the amount of calculation of the model.
[0056] In this application, recognizing the potential interference caused by directly applying edge prior features to deep features, an iterative encoding method is adopted, which uses element-by-element convolution to gradually reduce the channel dimension of the gating map, and then outputs it as an edge prior feature to the next edge prior gated feedforward neural network to promote its deep representation.
[0057] In this application, the self-attention mechanism is used to focus on global information, enhance key information, and restore and enhance edge information through edge prior guidance. This global and local fusion strategy provides strong support for the extraction of tampered image edges; the self-attention mechanism can be used to fully capture the global information in the image, and at the same time, the self-attention mechanism is used to enable the main branch network to no longer focus only on local features when processing images, but to grasp the internal connection of the image as a whole. In this process, the self-attention mechanism enhances key information through weight distribution, so that the network pays more attention to features that have an important impact on task performance.
[0058] The specific implementation methods of the present application are further described in detail below in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The accompanying drawings are part of this application and are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application, but do not constitute an improper limitation on this application. Obviously, the drawings described below are only some embodiments. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0060] In the drawings of the specification:
[0061] Figure 1 is a logical schematic diagram of the visual feature extraction process in this specific implementation mode;
[0062] Figure 2 is a logical schematic diagram of edge prior guided block extraction features in this specific implementation mode;
[0063] Figure 3 It is a logical schematic diagram of edge prior gated feedforward neural network feature extraction in this specific implementation method. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not used to limit the scope of the present application.
[0065] The present application provides a visual question-answering method based on remote sensing tampered images, constructs and trains a remote sensing copy-move tampered image visual question-answering network based on edge prior guidance, the question-answering network includes a main branch network and a prior branch network, the main branch network adopts a multi-layer encoder-decoder architecture, and constructing the question-answering network includes the following steps:
[0066] Visual feature extraction: Input the copy-move tampered image, and extract the visual features of the copy-move tampered image through the main branch network. Each layer of the encoder is composed of multiple edge prior guidance blocks connected together. The edge prior guidance blocks perform multi-scale feature extraction on the input features of each layer of encoder;
[0067] The edge prior features are obtained by performing edge detection on the copy-move tampered image through the prior branches. Each prior branch performs feature fusion on each edge prior guide block. The edge prior features are fused with the input features in each edge prior guide block and then output as the input features of the next edge prior guide block for visual feature extraction.
[0068] Cross-modal fusion features, guided by edge prior knowledge, extract features from the input text to obtain text features, and cross-modally fuse the extracted visual features and the extracted text features to obtain fused multi-modal features;
[0069] Based on the fused multimodal features, multimodal reasoning is performed to output question and answer results.
[0070] It should be noted that the executor of the visual question answering method in this embodiment is a visual question answering device based on remote sensing tampered images, which may be an electronic device, a component in an electronic device, an integrated circuit, or a chip. The electronic device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a wearable device, etc., and the non-mobile electronic device may be a server and a personal computer, etc., which are not specifically limited in this application. The following takes the execution subject as an example of a server to describe the visual question answering method based on remote sensing tampered images in this embodiment.
[0071] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.
[0072] In this embodiment, for the input copy-move tampered image, the shallow features are first extracted through convolution operation. At the same time, the edge detection of the input copy-move tampered image is performed through the Roberts edge operator to obtain the edge prior feature map (Edge Prior, EP). Deep feature extraction is performed through a three-layer encoder-decoder network. Each layer of the encoder-decoder is composed of a different number of edge prior guidance blocks connected together. Multiple edge prior guidance blocks have different spatial resolution domains and channel dimensions to extract shallow features. multi-scale features.
[0073] In each edge prior guided block, global modeling is performed through the self-attention mechanism, and then the edge prior gated feedforward neural network is used to further restore and enhance the features of edge structure details, where the lightweight prior branch (PB) encodes the shallow prior features and upsampling involves pixel reorganization and pixel shuffling operations. Skip connections are used to fuse encoder and decoder features from the same layer, and then channel compression is performed through point-by-point convolution, and the final visual features guided by edge priors are obtained. .
[0074] Specifically, edge detection is performed on the copy-move tampered image through the prior branch to obtain edge prior features, and edge detection is performed on the image through the Roberts edge operator to obtain edge prior features.
[0075] It should be noted that the Roberts operator is a gradient calculation method for oblique deviation points. The magnitude of the gradient represents the strength of the edge, and the direction of the gradient is perpendicular to the direction of the edge. The Roberts operator is usually expressed as follows:
[0076] ;
[0077] in, is the input image with integer pixel coordinates, is the horizontal coordinate of the input image, is the vertical coordinate of the input image.
[0078] See also Figure 1 and Figure 2 ,In this embodiment, the edge prior guidance block performs feature extraction on the input image, including global feature extraction and local feature extraction, and the edge prior guidance block feature extraction is output after passing through the self-attention mechanism and the edge prior gated feed-forward neural network;
[0079] In the process of extracting global features by the edge prior guidance block, the edge prior guidance block uses the self-attention mechanism to achieve global modeling;
[0080] In the process of extracting local features in the edge prior guided block, the edge prior features are used as gating features of the edge prior gated feedforward neural network. In the edge prior guided block, the input image is extracted and outputted through the edge prior gated feedforward neural network to obtain visual features.
[0081] Specifically, the Edge Prior-Guided Block (EPGB) follows the paradigm of the Transformer block, achieves long-distance feature perception through a token mixer, and then enhances local features through a feedforward neural network. The difference between this embodiment and the prior art is that in the global feature processing stage, the Edge Prior Guided Block uses a self-attention mechanism to achieve global modeling. For local feature processing, based on the edge prior knowledge of the copy-move tampering monitoring task, this embodiment introduces an Edge Prior-Gated Feed-Forward Network (EGFN), and uses edge prior features as gated features to help enhance the structural recovery capabilities of the feedforward neural network.
[0082] See also Figure 2 In this embodiment, the given feature is input , as the input of the xth marginal prior guided block, given the feature After batch normalization, the output is processed by the self-attention mechanism and combined with the given features Output features after addition operation ;
[0083] Output characteristics After batch normalization again, the features are input into the edge prior gated feedforward neural network. The features processed after batch normalization are used as edge prior features. As a gated feature Output after addition operation;
[0084] The edge prior-guided block feature extraction process is expressed as:
[0085] ;
[0086] ;
[0087] in, represents batch normalization, represents the self-attention mechanism, represents the output of the self-attention mechanism in the xth edge prior guided block, represents the output of the edge prior gated feedforward neural network in the xth edge prior guided block, represents an edge prior gated feedforward neural network, represents the x-1th marginal prior feature, Represents the feature map of the x-1th marginal prior.
[0088] Specifically, this embodiment obtains clear and accurate tampered image edges through the clever fusion of global and local features. This embodiment uses the self-attention mechanism to fully capture the global information in the image. The self-attention mechanism enables the main branch network to no longer focus only on local features when processing images, but to grasp the internal connection of the image as a whole. In this process, the self-attention mechanism enhances key information through weight distribution, so that the main branch network pays more attention to features that have an important impact on task performance.
[0089] In order to further improve the image processing effect, based on the self-attention mechanism, edge prior feature guidance is introduced. Through edge prior feature guidance, the edge information in the image can be effectively restored and enhanced. This global and local fusion mechanism provides strong support for the extraction of tampered image edges.
[0090] See also Figure 3 In this embodiment, in the edge prior gated feedforward neural network feature extraction process, it includes a first branch and a second branch, and the first branch is used to provide edge prior features to the second branch;
[0091] In the first branch, the edge prior features The high-dimensional features are obtained by pixel-by-pixel convolution operation, and then the inter-group gating weights are generated by group convolution with different convolution kernel sizes. , and then the inter-group gating weight Integrate into the second branch for feature extraction, and also pass the prior features of group convolution Through element-wise convolution operation and component gating weights After fusion, it is used as edge prior features Output, the edge prior feature As the input of the first branch in the next edge prior gated feedforward neural network;
[0092] The first branch extraction is expressed as:
[0093] ;
[0094] ;
[0095] in, represents group convolution with different kernel sizes, represents the gating weight between groups, represents the pixel-by-pixel convolution operation, express Activation function, represents the prior features after group convolution, represents the x-1th marginal prior feature.
[0096] Specifically, this embodiment recognizes the potential interference caused by directly applying edge prior features to deep features. This embodiment adopts an iterative coding method, which uses element-by-element convolution operations to gradually reduce the channel dimension of the gated graph, and then outputs it as an edge prior feature to the next edge prior gated feedforward neural network to promote its deep representation.
[0097] Please continue to see Figure 3 In this embodiment, in the second branch, for a given feature , firstly widen the channel dimension through pixel-by-pixel convolution operation, and then refine it through depth-wise convolution operation, using the inter-group gating weights generated in the first branch The split output features are weighted and then combined with the prior features after group convolution. After the Hadamard matrix product operation, the output is then pixel-by-pixel convolution operation. The second branch extraction process is expressed as:
[0098] ;
[0099] in, is the Hadamard matrix product, represents the depthwise convolution operation, Represents a given feature, represents the output of the edge prior gated feedforward neural network in the xth edge prior guided block, Represents the prior features after group convolution.
[0100] Specifically, the main branch extraction follows the structure of a deep convolutional feedforward neural network. For a given feature, the channel dimension is first widened by pixel-wise convolution (PConv), and then the local details are refined by depth-wise convolution (DConv). The split visual features are weighted using the inter-group weights generated in the prior branch, which significantly restores and strengthens the edge artifacts of the copy-move tampered images.
[0101] Specifically, most of the prior art uses standard feedforward neural networks based on deep features for local feature extraction, while ignoring the potential guidance of the prior features of the characteristic tasks. In actual operations, these prior features can be integrated into deep neural networks and have proven their effectiveness. The edge prior gated feedforward neural network in this embodiment integrates the edge prior features into the feedforward neural network in a gated manner to restore and enhance the edge details of the copy-move tampered image, enhance the edge artifacts of the copy-move tampered image, and further improve the accuracy of remote sensing tampered image analysis and understanding.
[0102] In this embodiment, the edge prior feature For input text Guided, marginal prior features After passing through a multi-layer perceptron, it is combined with the input text After matrix addition operation, we get text features , expressed as:
[0103] ;
[0104] in, Represents input text, represents a multi-layer perceptron, Represents edge prior features; represents the matrix addition operation, Represented by edge prior features Text features after induction;
[0105] Through the edge prior features Visual features obtained after guidance and text features After fusion, multimodal features are obtained for visual question answering prediction :
[0106] ;
[0107] in, Represents multimodal features, represents the visual feature representation, Represents text features, Indicates feature fusion.
[0108] Specifically, this embodiment uses the text features guided by the edge prior and visual features Fusion to obtain multimodal features , and use multimodal features Perform multimodal reasoning and then output question-answering results.
[0109] In this embodiment, the loss function consists of tampering detection loss and visual question answering loss. The tampering detection loss is calculated based on the root mean square error, and the visual question answering loss is determined by the cross entropy loss. The tampering detection loss is calculated as follows:
[0110] ;
[0111] in, represents the number of samples, represents the real mask, represents the prediction mask;
[0112] The cross entropy loss for visual question answering is expressed as:
[0113] ;
[0114] in, Indicates the true answer, Represented by multimodal features The predicted probability of
[0115] The final loss function is defined as follows:
[0116] ;
[0117] in, Represents the trade-off coefficient.
[0118] Based on the same inventive concept, the present application also provides a visual question answering system based on remote sensing tampered images, the system comprising:
[0119] A visual question-answering network for remote sensing copy-movement tampering images based on edge prior guidance. The question-answering network includes a main branch network and a prior branch network. The main branch network adopts a multi-layer encoder-decoder architecture and also includes:
[0120] Visual feature extraction network: The visual feature extraction network is used to input the copy-move tampered image, and the visual features of the copy-move tampered image are extracted through the main branch network. Each layer of the encoder is composed of multiple edge prior guidance blocks connected together, and the edge prior guidance blocks perform multi-scale feature extraction on the input features of each layer of encoder;
[0121] The edge prior features are obtained by performing edge detection on the copy-move tampered image through the prior branches. Each prior branch performs feature fusion on each edge prior guide block. The edge prior features are fused with the input features in each edge prior guide block and then output as the input features of the next edge prior guide block for visual feature extraction.
[0122] Cross-modal fusion feature network: The cross-modal fusion feature network is guided by edge prior knowledge to extract features from the input text to obtain text features, and cross-modally fuses the visual features extracted by the visual feature extraction network and the extracted text features to obtain fused multi-modal features;
[0123] The question and answer result output network performs multimodal reasoning based on the fused multimodal features and outputs the question and answer results.
[0124] Based on the same inventive concept, the present application also provides a visual question-answering device based on remote sensing tampered images, including a processor and a memory, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the visual question-answering method based on remote sensing tampered images as described above is implemented.
[0125] Based on the same inventive concept, the present application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the visual question-answering method based on remote sensing tampered images as described above is implemented.
[0126] The program product of the present application for implementing the above method may adopt a portable compact disk read-only memory and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present application is not limited thereto. In the present application, a readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, apparatus, or device.
[0127] It should be noted that a computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0128] The above are only preferred embodiments of the present application, and are not intended to limit the present application in any form. Although the present application has been disclosed as above with preferred embodiments, it is not intended to limit the present application. Any technician familiar with the present application can make some changes or modifications to equivalent embodiments of equivalent changes using the above-mentioned technical contents without departing from the scope of the technical solution of the present application. The implementation schemes in the above embodiments can also be further combined or replaced. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the solution of the present application.
Claims
1. A visual question answering method based on remote sensing tampered images, characterized in that: A remote sensing copy-move tampered image visual question-answering network based on edge prior guidance is constructed and trained. The question-answering network includes a main branch network and a prior branch network. The main branch network adopts a multi-layer encoder-decoder architecture. Constructing the question-answering network includes the following steps: Visual feature extraction, inputting the copy-move tampered image, performing visual feature extraction on the copy-move tampered image through the main branch network, each layer of the encoder is composed of a plurality of edge prior guide blocks connected, and the edge prior guide blocks perform multi-scale feature extraction on the input features of each layer of the encoder; Performing edge detection on the copy-move tampered image through a priori branches to obtain edge prior features, each of the priori branches corresponding to each edge prior guide block performs feature fusion, the edge prior features are fused with the input features in each edge prior guide block and then output as input features of the next edge prior guide block for visual feature extraction; The edge prior guidance block performs feature extraction on the input image, including global feature extraction and local feature extraction, and the feature extraction of the edge prior guidance block is output after passing through a self-attention mechanism and an edge prior gated feedforward neural network; In the process of extracting global features by the edge prior guidance block, the edge prior guidance block uses a self-attention mechanism to achieve global modeling; In the process of extracting local features by the edge prior guidance block, the edge prior features are used as gated features of an edge prior gated feedforward neural network, and the edge prior guidance block extracts features from an input image through the edge prior gated feedforward neural network and then outputs the extracted features to obtain visual features; Cross-modal fusion features, guided by edge prior knowledge, extract features from the input text to obtain text features, and cross-modally fuse the extracted visual features and the extracted text features to obtain fused multi-modal features; Based on the fused multimodal features, multimodal reasoning is performed to output question and answer results.
2. The visual question answering method based on remote sensing tampered images according to claim 1 is characterized in that: Input given features As the input of the xth edge prior guidance block, the given feature After batch normalization, the input is input into the self-attention mechanism and then output and combined with the given features Output features after addition operation The output characteristics After batch normalization again, the features are input into the edge prior gated feedforward neural network, and the features processed after batch normalization are used as edge prior features. As the output feature of the gated feature and the self-attention mechanism Output after addition operation; The edge prior guided block feature extraction process is expressed as: Among them, BN stands for batch normalization, Atten stands for self-attention mechanism, represents the output of the self-attention mechanism in the xth edge prior guided block, represents the output of the edge prior gated feedforward neural network in the xth edge prior guided block, EGFN represents the edge prior gated feedforward neural network, represents the x-1th marginal prior feature, represents the xth marginal prior feature.
3. The visual question answering method based on remote sensing tampered images according to claim 2 is characterized in that: In the edge prior gated feedforward neural network feature extraction process, a first branch and a second branch are included, and the first branch is used to provide edge prior features to the second branch; In the first branch, the edge prior features The high-dimensional features are obtained through pixel-by-pixel convolution operations, and then the inter-group gating weights σ are generated through group convolution with different convolution kernel sizes. The inter-group gating weights σ are then integrated into the main branch for feature extraction, and the prior features after group convolution are also Through the element-by-element convolution operation and the component gating weight σ, it is fused as the edge prior feature Output, the edge prior feature As the input of the first branch in the next edge prior gated feedforward neural network; The first branch extraction is expressed as: Among them, Conv group represents group convolution with different convolution kernel sizes, σ represents the inter-group gating weight, PConv represents the pixel-by-pixel convolution operation, Sigmoid represents the Sigmoid activation function, represents the prior features after group convolution, represents the x-1th marginal prior feature.
4. The visual question answering method based on remote sensing tampered images according to claim 3 is characterized in that: In the second branch, for a given feature First, the channel dimension is widened by pixel-by-pixel convolution operation, and then refined by depth convolution operation. The split output features are weighted by the inter-group gating weights σ generated in the first branch, and then combined with the prior features after group convolution. After the Hadamard matrix product operation, the output is then pixel-by-pixel convolution operation. The second branch extraction process is expressed as: Where ⊙ is the Hadamard matrix product, DConv represents the deep convolution operation, Represents a given feature, represents the output of the edge prior gated feedforward neural network in the xth edge prior guided block, Represents the prior features after group convolution.
5. The visual question answering method based on remote sensing tampered images according to claim 1 is characterized in that: Through edge prior features Guide the input text T, edge prior features After passing through the multi-layer perceptron, the matrix is added to the input text T to obtain the text features. It is expressed as: Where T represents input text, MLP represents multi-layer perceptron, Represents edge prior features; represents the addition matrix addition operation, Represented by edge prior features Text features after induction; Through the edge prior features Visual features obtained after guidance and text features After fusion, multimodal features are obtained for visual question answering prediction : in, Represents multimodal features, represents the visual feature representation, Represents text features, and Fusion represents feature fusion.
6. The visual question answering method based on remote sensing tampered images according to claim 4 is characterized in that: The loss function consists of tampering detection loss and visual question answering loss. The tampering detection loss is calculated based on the root mean square error, and the visual question answering loss is determined by the cross entropy loss. The tampering detection loss is calculated as follows: Where n represents the number of samples, represents the true mask, F v represents the prediction mask; The cross entropy loss of the visual question answering is expressed as: Among them, y i Indicates the true answer, Represented by multimodal features The predicted probability of The final loss function is defined as follows: Among them, α represents the trade-off coefficient.
7. A visual question answering system based on remote sensing tampered images, characterized in that: The system comprises: A remote sensing copy-move tampered image visual question answering network based on edge prior guidance, the question answering network includes a main branch network and a prior branch network, the main branch network adopts a multi-layer encoder-decoder architecture, and also includes: A visual feature extraction network, wherein the visual feature extraction network is used to input a copy-move tampered image, and the main branch network is used to extract visual features of the copy-move tampered image, wherein each layer of the encoder is composed of a plurality of edge prior guide blocks connected, and the edge prior guide blocks perform multi-scale feature extraction on the input features of each layer of the encoder; Performing edge detection on the copy-move tampered image through a priori branches to obtain edge prior features, each of the priori branches corresponding to each edge prior guide block performs feature fusion, the edge prior features are fused with the input features in each edge prior guide block and then output as input features of the next edge prior guide block for visual feature extraction; The edge prior guidance block performs feature extraction on the input image, including global feature extraction and local feature extraction, and the feature extraction of the edge prior guidance block is output after passing through a self-attention mechanism and an edge prior gated feedforward neural network; In the process of extracting global features by the edge prior guidance block, the edge prior guidance block uses a self-attention mechanism to achieve global modeling; In the process of extracting local features by the edge prior guidance block, the edge prior features are used as gated features of an edge prior gated feedforward neural network, and the edge prior guidance block extracts features from an input image through the edge prior gated feedforward neural network and then outputs the extracted features to obtain visual features; A cross-modal fusion feature network, wherein the cross-modal fusion feature network is guided by edge prior knowledge to extract features from input text to obtain text features, and cross-modally fuses the visual features extracted by the visual feature extraction network and the extracted text features to obtain fused multi-modal features; The question and answer result output network performs multimodal reasoning based on the fused multimodal features and outputs the question and answer result.
8. A visual question-answering device based on remote sensing tampered images, characterized in that: It includes a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the visual question answering method based on remote sensing tampered images as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the visual question answering method based on remote sensing tampered images as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Social network image tampering positioning method based on multi-scale feature intelligent perception
CN115063373A
Weak feature target detection method based on multi-prior guidance
CN116883822A