Steganography detection method and system based on multi-source domain adaptation
Through the multi-source domain-adapted steganography detection method, the cross-domain detection accuracy of the steganography detection model is improved by using feature-level mixing and domain alignment losses, solving the problem of poor detection performance in multi-source domain data scenarios, and achieving higher detection accuracy.
Patent Information
- Application Number
- CN202510747812.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The existing steganography analysis methods have reduced detection accuracy under cross-domain conditions, especially in multi-source domain data scenarios, which is difficult to effectively adapt to the steganography noise distribution of the target domain, resulting in poor detection performance.
The multi-source domain adaptation steganography detection method is adopted, and multiple source domain data are mapped to the common feature space through a common feature extractor, and weighted mixing is used to use feature level hybrid modules, and the two-stage alignment is performed through domain alignment loss and classifier prediction of differential losses, improving the generalization ability of the detection model.
It effectively improves the detection accuracy and generalization ability of the steganographic detection model under cross-domain conditions, can better adapt to multi-source domain data scenarios, and improves the detection accuracy on the target domain.
Smart Images

Figure CN120278869A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security technology, and particularly to a steganography detection method and system based on multi-source domain adaptation. Background Technique
[0002] With the rapid development of digital technology, multimedia carriers are increasingly widely used in information dissemination. Steganography, as a technology for embedding secret information into multimedia carriers in a concealed manner, is widely applied in fields such as secret communication, copyright protection, and data security, especially image steganography that uses images widely used in digital media as carriers. However, the concealment of steganography also brings potential security risks. For example, steganographic information may be used for illegal purposes or unauthorized secret transmission. Therefore, steganalysis technology has emerged, aiming to detect whether there is hidden secret information in multimedia carriers, so as to ensure information security and privacy protection. Early image steganalysis mainly relied on manually designed feature extraction methods to detect whether there was secret information by analyzing the statistical characteristics of digital images. With the continuous progress of steganography technology, traditional methods based on handcrafted features have gradually become difficult to cope with complex steganographic algorithms and diverse image contents. In recent years, the booming development of deep learning technology has brought new opportunities for steganalysis. Steganalysis methods based on deep learning significantly improve the detection performance by automatically learning image features. Existing methods have achieved remarkable detection results in steganalysis tasks. These methods usually assume that the training data and test data come from the same distribution, that is, the source domain and the target domain are consistent. However, in practical applications, the target images may come from data sources different from the training samples, and the labeled data in the target domain is usually unavailable. This inconsistency between the training data and the test data is called carrier source mismatch (CSM). The CSM problem is very common in practical scenarios and may lead to a significant decline in the detection accuracy of steganalysis. Therefore, how to improve the detection accuracy of steganalysis models under cross-domain conditions is the key challenge for applying steganalysis technology to practical scenarios. In recent years, a large number of explorations have been carried out by researchers on the carrier source mismatch problem in steganalysis. The CSM problem in steganalysis is similar to the unsupervised domain adaptation (UDA) task in the field of computer vision. Therefore, some UDA methods have been introduced into the steganalysis task to alleviate the performance degradation caused by the distribution difference between the source domain and the target domain. J-Net is the first network specifically for the CSM problem in deep steganalysis. By integrating the joint maximum mean discrepancy (JMMD) into deep steganalysis, it minimizes the joint maximum average difference between the source and the target, thus alleviating the performance degradation caused by CSM. Based on J-Net, subsequent studies further considered the domain difference and class difference, and achieved sub-domain alignment by minimizing the local maximum mean discrepancy, further reducing the performance degradation of steganalysis caused by different carrier sources. ISNet generates an intermediate domain by designing a local mixing module, enabling the network to gradually adapt to the steganographic noise distribution in the target domain, thereby improving the discriminability on the target domain. Although existing methods have alleviated the CSM problem to a certain extent, most of them rely on data from a single source domain and have limited performance in complex scenarios with multi-source domain data.In addition, existing methods still have deficiencies in adapting to the steganographic noise distribution in the target domain, especially when the steganographic noise is weak and difficult to distinguish. Summary of the Invention
[0003] To solve the deficiencies mentioned in the above background art, the purpose of the present invention is to provide a steganography detection method and system based on multi-source domain adaptation.
[0004] In a first aspect, the purpose of the present invention can be achieved through the following technical solutions: A steganography detection method based on multi-source domain adaptation, the method comprising the following steps: Receiving multiple source domain and target domain data, and inputting the multiple source domain and target domain data into a pre-established common feature extractor to extract the feature representations of the source domain and the feature representations of the target domain; Inputting the feature representations of the source domain and the feature representations of the target domain into a pre-established feature-level mixing module for weighted mixing to obtain mixed features, obtaining target domain features, and calculating a mixed class loss based on the target domain features and the mixed features; Inputting the data of each source domain and target domain into a pre-established domain feature extractor with non-shared weights, outputting a domain alignment loss, and inputting the mixed class loss and the domain alignment loss into a pre-established domain-specific classifier to output a final prediction result.
[0005] In combination with the first aspect, in some implementation manners of the first aspect, the method further includes: The data sources of the multiple source domain and target domain data are: labeled data from N source domains , where and respectively represent the data of the domain and its corresponding label, and the unlabeled data of the target domain . .
[0006] In combination with the first aspect, in some implementation manners of the first aspect, the method further includes: The process of inputting the multiple source domain and target domain data into a pre-established common feature extractor: Passing through a common feature extractor G, which maps the multiple source domain and target domain data from the original feature space to the common feature space; The common feature extractor is composed of the first three types of convolutional layers designed in SRNet. For a batch of data from the source domain and the data of the target domain , the feature representations obtained after passing through the common feature extractor and are given by the formula: The feature representations obtained from each source domain will all be paired up in pairs and input into the feature-level mixing module to additionally generate mixed features.
[0007] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the input feature-level mixing module receives the output features from the common feature extractor and , and then calculates the mixing ratio of the target domain according to the current training epoch , and the update of the mixing ratio follows the following formula: where are respectively the maximum values of the mixing ratio, is the parameter controlling the exponential growth rate, current_epoch and total_epoch respectively represent the current training epoch and the total number of training epochs; According to the calculated mixing ratio , the source domain features and the target domain features are weighted and summed according to the mixing ratio to obtain the mixed features , and its formula is: where represents the addition operation between matrices, represents the th sample in this batch, and respectively represent the source domain and target domain output features from the common feature extractor and the th sample; In addition to feature mixing, the labels also need to be mixed, and the label of the mixed feature is calculated using the same mixing ratio as the feature. Here is the prediction result of the network on the target domain, is the label of the jth sample: Combined with the first aspect, in some implementations of the first aspect, the method further includes: aligning the data of the source domain and the target domain, and the domain feature extractor F i receives the features from a batch of data from the source domain and the features from the target domain and , and map each pair of source domain and target domain into a specific feature space, then execute the first-stage alignment strategy to align the feature distributions of the two domains. Use the Maximum Mean Discrepancy (MMD) to perform the first-stage alignment, so as to judge whether the two distributions and are the same. Based on the observed sample data, MMD defines the difference measure between distributions through the following formula: where represents the distance measure between two probability distributions and . H is the Reproducing Kernel Hilbert Space (RKHS) with the feature kernel k. and are samples from different distributions respectively. represents mapping the original sample to a certain feature map in the RKHS. represents sampling from the distribution and calculating the expected value of the function . represents sampling from the distribution and calculating the expected value of the function . The feature kernel k represents . . , represents the inner product of vectors. When and only when , and have the same distribution. MMD estimates the difference between distributions by comparing the squared distance of the empirical kernel mean embeddings. Combining with the input of the domain feature extractor layer, the formula is as follows: Here is the unbiased estimator of . H is the Reproducing Kernel Hilbert Space (RKHS) with the feature kernel k. and are the number of samples in a batch. and are the outputs of the source domain and target domain processed by the common feature extractor. and represent the samples among them respectively. Obtain the estimate of the difference between each source domain and target domain. According to MMD, design the domain alignment loss MMD Loss of this layer as: where F i represents the i-th domain feature extractor. and The outputs of the source domain and the target domain processed by the common feature extractor. N domain feature extractors that do not share weights map each pair of source domain and target domain data into a specific feature space, and the resulting domain alignment loss , by minimizing during the backpropagation process to achieve feature distribution alignment.
[0008] Combined with the first aspect, in some implementations of the first aspect, the method further includes: inputting the processed features into a pre-established domain-specific classifier to obtain the predicted outputs of the source, target, and mixed features: Based on the alignment strategy of the loss, measure the prediction differences of different classifiers on the target domain data, The form of the loss function is as follows: where and respectively represent the predicted probabilities of the classifiers corresponding to the source domain and branches for the th target domain sample. By minimizing the loss, is the number of samples in a batch; finally, the mean of the prediction results of all classifiers on the target domain data is used as the final prediction result for output.
[0009] Second aspect, to achieve the above object, the present invention discloses a steganography detection system based on multi-source domain adaptation, including: A feature extraction module, configured to receive multiple source domain and target domain data, input the multiple source domain and target domain data into a pre-established common feature extractor, and extract the feature representations of the source domain and the feature representations of the target domain; A feature mixing module, configured to pair the feature representations of the source domain and the feature representations of the target domain and input them into a pre-established feature-level mixing module for weighted mixing to obtain mixed features; A feature prediction module, configured to input the mixed features and the paired feature representations of the source domain and the feature representations of the target domain into a domain feature extractor to obtain processed features, and input the processed features into a pre-established domain-specific classifier to output the predicted outputs of the source, target, and mixed features.
[0010] In another aspect of the present invention, to achieve the above object, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, it adopts a steganography detection method based on multi-source domain adaptation as described above.
[0011] In another aspect of the present invention, in order to achieve the above object, a computer-readable storage medium is disclosed. A computer program is stored in the computer-readable storage medium. When the computer program is loaded and executed by a processor, a steganography detection method based on multi-source domain adaptation as described above is adopted.
[0012] Advantages of the present invention: The present invention aims to solve the problem of the decline in detection performance of image steganalysis in practical applications due to different sources of the sample datasets to be detected and the lack of label information. Through the multi-source two-stage alignment strategy and the gradually changing feature-level hybrid module, the detection accuracy and generalization ability of the steganography detection model under cross-domain conditions are effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings; Figure 1 It is a schematic flowchart of the method of the present invention; Figure 2 It is a schematic diagram of the overall structure of the steganography detection network based on multi-source domain adaptation of the present invention; Figure 3 It is a schematic diagram of the structure of the progressive feature fusion module of the present invention; Figure 4 It is a schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0015] Embodiment 1: As Figure 1 shown, a steganography detection method based on multi-source domain adaptation, the method includes the following steps: S101: Receive data of multiple source domains and a target domain, input the data of the multiple source domains and the target domain into a pre-established common feature extractor, and extract the feature representations of the source domains and the feature representation of the target domain; As Figure 2As shown, source domain 1 to source domain N respectively represent the labeled samples of N different domains for domain adaptation and network training. The target domain refers to the unlabeled samples from the target domain. The common feature extractor is a convolutional neural network composed of the first three convolutional layers of SRNet, which is used to initially extract the features of the samples. The feature-level mixing module in the figure is specifically as Figure 3 shown, and is used to perform weighted mixing on each pair of samples from the source and the target to generate mixed features. In the figure, the domain-specific extractors are all composed of the last convolutional layer of SRNet, but are independent of each other and do not share parameters. This extractor is used to further perform convolutional processing on the features and calculate the feature alignment loss using the processed features. The domain-specific classifier in the figure is composed of a linear classification layer, and each classifier is independent of each other and does not share parameters. It is used to predict the feature category and provide the prediction output for the calculation of the cross-entropy loss. The average in the figure refers to the operation of taking the average of the prediction outputs of each subnet classifier and using this average as the prediction output of the sample. The domain alignment loss in the figure is the MMD Loss mentioned above, which refers to the domain alignment loss calculated by the maximum mean discrepancy. The classifier loss refers to the classification alignment loss calculated by the L1 norm for the prediction outputs of different classifiers. The average in the figure refers to taking the average of the prediction outputs of different domain-specific classifiers as the final prediction output of the network. The carrier image and the stego image in the figure respectively represent the classification results of the network. The cross-entropy loss in the figure refers to the classification loss calculated by combining the classification result of the network with the true label using the cross-entropy loss function.
[0016] As Figure 3 shown, this figure takes the first branch of the network as an example and is used to represent the process of feature mixing. Among them, G represents the Figure 2 common feature extractor in this and respectively represent the features of the samples from the first source domain and the target domain after being processed by the common feature extractor. Here and respectively represent the mixing weights of the target domain features and the source domain features. represents the features after mixing by weight. H1 represents the domain-specific extractor of the first branch network.
[0017] Assume that the present invention has labeled data from N source domains , where and respectively represent the data of domain and its corresponding label, as well as the unlabeled data of the target domain . After these data are input into the network, they first pass through a common feature extractor G, which maps the data from the original feature space to the common feature space.
[0018] The common feature extractor used in this invention consists of the first three types of convolutional layers designed in SRNet. For a batch of data from the source domain and the data of the target domain , the feature representations obtained through the common feature extractor and and are formulated as follows: Each feature representation obtained from a source domain will be paired with and input into the feature-level mixing module in pairs to additionally generate a mixed feature.
[0019] S102: Pair the feature representations of the source domain and the target domain and input them into a pre-established feature-level mixing module for weighted mixing to obtain a mixed feature; For the progressive feature fusion module as shown in Figure 3 , fuse the feature distribution information of samples from the source domain and the target domain to enable the network to more directly learn the steganographic traces of the target domain: As shown in Figure 3 , taking the source domain and the target as an example, the core idea of the feature-level mixing module is to perform weighted mixing on the features from the source domain and the target domain in a certain proportion. This proportion changes progressively during the training process. The mixing proportion is adjusted in an exponential growth manner in each training round, and the features of the target domain gradually occupy a larger proportion in the mixed feature, enabling the network to gradually transition to the feature distribution of the target domain during the training process. Specifically, the domain feature-level mixing module receives the output features and from the common feature extractor. Subsequently, this invention calculates the mixing proportion of the target domain according to the current training round. The update of the mixing proportion follows the following formula: are the maximum values of the mixing proportion, respectively, with the value set to 0.9. is the parameter controlling the exponential growth rate. current_epoch and total_epoch represent the current training round and the total number of training rounds, respectively.
[0020] According to the calculated mixing proportion , the source domain features and the target domain features are weighted and summed according to this proportion to obtain the mixed feature , and its formula is: Represents the addition operation between matrices, indicating weighted fusion at the feature level. Indicates the th sample in this batch. And respectively represent the source domain and target domain output features from the common feature extractor. And The th sample in. In addition to feature mixing, labels also need to be mixed accordingly. The label of the mixed feature is calculated using the same mixing ratio as the feature: Since the target domain samples are unlabeled, the prediction result of the network on the target domain is used here as the target domain sample label. The target domain feature and the mixed feature are jointly fed into the subsequent network and the cross-entropy loss is calculated according to the mixed label as the classification loss of the mixed domain. Among them, is the prediction output of the network for the mixed feature, is the number of samples in a batch.
[0021] Since there are large distribution differences between different carrier sources themselves, and it is difficult to align data from multiple domains, which is prone to negative transfer, this network designs N domain feature extractors for N source-target pairs to align each source domain with the target domain respectively. The domain feature extractor F i Receives the features from a batch of source domains and the features from the target domain, and maps each pair of source domain and target domain to a specific feature space, and then executes the first-stage alignment strategy to align the feature distributions of the two domains. The present invention uses the maximum mean discrepancy MMD to perform the first-stage alignment. The maximum mean discrepancy (MMD) is a kernel-based two-sample test method for judging whether two distributions and are the same, based on the observed sample data. MMD defines the difference measure between distributions through the following formula: Among them, represents the distance measure between two probability distributions and , H is the reproducing kernel Hilbert space RKHS with feature kernel k, and are samples from different distributions respectively. denotes mapping the original sample to a certain feature map in the RKHS, denotes sampling from the distribution and calculating the expected value of the function ; denotes sampling from the distribution and calculating the expected value of the function ; The feature kernel k denotes where denotes the inner product of vectors. Theoretically, when and only when , the distributions of and are the same. In practical applications, MMD estimates the difference between distributions by comparing the squared distances of empirical kernel mean embeddings. Combining the input of the domain feature extractor layer, the formula is as follows: where is an unbiased estimator of , and H is the reproducing kernel Hilbert space RKHS with the feature kernel k. and are the number of samples in a batch, and are the outputs of the source domain and the target domain processed by the common feature extractor, and represent the samples therein respectively. The estimation of the difference between each source domain and the target domain is obtained according to this formula. According to MMD, the domain alignment loss MMD Loss of this layer is designed as: where Fi represents the i-th domain feature extractor, and are the outputs of the source domain and the target domain processed by the common feature extractor. S103: Input the mixed features, the paired feature representations of the source domain and the target domain into the domain feature extractor to obtain the processed features, and input the processed features into the pre-established domain-specific classifier to obtain the predicted outputs of the source, target, and mixed features. N domain feature extractors without shared weights map each pair of source domain and target domain data to a specific feature space, and the obtained domain alignment loss
[0022] . Feature distribution alignment is achieved by minimizing in the backpropagation process, and feature representations unaffected by the domain are learned. The present invention uses the last convolutional layer of SRNet as the domain feature extractor.
[0023] These features are then fed into their respective domain-specific classifiers for classification. To reduce the prediction bias caused by classifier differences, the present invention introduces a strategy based on prediction probability alignment, which enhances the model's generalization ability for target-domain samples by constraining the outputs of different classifiers on the target domain.
[0024] To achieve this goal, the present invention designs an alignment strategy based on loss to measure the prediction differences of different classifiers on target-domain data. The form of this loss function is as follows: where and respectively represent the prediction probabilities of the classifiers corresponding to the source-domain and branches for the j-th target-domain sample. By minimizing this loss, the output bias of different classifiers on the target domain can be significantly reduced, making the decision boundaries more consistent.
[0025] Finally, the mean of the prediction results of all classifiers on the target-domain data is used as the final prediction output. Through the two-stage alignment steganalysis framework of the present invention, the distribution differences between different source domains and target domains can be well aligned, and more discriminative domain-invariant feature representations can be learned from multiple domains.
[0026] Finally, all losses are weighted and summed as the total loss for network training. By performing domain adaptation optimization on the network, the detection accuracy on the target domain is improved.
[0027] Experiments were conducted using the common steganographic algorithm S - UNIWARD with an embedding rate of 0.4bpp on four different datasets: BOSSbase 1.01, ALASKA#2, MIRFlickr 25k, and ImageNet. Among them, BOSSbase contains 10,000 grayscale images in PGM format with a size of 512×512. ALASKA#2 contains 80,000 images in various formats. MIRFlickr 25 contains 25,000 color JPEG images of different sizes. ImageNet contains pictures of different sizes from 1000 categories in the real world. Subsequently, in this invention, B, A, M, and I are used to represent the four datasets. 10,000 pictures were randomly selected from each dataset for experiments. The network was pre - trained with 8000 pairs of pictures for training and 2000 pairs of pictures for validation. In the domain adaptation stage, 500 pairs of pictures were used for training and testing respectively. The batch size in the domain adaptation stage was set to 8, and a total of 60 rounds of training were carried out in the domain adaptation stage. A comparison was made with the backbone network SRNet of this invention. First, the problem of the detection performance decline in the CSM situation was tested through experiments. For example, here B - B represents training using the data on BOSSbase and testing on the data from the same source. While B - A represents the CSM situation, that is, the result of training on BOSSbase and testing on ALASKA#2.
[0028] Table 1 Detection results of the backbone network SRNet for stego - images detected by the S - UNIWARD algorithm when CSM occurs The occurrence of the CSM situation leads to a high probability of a significant decline in the performance of the network trained on the source domain when detecting on the target domain. The solution proposed in this invention can effectively improve the cross - domain detection performance by using multiple source domains for domain - adaptation training. When using two source domains, experiments were conducted under the same scenario as shown in Table 2: Table 2 Detection effect of the network proposed in this paper in the case of dual - source domains and comparison of detection accuracy with the backbone network Here, the "best accuracy on the source domain" refers to the maximum value among the detection accuracies obtained when SRNet detects target - domain samples on two source domains respectively. From the experimental results, it can be seen that the method proposed in this invention can effectively extract invariant features of multiple domains through operations such as feature mixing and aligning the feature distributions of multiple domains, and improve the image steganography detection accuracy of the network in the CSM scenario.
[0029] Example 2: As Figure 4 shown, to achieve the above - mentioned purpose, this invention discloses a steganography detection system based on multi - source domain adaptation, including: A feature extraction module 11, configured to receive multiple source domain and target domain data, input the multiple source domain and target domain data into a pre-established common feature extractor, and extract the feature representations of the source domain and the target domain; A feature mixing module 12, configured to pair the feature representations of the source domain and the target domain and input them into a pre-established feature-level mixing module for weighted mixing to obtain mixed features; A feature prediction module 13, configured to input the mixed features, and the paired feature representations of the source domain and the target domain into a domain feature extractor to obtain processed features, and input the processed features into a pre-established domain-specific classifier to output the prediction outputs of the source, target, and mixed features.
[0030] Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions, specifically for loading and executing one or more instructions in the computer storage medium to implement the above method.
[0031] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.
[0032] In the description of this specification, the descriptions referring to the terms "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0033] The above has shown and described the basic principles, main features, and advantages of the present disclosure. Those skilled in the art of this industry should understand that the present disclosure is not limited by the above embodiments, and the above embodiments and the descriptions in the specification only illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements all fall within the scope of the present disclosure claimed.
Claims
1. A steganography detection method based on multi-source domain adaptation, characterized in that The method includes the following steps: Receiving a plurality of source domain and target domain data, inputting the plurality of source domain and target domain data into a pre-established common feature extractor, and extracting the feature representations of the source domain and the target domain; Pairing the feature representations of the source domain and the target domain and inputting them into a pre-established feature-level mixing module for weighted mixing to obtain mixed features; Inputting the mixed features, as well as the paired feature representations of the source domain and the target domain, into a domain feature extractor to obtain processed features, and inputting the processed features into a pre-established domain-specific classifier to obtain the predicted outputs of the source, target, and mixed features.
2. The steganography detection method based on multi-source domain adaptation according to claim 1, wherein The data sources of the multiple source domain and target domain data are: labeled data from N source domains , where and represent the data and its corresponding labels of domain , and the unlabeled data of the target domain , where i represents the i-th source domain .
3. A steganography detection method based on multi-source domain adaptation according to claim 1, characterized in that The process of inputting the plurality of source domain and target domain data into a pre-established common feature extractor: Passing through a common feature extractor G, which maps the plurality of source domain and target domain data from the original feature space to the common feature space; The common feature extractor is composed of the first three types of convolutional layers designed in SRNet. For a batch of data from the source domain , and the data of the target domain , the feature representations obtained through the common feature extractor and are formulated as follows: Feature representations obtained from each source domain will be paired with in pairs and input into the feature-level mixing module to additionally generate mixed features.
4. A steganography detection method based on multi-source domain adaptation according to claim 3, characterized in that The input feature-level mixing module receives the source domain and target domain output features from the common feature extractor and , and then calculates the mixing ratio of the target domain according to the current number of training epochs , and the update of the mixing ratio follows the following formula: Among them, are respectively the maximum values of the mixing ratio, is a parameter that controls the exponential growth rate, where current_epoch and total_epoch represent the current training epoch and the total number of training epochs respectively; According to the calculated mixing ratio , the source domain features and the target domain features are weighted and summed according to the mixing ratio to obtain the mixed features , and its formula is: Among them, represents the addition operation between matrices, indicates the th sample in this batch, and respectively represent the source domain and target domain output features from the common feature extractor and the th sample; In addition to feature mixing, labels also need to be mixed, and the labels of the mixed features are calculated using the same mixing ratio as the features, which is the prediction result of the network on the target domain, and is the label of the j-th sample: Use as the mixed-domain sample label, the target-domain feature and the mixed feature are jointly fed into the subsequent network and the cross-entropy loss is calculated according to the mixed label as the classification loss of the mixed domain, where is the predicted output of the network for the mixed feature, is the number of samples in a batch: 。 5. A steganography detection method based on multi-source domain adaptation according to claim 1, characterized in that Align the data of the source domain and the target domain, and the domain feature extractor F i receives features from a batch of source domain features and features from the target domain features , and maps each pair of the source domain and the target domain into a specific feature space. Subsequently, execute the first-stage alignment strategy to align the feature distributions of the two domains. Use the maximum mean discrepancy (MMD) to perform the first-stage alignment, thereby judging whether the two distributions and are the same, based on the observed sample data. MMD defines the difference measure between distributions through the following formula: Among them, represents the distance metric between two probability distributions and where \(H\) is the Reproducing Kernel Hilbert Space (RKHS) with characteristic kernel \(k\), and are samples from different distributions respectively, represents mapping the original samples to a certain feature map in the RKHS, represents sampling from the distribution and calculating the expected value of the function ; represents sampling from the distribution and calculating the expected value of the function ; the characteristic kernel \(k\) represents , represents the inner product of vectors, and when and only when , and have the same distribution. MMD estimates the difference between distributions by comparing the squared distance of the empirical kernel mean embeddings, combined with the input of the domain feature extractor layer, and the formula is as follows: Here is an unbiased estimator of, where H is a reproducing kernel Hilbert space RKHS with characteristic kernel k; and is the number of samples in a batch, and are the outputs of the source domain and the target domain processed by a common feature extractor, and represent the samples therein respectively, so as to obtain an estimate of the difference between each source domain and the target domain. According to MMD, the domain alignment loss MMD Loss of this layer is designed as: Among which F i represents the i-th domain feature extractor, and are the outputs of the source domain and the target domain processed by the common feature extractor; N domain feature extractors without shared weights map each pair of source domain and target domain data into a specific feature space, and the obtained domain alignment loss , by minimizing during the backpropagation process to achieve feature distribution alignment.
6. A steganography detection method based on multi-source domain adaptation according to claim 1, characterized in that, The process of inputting the processed features into a pre-established domain-specific classifier to obtain the predicted outputs of the source, target, and mixed features: Based on the loss-based alignment strategy to measure the prediction differences of different classifiers on the target domain data, the loss function has the following form: Among them and represent the prediction probabilities of the classifiers corresponding to the source domain and for the th target domain sample respectively, is the number of samples in a batch; by minimizing the loss, finally, the mean value of the prediction results of all classifiers for the target domain data is used as the final prediction result for output.
7. A steganography detection system based on multi-source domain adaptation, which adopts a steganography detection method based on multi-source domain adaptation described in any one of claims 1 to 6, characterized in that Includes: A feature extraction module, configured to receive a plurality of source domain and target domain data, input the plurality of source domain and target domain data into a pre-established common feature extractor, and extract the feature representations of the source domain and the target domain; A feature mixing module, configured to pair the feature representations of the source domain and the target domain and input them into a pre-established feature-level mixing module for weighted mixing to obtain mixed features; A feature prediction module, configured to input the mixed features, as well as the paired feature representations of the source domain and the target domain, into a domain feature extractor to obtain processed features, and input the processed features into a pre-established domain-specific classifier to output the predicted outputs of the source, target, and mixed features.
8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, The memory stores a computer program that can run on a processor. When the processor loads and executes the computer program, it adopts a steganography detection method based on multi-source domain adaptation according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is loaded and executed by the processor, it adopts a steganography detection method based on multi-source domain adaptation according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image steganalysis method for carrier source mismatch problem
CN115439447A
CSM-oriented steganalysis network training method and device based on progressive intermediate domain
CN117115552A
Adversarial dual-classifier deep steganalysis network training method and device
CN118799620A
Cited By
Cross-domain steganography text recognition and analysis method and system based on multi-adversarial domain adaptation
CN121960500A