Image steganography information detection method and device
Through the image steganography information detection method based on twin neural networks, color channel similarity calculation and dynamic threshold adjustment are used to solve the accuracy and applicability of LSB steganography information detection in the prior art, and efficient identification and extraction of steganography information is realized.
Patent Information
- Application Number
- CN202510376810.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
AI Technical Summary
The existing LSB steganographic information detection methods have insufficient identification accuracy and applicability, and it is difficult to effectively distinguish images with sensitive information from normal images. Conventional detection methods cannot recognize diversified steganographic data.
The image steganography information detection method based on twin neural network is adopted. By separating the images into multiple color channels, image vector pairs are constructed, and the similarity between color channels is calculated using twin neural networks, and the similarity threshold is dynamically adjusted to identify and extract steganography information.
It improves the accuracy of identification of steganographic information, expands the scope of application of detection, and can flexibly supervise the dissemination of sensitive data on the network.
Smart Images

Figure CN120298722A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular, to a method and apparatus for identifying steganographic information in images. Background Art
[0002] Among numerous image steganography techniques, the least significant bit (LSB) steganography in the spatial domain is widely used due to its advantages such as good concealment, large information hiding capacity, and easy implementation. Therefore, it is of great significance to effectively, accurately, and reliably detect LSB information steganography for protecting information security and preventing covert communication. After using the LSB steganography technique, it is difficult to distinguish an image with sensitive information from a normal image, and it is impossible to use conventional means (for example, data leakage prevention (DLP) file detection, optical character recognition (OCR) image scanning) for detection, and it is impossible to monitor the leakage of sensitive information.
[0003] Currently, the methods for identifying LSB hiding information technology include:
[0004] 1. Methods implemented based on statistics. The problem with this method is that the image content varies greatly, and it is impossible to obtain accurate theoretical distribution data with universal applicability. In addition, too much normal data is cited during sampling, masking the feature changes caused by abnormal data and reducing the accuracy of identification.
[0005] 2. Methods for establishing an abnormal image feature model based on machine learning. Since the content of steganographic data is diverse, the established feature model can only be effective for a certain type of steganographic data. It cannot identify abnormal images that have been steganographed with other types of data that have not been trained. This method has limitations in applicability. Summary of the Invention
[0006] The present disclosure provides a method and apparatus for identifying steganographic information in images. Through this method and apparatus, the accuracy of identifying abnormal images including steganographic information can be greatly improved, and / or the detection intensity can be dynamically adjusted, thereby expanding the applicable range of this identification method.
[0007] According to one aspect of the present disclosure, there is provided a method for detecting steganographic information in an image, which includes: separating the image to be recognized into at least three color channels, and respectively constructing image vectors based on each color channel based on the at least three color channels; combining the image vectors of the at least three color channels in pairs to form pairs of image vectors of multiple color channels; respectively inputting the image vectors of each color channel in the pairs of image vectors of multiple color channels into a trained siamese neural network to generate the similarity between the pairs of image vectors of each color channel; and determining whether there is steganographic information in the image to be recognized based on the similarity between the pairs of image vectors of each color channel generated.
[0008] In one embodiment, separating the image to be recognized into at least three color channels includes separating the image to be recognized into three color channels of R, G, and B based on the RGB mode or separating the image to be recognized into four color channels of C, M, Y, and K based on the CMYK mode.
[0009] In one embodiment, the siamese neural network includes two siamese neural sub-networks with shared weights.
[0010] In one embodiment, the method further includes: when it is determined that there is steganographic information in the image to be recognized, determining the color channels including steganographic information based on the similarity between the pairs of image vectors of each color channel; and extracting the steganographic information based on the determined color channels including steganographic information.
[0011] In one embodiment, extracting the steganographic information based on the determined color channels including steganographic information includes: extracting the steganographic information based on the information of the least significant bit in the determined color channels including steganographic information.
[0012] In one embodiment, determining whether there is steganographic information in the image to be recognized based on the similarity between the pairs of image vectors of each color channel may include: determining whether there is steganographic information in the image to be recognized based on the comparison between the similarity between the pairs of image vectors of each color channel and a predetermined threshold.
[0013] In one embodiment, the output layer of the siamese neural network uses the sigmoid function or the softmax function as the activation function.
[0014] In one embodiment, the predetermined threshold is an adaptively adjustable threshold.
[0015] In one embodiment, the similarity is determined based on the Euclidean distance of the outputs of the two siamese neural sub-networks.
[0016] According to another aspect of the present disclosure, there is provided a system for detecting steganographic information in images, including: a feature extraction module configured to separate an image to be recognized into at least three color channels, construct image vectors based on each color channel respectively based on the at least three color channels, and combine the image vectors of the at least three color channels in pairs to form pairs of image vectors of multiple color channels; a similarity determination module configured to receive the image vectors of each color channel in the pairs of image vectors of multiple color channels as inputs and output the similarity between the image vectors of each pair of color channels; and a steganography analysis module configured to determine whether there is steganographic information in the image to be recognized based on the similarity between the generated pairs of image vectors of each color channel.
[0017] In one embodiment, the similarity determination module includes a trained siamese neural network.
[0018] In a further embodiment, the system for detecting steganographic information in images may further include: a similarity threshold adjustment module configured to adjust the similarity threshold, and the steganography analysis module may also be configured to determine whether there is steganographic information in the image to be recognized based on the comparison between the similarity between the generated pairs of image vectors of each color channel and the similarity threshold.
[0019] According to another aspect of the present disclosure, there is provided a siamese neural network system, which includes two siamese neural sub-networks with shared weights and a third output layer. Among them, each siamese neural sub-network includes: an input layer configured to receive pairs of image vectors of color channels constructed based on multiple color channels constituting the image to be recognized; a hidden layer configured to identify features in the input pairs of image vectors of color channels and perform non-linear transformation; and among them, the input layer, hidden layer and output layer of each siamese neural sub-network are connected in a fully connected manner, and an output layer configured to output the features of the input pairs of image vectors of color channels, where the third output layer is configured to output the similarity between the input pairs of image vectors of color channels based on the outputs of the first output layers of the two siamese neural sub-networks.
[0020] In one embodiment, the second output layer uses the sigmoid function or the softmax function as the activation function.
[0021] In one embodiment, the input layer and the hidden layer include neurons with the number of pixels in the image vectors of each color channel, and the first output layer includes neurons with half the number of pixels in the image vectors of each color channel.
[0022] According to another aspect of the present disclosure, there is provided a method for training a siamese neural network for detecting steganographic information in images, which includes: a. Separating image samples in an image sample set into at least three color channels, and constructing image vectors based on each color channel respectively based on the at least three color channels; b. Pairing the image vectors of at least three color channels pairwise to form pairs of image vectors of multiple color channels; c. Receiving the image vectors of each color channel in the pairs of image vectors of multiple color channels and inputting them into the siamese neural network respectively to generate the similarity between the pairs of image vectors of each color channel; and d. Adjusting the parameters of the siamese neural network based on the loss between the generated similarity between the pairs of image vectors of each color channel and the true value of the similarity of the image samples.
[0023] In one embodiment, the image sample set includes: an original image sample set and a reverse image sample set constructed based on the original sample image set in which the least significant bits of pixel values are modified.
[0024] According to another aspect of the present disclosure, there is provided a system for training a siamese neural network, including: a sample feature extraction module configured to separate image samples in an image sample set into at least three color channels, construct image vectors based on each color channel respectively based on the at least three color channels, and pair the image vectors of at least three color channels pairwise to form pairs of image vectors of multiple color channels; a sample similarity determination module configured to receive the image vectors of each color channel in the pairs of image vectors of multiple color channels as inputs to generate the similarity between the pairs of image vectors of each color channel; and a loss calculation and parameter adjustment module configured to adjust the parameters of the siamese neural network based on the loss between the generated similarity between the pairs of image vectors of each color channel and the true value of the similarity of the image samples.
[0025] In a further embodiment, the image sample set includes: an original image sample set and a reverse image sample set constructed based on the original sample image set in which the least significant bits of pixel values are modified.
[0026] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing the above-mentioned software or computer program, wherein when the computer program is executed by at least one processor, it causes the at least one processor to execute any one of the above-mentioned methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 FIG. shows a schematic flowchart of a method for detecting steganographic information in an image according to an embodiment of the present disclosure.
[0028] Figure 2Shows a schematic block diagram of an image steganography information detection system according to an embodiment of the present disclosure.
[0029] Figure 3 Shows a schematic block diagram of a twin neural network system according to an embodiment of the present disclosure.
[0030] Figure 4 Shows a training method of a twin neural network for a steganography detection system according to an embodiment of the present disclosure.
[0031] Figure 5 Shows a system for training a twin neural network according to an embodiment of the present disclosure. Detailed implementation manners
[0032] Before describing the following detailed implementation manners, it may be advantageous to elaborate on the definitions of certain words and phrases used throughout this patent document. The terms "send", "receive", and "communicate" and their derivatives cover direct and indirect communication. The terms "include" and "comprise" and their derivatives mean including but not limited to. The term "or" is inclusive and means and / or. The phrase "at least one of..." when used with a list of items means that different combinations of one or more of the listed items can be used and may only require one item from the list.
[0033] In addition, the various functions described below can be implemented or supported by one or more computer programs, each of which is formed by computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, processes, functions, objects, classes, instances, related data, or portions thereof suitable for implementation in a suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as a read-only memory (ROM), a random access memory (RAM), a hard disk drive, a compact disc (CD), a digital video disc (DVD), or any other type of memory. "Non-transitory" computer-readable media excludes wired, wireless, optical, or other communication links that transmit transitory electrical signals or other signals. Non-transitory computer-readable media include media that can permanently store data and media that can store and later rewrite data, such as rewritable compact discs or erasable memory devices.
[0034] It should be understood that the "first", "second" and similar terms used in this disclosure do not denote any order, quantity or importance, but are merely used to distinguish different components. Unless the context clearly indicates otherwise, the singular forms "a", "an" or "the" and similar terms do not denote a limitation of quantity, but rather indicate the presence of at least one.
[0035] The various embodiments discussed below for describing the principles of this disclosure in this patent document are for illustration only and should not be construed in any way as limiting the scope of this disclosure.
[0036] The following description with reference to the accompanying drawings is provided to facilitate a comprehensive understanding of the various embodiments of this disclosure defined by the claims and their equivalents. This description includes various specific details to facilitate understanding but should only be considered exemplary. Accordingly, those of ordinary skill in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of this disclosure. In addition, descriptions of well-known functions and structures may be omitted for clarity and conciseness.
[0037] Unless otherwise defined, all terms (including technical or scientific terms) used in this disclosure have the same meaning as understood by those of ordinary skill in the art to which this disclosure pertains. Ordinary terms defined in a dictionary are interpreted to have a meaning consistent with the context in the relevant technical field and should not be interpreted idealistically or overly formally unless explicitly so defined in this disclosure.
[0038] The following discussion Figures 1 to 5 and the various embodiments for describing the principles of this disclosure in this patent document are for illustration only and should not be construed in any way as limiting the scope of this disclosure. Those skilled in the art will understand that the principles of this disclosure can be implemented in any suitably arranged system or device.
[0039] In order not to affect the overall visual perception of the image, the information steganographically embedded by the common LSB steganography method basically only modifies the data of one color channel of the image, so that the image of the modified color channel will have slight differences from the images of other normal channels in terms of light changes, line changes, etc.
[0040] Based on the above observations, the present disclosure provides an image steganography information detection method based on a siamese neural network that can dynamically adjust the recognition intensity. The method is divided into two working stages: in the first stage, a siamese neural network model is constructed using deep learning techniques, and the model is trained to calculate the similarity of images in each color channel; in the second stage, the steganography information is detected and extracted. When an image contains steganography information, the similarity of the images in two of its color channels will be significantly higher than that of the other. By calculating the similarity of the color channel images, the color channel where the steganography information is located can be located, and thus the steganography information can be extracted from the least significant bit data of the color channel. The method can also dynamically adjust the detection intensity. Specifically, the maximum value of the similarity difference between color channel images when determining the existence of steganography information can be adjusted (this adjustment can be based on the analysis of image samples or based on empirical values), so as to improve the accuracy of identifying images including steganography information and expand the applicable range of the method for images with different steganography information, thereby better monitoring the spread of sensitive data in the network.
[0041] Figure 1 FIG. shows a flowchart of a method for detecting image steganography information according to an embodiment of the present disclosure.
[0042] The method 100 begins at step 110. In step 110, a plurality of image vectors corresponding to each color channel among the plurality of color channels constituting the image to be recognized can be constructed for the image to be recognized.
[0043] In one embodiment, the image vectors constructed based on each color channel can reflect the respective pixel values of the image based on each color channel.
[0044] In one embodiment, the image to be recognized can be separated into three color channels R, G, and B according to the RGB mode, and image vectors can be constructed for the images under the three color channels R, G, and B respectively. In another embodiment, the image to be recognized can be separated into four color channels C, M, Y, and K according to the CMYK mode, and image vectors can be constructed for the images under the four color channels C, M, Y, and K respectively.
[0045] In step 120, a pair of image vectors of two color channels of the image to be recognized can be generated based on the constructed image vectors of the plurality of color channels. More specifically, this step involves pairing the image vectors of different color channels in pairs to form a plurality of image vector pairs, and each image vector pair is composed of image vectors from two different color channels.
[0046] In one example, if the image to be recognized is separated into three color channels: R (Red), G (Green), and B (Blue), then pairing the image vectors of different color channels pairwise to form multiple image vector pairs may include forming an image vector pair of the R and G color channels (which can be represented as an image vector pair composed of the image vector "V R " under the R color channel and the image vector "V G " under the G color channel); an image vector pair of the R and B color channels (which can be represented as an image vector pair composed of the image vector "V R " under the R color channel and the image vector "V B " under the B color channel); and an image vector pair of the B and G color channels (which can be represented as an image vector pair composed of the image vector "V B " under the B color channel and the image vector "V G " under the G color channel). That is to say, three image vector pairs are formed.
[0047] In one example, if the image to be recognized is separated into four color channels: C (Cyan), M (Magenta), Y (Yellow), and K (Key), then pairing the image vectors of different color channels pairwise to form multiple image vector pairs may include forming an image vector pair of the C and M color channels (which can be represented as an image vector pair composed of the image vector "V C " under the C color channel and the image vector "V M " under the M color channel); an image vector pair of the C and Y color channels (which can be represented as an image vector pair composed of the image vector "V C " under the C color channel and the image vector "V B " under the Y color channel); an image vector pair of the C and K color channels (which can be represented as an image vector pair composed of the image vector "V C " under the C color channel and the image vector "V K " under the K color channel); an image vector pair of the M and Y color channels (which can be represented as an image vector pair composed of the image vector "V M " under the M color channel and the image vector "V Y " under the Y color channel); and an image vector pair of the M and K color channels (which can be represented as an image vector pair composed of the image vector "V M " under the M color channel and the image vector "V K " under the K color channel). That is to say, six image vector pairs are formed.
[0048] In step 130, the image vectors in each pair of image vectors are respectively input into the trained siamese neural network. The siamese neural network can be used to generate a similarity measure between the image vector pairs of color channels. The siamese neural network may include two siamese neural sub-networks with shared weights.
[0049] In one example, if an image is separated into three color channels R, G, and B, inputting the image vectors in each pair of image vectors into the siamese neural network respectively may include inputting the image vectors "V R " and the image vector "V G " in the image vectors of the R and G color channels into the siamese neural network respectively; inputting the image vectors "V R " and the image vector "V B " in the image vector pair of the R and B color channels into the siamese neural network respectively; and inputting the image vectors "V B " and the image vectors in the image vector pair of the B and G color channels into the siamese neural network respectively.
[0050] Step 140 involves generating a similarity measure between all pairs of image vectors of two color channels. As described above, generating a similarity measure between pairs of image vectors of two color channels can be achieved through a siamese neural network.
[0051] In one example, if an image is separated into three color channels R, G, and B, generating a similarity measure between pairs of image vectors of two color channels includes generating a similarity measure between the image vectors "V R " and the image vector "V G "; generating a similarity measure between the image vectors "V R " and the image vector "V B "; and generating a similarity measure between the image vectors "V B " and the image vectors.
[0052] In one embodiment, the similarity is determined based on the Euclidean distance of the outputs of the siamese neural sub-networks included in the siamese neural network. In a specific example, the numerical range of the output similarity may be {0, 1}. A similarity value close to 1 can be used to indicate a large feature difference between two image vectors, and a similarity value close to 0 can be used to indicate a small feature difference between two image vectors. In a specific embodiment, the siamese neural network uses the sigmoid function or the softmax function as the activation function.
[0053] In the above-mentioned twin neural network, the image features used are extracted from the image vectors of the color channels that make up the image, thereby making the features to be extracted in the neural network clearer and making the image steganography information detection method according to the embodiments of the present disclosure more sensitive.
[0054] In step 150, it is possible to determine whether there is steganography information in the image to be recognized based on the similarity between the image vector pairs of each generated color channel. In a further example, when it is determined that there is steganography information in the image to be recognized, the detection method may further include extracting the steganography information in the image to be recognized.
[0055] When the image to be recognized contains steganography information implemented by the common LSB steganography method, the similarity of the image vector pairs of some color channels will be significantly higher than that of the image vector pairs of other color channels. In the case where a similarity value close to 1 indicates a large feature difference between two image vectors and a similarity value close to 0 indicates a small feature difference between two image vectors, if the R color channel contains steganography information, then the similarity of the image vectors between the R and B color channels and the similarity of the image vectors between the R and G color channels are both close to 1, indicating a large feature difference between the R and B color channels and between the R and G color channels, and the similarity of the image vectors between the B and G color channels is close to 0, indicating that the features between the B and G color channels are similar. Furthermore, the channel R where the steganography information is located can be located, and then the steganography information can be extracted from the least significant bit data of the located R color channel.
[0056] Therefore, as described above, when the similarity between the image vector pairs of some color channels is significantly higher than that of the image vector pairs of other color channels, it can be determined that there is steganography information in the image to be recognized; otherwise, it is considered that there is no steganography information in the image to be recognized.
[0057] More specifically, when dividing the color channels of the image to be recognized in the RGB mode, when the similarity between two image vector pairs (for example, the similarity between V R and V G and the similarity between V R and V B ) is higher than the first threshold, while the similarity between the image vector pairs of one color channel (for example, the similarity between V B and V G ) is lower than the first threshold, it can be determined that there is steganography information in the image to be recognized; otherwise, it is considered that there is no steganography information in the image to be recognized.
[0058] In one embodiment, when it is determined that the image to be recognized has steganographic information, the detection method more specifically includes: determining the color channels including steganographic information based on the similarity between image vector pairs of each color channel; and extracting the steganographic information based on the determined color channels including steganographic information. In the above example, when dividing the color channels of the image to be recognized in the RGB mode, when the similarity between two image vector pairs (e.g., the similarity between V R and V G and the similarity between V R and V B ) is higher than the first threshold (i.e., the feature difference is larger), while the similarity between image vector pairs of one color channel (e.g., the similarity between V B and V G ) is lower than the first threshold (i.e., the feature difference is smaller), it can be determined that the common color channel involved in the two image vector pairs with lower similarity (i.e., larger feature difference) includes steganographic information, that is, the R color channel includes steganographic information.
[0059] The first threshold can be a default value. For example, it can be determined via the average value of the similarity true values obtained during the training process of the image sample set. In addition, the first threshold can also be determined based on the experience value of the detector.
[0060] In a preferred embodiment, the above first threshold can be dynamically adjusted, so as to adjust the recognition intensity of the steganographic information in the image to be recognized. Specifically, the first threshold can be adjusted manually and dynamically according to needs through an external interface module. By dynamically adjusting the first threshold, the detection accuracy of images with steganographic information can be improved without retraining the siamese neural network, and the range of images with steganographic information that can be detected is expanded, so that the present detection method can adapt to more network scenarios, and thus can more flexibly and accurately monitor the spread of possible sensitive information in the network.
[0061] In a specific example, if the value of the first threshold is adjusted to a larger value, then the similarity between the image vectors of two color channels can only be considered that the image of the color channel involved includes steganographic information when it is higher than a specific size.
[0062] Through the above method, the process of generating similarity can be separated from the process of determining the steganographic information of the image. On the one hand, the artificial neural network is used more precisely to extract the features related to the color channels that may include steganographic information. On the other hand, the similarity threshold for determining the steganographic information of the image can be adjusted manually and dynamically according to needs, so that the method according to the present disclosure can adapt to more network scenarios and more flexibly and accurately monitor the spread of sensitive information in the network.
[0063] In a further embodiment, extracting steganographic information based on the determined color channel including steganographic information may include: extracting steganographic information based on the information of the least significant bit in the determined color channel including steganographic information. Continuing with the above example, steganographic information can be extracted from the least significant bit of the R color channel.
[0064] Figure 2 FIG. 4 shows a schematic block diagram of an image steganographic information detection system according to an embodiment of the present disclosure.
[0065] The image steganographic information detection system (or simply referred to as the "steganography detection system") 200 includes three functional modules, namely, a feature extraction module 210, a similarity calculation module 220, and a steganography analysis module 230.
[0066] The feature extraction module 210 may be configured to separate the image to be recognized into at least three color channels, construct an image vector according to each color channel, and combine two by two among the multiple color channels to generate an image vector pair. In a more specific example, the feature extraction module 210 may be configured to construct multiple image vectors of multiple color channels constituting the image for the image sample to be recognized. This involves separating the image to be recognized into at least three color channels and constructing an image vector based on each color channel.
[0067] In one embodiment, the image vector constructed based on each color channel reflects the respective pixel values of the image based on each color channel.
[0068] In one embodiment, the image to be recognized may be separated into three color channels R, G, and B according to the RGB mode, and image vectors are constructed for the images under the three color channels R, G, and B respectively. In another embodiment, the image to be recognized may be separated into four color channels C, M, Y, and K according to the CMYK mode, and image vectors are constructed for the images under the four color channels C, M, Y, and K respectively.
[0069] The feature extraction module 210 may also be configured to generate an image vector pair of two color channels of the image to be recognized based on the constructed image vectors of multiple color channels. This step involves pairing the image vectors of different color channels two by two to form multiple image vector pairs, and each image vector pair is composed of image vectors from two different color channels.
[0070] In one example, if the image is separated into three color channels R, G, and B, pairing the image vectors of different color channels two by two to form multiple image vector pairs may include forming an image vector pair of the two color channels R and G (which can be expressed as the image vector "V R " under the R color channel and the image vector "V G” (image vector pairs formed); image vector pairs of the R and B color channels (which can be represented as the image vector “V R ” under the R color channel and the image vector “V B ” under the B color channel, forming an image vector pair), and image vector pairs of the B and G color channels (which can be represented as the image vector “V B ” under the B color channel and the image vector “V G ” under the G color channel, forming an image vector pair). That is to say, three image vector pairs are formed.
[0071] In one example, if an image is separated into four color channels C, M, Y, and K, pairing the image vectors of different color channels two by two to form multiple image vector pairs can include forming an image vector pair of the C and M color channels (which can be represented as the image vector “V C ” under the C color channel and the image vector “V M ” under the M color channel, forming an image vector pair); an image vector pair of the C and Y color channels (which can be represented as the image vector “V C ” under the C color channel and the image vector “V B ” under the Y color channel, forming an image vector pair); an image vector pair of the C and K color channels (which can be represented as the image vector “V C ” under the C color channel and the image vector “V K ” under the K color channel, forming an image vector pair); an image vector pair of the M and Y color channels (which can be represented as the image vector “V M ” under the M color channel and the image vector “V Y ” under the Y color channel, forming an image vector pair); and an image vector pair of the M and K color channels (which can be represented as the image vector “V M ” under the M color channel and the image vector “V K ” under the K color channel, forming an image vector pair). That is to say, six image vector pairs are formed.
[0072] The similarity calculation module 220 can be configured to receive image vector pairs of multiple color channels as input and output the similarity between each pair of image vectors. The similarity calculation module 220 can be specifically implemented as a trained siamese neural network. The similarity calculation module can be configured to: input the image vectors in each pair of image vectors into the trained siamese neural network respectively.
[0073] In one example, if an image is separated into three color channels R, G, and B, inputting the image vectors in each pair of image vectors into the siamese neural network respectively can include inputting the image vector “V R” and the image vector “V G ” are respectively input into the siamese neural network; the image vectors “V R ” and the image vector “V B ” in the image vector pairs of the R and B color channels are respectively input into the siamese neural network; and the image vectors “V B ” and the image vectors are respectively input into the siamese neural network.
[0074] The similarity calculation module 220 can also be configured to generate a similarity metric between the image vector pairs of two color channels.
[0075] In one example, if the image is separated into three color channels of R, G, and B, generating the similarity metric between the image vector pairs of two color channels includes generating the similarity metric between the image vector “V R ” and the image vector “V G ”; generating the similarity metric between the image vector “V R ” and the image vector “V B ”; and generating the similarity metric between the image vector “V B ” and the image vector.
[0076] In one embodiment, the similarity is determined based on the Euclidean distance of the outputs of the siamese neural sub-networks included in the siamese neural network. In a specific example, the numerical range of the output similarity can be {0, 1}. A similarity value close to 1 can be used to indicate a large feature difference between two image vectors, and a similarity value close to 0 can be used to indicate a small feature difference between two image vectors. In a specific embodiment, the siamese neural network uses the sigmoid function or the softmax function as the activation function.
[0077] The steganalysis module 230 can be configured to determine whether the image contains steganographic information according to the similarity of the generated image vector pairs.
[0078] As described above, when the similarity between the image vector pairs of some color channels is significantly higher than that of the other color channel image vector pairs, it can be determined that the image to be recognized contains steganographic information; otherwise, it is considered that the image to be recognized does not contain steganographic information.
[0079] More specifically, when dividing the color channels of the image to be recognized in the RGB mode, when the similarity between two image vector pairs (e.g., the similarity between V R and V G and the similarity between V R and V Bthe similarity between) is higher than the first threshold, and when the similarity between the image vector pairs of one color channel (e.g., V B and V G the similarity between) is lower than the first threshold, it can be determined that there is steganographic information in the image to be recognized; otherwise, it is considered that there is no steganographic information in the image to be recognized.
[0080] In one embodiment, the steganalysis module 230 can also be configured to, when it is determined that there is steganographic information in the image to be recognized, determine the color channels including steganographic information based on the similarity between the image vector pairs of each color channel; and extract the steganographic information based on the determined color channels including steganographic information. In the above example, when dividing the color channels of the image to be recognized in the RGB mode, when the similarity between two image vector pairs (e.g., V R and V G the similarity between and V R and V B the similarity between) is higher than the first threshold (i.e., the feature difference is larger), and the similarity between the image vector pairs of one color channel (e.g., V B and V G the similarity between) is lower than the first threshold (i.e., the feature difference is smaller), it can be determined that the common color channel involved in the two image vector pairs with lower similarity (i.e., larger feature difference) includes steganographic information, that is, the R color channel includes steganographic information.
[0081] In a preferred embodiment, the image steganographic information detection system (or simply referred to as "steganalysis system") 200 can further include a similarity threshold adjustment module 250. The similarity threshold adjustment module 250 can be configured to dynamically adjust the above-mentioned first threshold, thereby adjusting the recognition intensity of the steganographic information in the image to be recognized. By dynamically adjusting the first threshold, the detection accuracy of images with steganographic information can be improved without retraining the siamese neural network, the range of images with steganographic information that can be detected is expanded, and different network scenarios are adapted, so that the possible sensitive information in the network can be supervised more flexibly and accurately during transmission.
[0082] In a further embodiment, extracting the steganographic information based on the determined color channels including steganographic information can include: extracting the steganographic information based on the information of the least significant bit in the determined color channels including steganographic information. Continuing the above example, in the case where it is determined that the R color channel includes steganographic information, the steganographic information can be extracted from the least significant bit of the R color channel.
[0083] Figure 3 FIG. shows a schematic block diagram of a siamese neural network system according to an embodiment of the present disclosure.
[0084] The Siamese neural network system 300 includes two Siamese neural sub-networks 310 and 320, and a third output layer 330. The Siamese neural sub-networks 310 and 320 can be identical Siamese neural sub-networks, and the two can share weights.
[0085] The Siamese neural sub-network 310 can include a first input layer 311, a first hidden layer 312, and a first output layer 313. A fully connected connection can be adopted between the first input layer 311 and the first hidden layer 312, and between the first hidden layer 312 and the first output layer 313. More specifically, the first input layer 311 and the first hidden layer 312 can include the same number of neurons as the number of pixels of the image vector of one color channel mentioned above as the input, and the first output layer 313 can include half the number of neurons as the number of pixels of the image vector of one color channel mentioned above as the input.
[0086] Similarly, the Siamese neural sub-network 320 includes a second input layer 321, a second hidden layer 322, and a second output layer 323. A fully connected connection can be adopted between the second input layer 321 and the second hidden layer 322, and between the second hidden layer 322 and the second output layer 323. More specifically, the second input layer 321 and the second hidden layer 322 can include the same number of neurons as the number of pixels of the image vector of one color channel mentioned above as the input, and the second output layer 323 can include half the number of neurons as the number of pixels of the image vector of one color channel mentioned above as the input.
[0087] The input of the third output layer 330 can be based on the output of the first output layer 313 of the Siamese neural sub-network 310 and the output of the second output layer 323 of the Siamese neural sub-network 320. Specifically, the input of the third output layer 330 can be the Euclidean distance between the output of the first output layer 313 of the Siamese neural sub-network 310 and the output of the second output layer 323 of the Siamese neural sub-network 320. In addition, an activation function such as the sigmoid or softmax function can be used in the third output layer 330 to generate the similarity between the outputs of the Siamese neural sub-network 310 and the Siamese neural sub-network 320, that is, the similarity between the two input image vectors of color channels.
[0088] The Siamese neural sub-network 320 can be used to implement Figure 1 step 140 described in Figure 2 and the similarity calculation module in
[0089] Figure 4A schematic flowchart showing a twin neural network system for training to detect image steganographic information according to an embodiment of the present disclosure is shown.
[0090] All image samples in the image sample set for training will be used in this training process 400.
[0091] The training process 400 starts at step 410. In step 410, multiple image vectors corresponding to multiple color channels that make up the image sample can be constructed based on the image samples in the image sample set. This involves separating the image sample into at least three color channels and constructing image vectors based on each channel.
[0092] In one embodiment, the image vectors constructed based on each color channel can reflect the respective pixel values of the image based on each color channel.
[0093] In one embodiment, the image sample can be separated into three color channels of R, G, and B according to the RGB mode, and image vectors are constructed for the images under the three color channels of R, G, and B respectively. In another embodiment, the image sample can be separated into four color channels of C, M, Y, and K according to the CMYK mode, and image vectors are constructed for the images under the four color channels of C, M, Y, and K respectively.
[0094] In step 420, an image vector pair of two color channels of the image sample can be generated based on the constructed image vectors of multiple color channels. This step involves pairing the image vectors of different color channels in pairs to form multiple image vector pairs, and each image vector pair consists of image vectors from two different color channels.
[0095] In one example, if the image sample is separated into three color channels of R, G, and B, pairing the image vectors of different color channels in pairs to form multiple image vector pairs can include forming an image vector pair of the R and G color channels (which can be represented as an image vector pair composed of the image vector "V R " under the R color channel and the image vector "V G " under the G color channel); an image vector pair of the R and B color channels (which can be represented as an image vector pair composed of the image vector "V R " under the R color channel and the image vector "V B " under the B color channel); and an image vector pair of the B and G color channels (which can be represented as an image vector pair composed of the image vector "V B " under the B color channel and the image vector "V G " under the G color channel), that is, forming three image vector pairs.
[0096] In one example, if an image sample is separated into four color channels C, M, Y, and K, then pairing the image vectors of different color channels in pairs to form a plurality of image vector pairs may include forming an image vector pair of the C and M color channels (which can be represented as the image vector "V C " under the C color channel and the image vector "V M " under the M color channel); an image vector pair of the C and Y color channels (which can be represented as the image vector "V C " under the C color channel and the image vector "V B " under the Y color channel); an image vector pair of the C and K color channels (which can be represented as the image vector "V C " under the C color channel and the image vector "V K " under the K color channel); an image vector pair of the M and Y color channels (which can be represented as the image vector "V M " under the M color channel and the image vector "V Y " under the Y color channel); and an image vector pair of the M and K color channels (which can be represented as the image vector "V M " under the M color channel and the image vector "V K " under the K color channel), that is, six image vector pairs are formed.
[0097] In step 430, the image vectors in each pair of image vectors are respectively input into the twin neural network to be trained. The twin neural network can be used to generate a similarity measure between the image vector pairs of the color channels.
[0098] In one example, for an image sample, if an image sample is separated into three color channels R, G, and B, then inputting the image vectors in each pair of image vectors into the twin neural network respectively may include inputting the image vector "V R " and the image vector "V G " in the image vectors of the R and G color channels into the twin neural network respectively; inputting the image vector "V R " and the image vector "V B " in the image vector pair of the R and B color channels into the twin neural network respectively; and inputting the image vector "V B " and the image vector into the twin neural network in the image vector pair of the B and G color channels respectively.
[0099] Step 440 involves generating a similarity measure between all the image vector pairs of the two color channels by the twin neural network.
[0100] In one example, for an image sample, if the image sample is separated into three color channels of R, G, and B, the similarity measure between the generated image vector pairs of two color channels includes generating the similarity measure between the image vector "V R " and the image vector "V G "; generating the similarity measure between the image vector "V R " and the image vector "V B "; and generating the similarity measure between the image vector "V B " and the image vector.
[0101] In one embodiment, the similarity is determined based on the Euclidean distance of the outputs of the twin neural sub-networks included in the twin neural network. In a specific example, the numerical range of the output similarity can be from 0 to 1. A similarity value close to 1 can be used to indicate a large feature difference between two image vectors, and a similarity value close to 0 can be used to indicate a small feature difference between two image vectors. In a specific embodiment, the twin neural network uses the sigmoid function or the softmax function as the activation function.
[0102] Finally, step 450 involves adjusting the parameters of the twin neural network based on the loss between the similarity between the generated image vector pairs of each color channel in step 440 and the ground truth similarity of the image sample.
[0103] In one example, the loss between the similarity generated by the neural network and the ground truth similarity is calculated by a Contrastive Loss function or a Triplet Loss function.
[0104] This training process 400 will be performed based on all the image samples in the image sample set. In addition, another image sample set can also be generated based on the original image sample set, so as to increase the sample base and achieve training the twin neural network with an increased sample size to obtain more accurate adjusted model parameters.
[0105] In one example, the image sample set for training can include: the original image sample set and the reverse image sample set constructed based on the original sample images whose least significant bits of pixel values are modified.
[0106] In the scenario of deeply detecting network traffic, the sample features constructed by using this method are more obvious. The use of these sample features makes the twin neural network model more targeted and sensitive, and can greatly improve the accuracy of abnormal image recognition.
[0107] Figure 5 Shows a system for training a twin neural network according to an embodiment of the present disclosure.
[0108] The neural network training system 500 can be configured to execute the training method as shown in Figure 4 .
[0109] In one embodiment, the neural network training system 500 can include: a sample feature extraction module 510; a sample similarity determination module 520; and a loss calculation and parameter adjustment module 530.
[0110] The sample feature extraction module 510 can be configured to separate the image samples in the image sample set into at least three color channels, construct image vectors based on each color channel respectively based on the at least three color channels, and combine the image vectors of the at least three color channels in pairs to form pairs of image vectors of multiple color channels.
[0111] The sample similarity determination module 520 can be configured to receive each image vector of the pairs of image vectors of multiple color channels as input to generate the similarity between each pair of image vectors of multiple color channels.
[0112] The loss calculation and parameter adjustment module 530 can be configured to adjust the parameters of the siamese neural network based on the loss between the generated similarity between each pair of image vectors of multiple color channels and the ground truth similarity of the image samples.
[0113] Wherein, the functions of the sample feature extraction module 510 and the sample similarity determination module 520 are similar to those of the feature extraction module 210 and the similarity calculation module 220 shown in Figure 2 and described with respect to Figure 2 . The only difference is that the inputs of the sample feature extraction module 510 and the sample similarity determination module 520 correspond to the sample image set for training, while the inputs of the feature extraction module 210 and the similarity calculation module 220 correspond to the images to be recognized. Therefore, the same parts will not be elaborated here.
[0114] In one embodiment, the sample image set for training as described above can include: an original image sample set and a reverse image sample set constructed based on the original sample image set in which the least significant bits of the pixel values are modified.
[0115] In an embodiment of the present disclosure, by first generating an image vector based on each color channel, then the image vectors of different color channels can be combined to form an image vector pair, the combined image vector pair is input into a neural network to generate a similarity metric, and then the steganographic information is detected and extracted based on the similarity metric. Through the above embodiment, by using a siamese neural network to compare the similarities between the image vectors of multiple color channels, the features extracted for the image to be recognized can be clarified, and the recognition accuracy of abnormal images containing steganographic information can be improved. In addition, the embodiments of the present disclosure can also dynamically adjust the detection intensity for detecting image steganographic information, thereby expanding the applicable range of detection and better supervising the spread of sensitive data in the network.
[0116] According to an embodiment of the present disclosure, a computer-readable storage medium storing the above-mentioned software or computer program may also be provided. When the computer program is executed by at least one processor, the at least one processor is caused to execute or assist in executing any one of the above methods according to the exemplary embodiments of the present disclosure. Examples of such computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store the computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The instructions or computer programs in the above computer-readable storage media can run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0117] Those skilled in the art will understand that the above-described illustrative embodiments are described herein and are not intended to be limiting. It should be understood that any two or more of the embodiments disclosed herein can be combined in any combination. In addition, other embodiments can be utilized and other changes can be made without departing from the spirit and scope of the subject matter presented herein. It will be readily understood that the aspects of the invention of the present disclosure, as generally described herein and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are contemplated herein.
[0118] Those skilled in the art will understand that the various illustrative logical blocks, modules, circuits, and steps described in this application can be implemented as hardware, software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in terms of their functional sets. Whether such a functional set is implemented as hardware or software depends on the particular application and the design constraints imposed on the overall system. A person skilled in the art can implement the described functional sets in different ways for each particular application, but such design decisions should not be construed as causing a departure from the scope of this application.
[0119] The various illustrative logical blocks, modules, and circuits described in this application can be implemented or executed using a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gates or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor can be a microprocessor, but in an alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.
[0120] The steps of the methods or algorithms described in this application can be embodied directly in hardware, in software modules executed by a processor, or in a combination of both. The software modules can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read from and write to the storage medium. In an alternative, the storage medium can be integrated into the processor. The processor and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In an alternative, the processor and the storage medium can reside in the user terminal as discrete components.
[0121] In one or more exemplary designs, the functions can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted via a computer-readable medium as one or more instructions or code. The computer-readable medium includes both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. The storage medium can be any available medium that can be accessed by a general purpose or special purpose computer.
[0122] The foregoing are only exemplary embodiments of this application, and are not intended to limit the scope of protection of this application. The scope of protection of this application is determined by the appended claims.
Claims
1. An image steganography information detection method, comprising: Separating the image to be recognized into at least three color channels, and respectively constructing image vectors based on each color channel based on the at least three color channels; Pairwise combining the image vectors of the at least three color channels to form pairs of image vectors of two color channels among the multiple color channels; Respectively inputting the image vectors of each color channel in the pairs of image vectors of the multiple color channels into a trained siamese neural network to generate the similarity between the pairs of image vectors of each color channel; And Determining whether there is steganography information in the image to be recognized based on the similarity between the pairs of image vectors of each color channel generated.
2. The method according to claim 1, wherein, Separating the image to be recognized into at least three color channels includes separating the image to be recognized into three color channels of R, G, and B based on the RGB mode or separating the image to be recognized into four color channels of C, M, Y, and K based on the CMYK mode.
3. The method according to claim 1, wherein, The siamese neural network includes two siamese neural subnets with shared weights.
4. The method according to claim 1, further comprising: When it is determined that there is steganography information in the image to be recognized, determining the color channels including steganography information based on the similarity between the pairs of image vectors of each color channel; And Extracting steganography information based on the determined color channels including steganography information.
5. The method according to claim 4, wherein, Extracting steganography information based on the determined color channels including steganography information includes: Extracting steganography information based on the information of the least significant bit in the determined color channels including steganography information.
6. The method according to claim 1, wherein Determining whether there is steganography information in the image to be recognized based on the similarity between the pairs of image vectors of each color channel includes: Determining whether there is steganography information in the image to be recognized based on the comparison between the similarity between the pairs of image vectors of each color channel and a predetermined threshold.
7. The method according to claim 6, wherein, The predetermined threshold is an adaptively adjustable threshold.
8. The method according to claim 3, wherein, The similarity is determined based on the Euclidean distance of the outputs of the two siamese neural subnets.
9. The method according to claim 8, wherein, The output layer of the siamese neural network uses the sigmoid function or the softmax function as the activation function.
10. A system for image steganography information detection, comprising: A feature extraction module configured to separate the image to be recognized into at least three color channels, respectively construct image vectors based on each color channel based on the at least three color channels, and pairwise combine the image vectors of the at least three color channels to form pairs of image vectors of two color channels among the multiple color channels; A similarity determination module configured to receive the image vectors of each color channel in the pairs of image vectors of the multiple color channels as inputs and output the similarity between the pairs of image vectors of each color channel; And A steganography analysis module configured to determine whether there is steganography information in the image to be recognized based on the similarity between the pairs of image vectors of each color channel generated.
11. The system according to claim 10, wherein the similarity determination module includes a trained siamese neural network.
12. The system according to claim 11, further comprising: A similarity threshold adjustment module configured to adjust the similarity threshold, and The steganography analysis module is further configured to determine whether there is steganographic information in the image to be recognized based on the comparison between the similarity of the image vectors of each generated color channel and the similarity threshold.
13. A twin neural network system for image steganography information detection, which includes two twin neural sub-networks with shared weights and a third output layer, wherein, Each twin neural sub-network includes: An input layer configured to receive pairs of image vectors based on two color channels among the multiple color channels constituting the image to be recognized; A hidden layer configured to recognize features in the input pairs of image vectors of the color channels and perform non-linear transformation; and An output layer configured to output the features of the input pairs of image vectors of the color channels, wherein, the input layer and the hidden layer of each twin neural sub-network are connected by a fully connected manner, and between the hidden layer and the output layer, and wherein, the third output layer is configured to output the similarity between the input pairs of image vectors of the color channels based on the outputs of the output layers of the two twin neural sub-networks.
14. The system according to claim 13, wherein, The second output layer uses the sigmoid function or the softmax function as the activation function.
15. The system according to claim 13, wherein, The input layer and the hidden layer include neurons corresponding to the number of pixels in the image vector of each color channel, and the first output layer includes neurons corresponding to half of the number of pixels in the image vector of each color channel.
16. A method for training a twin neural network for detecting steganographic information in images, including: Separating the image samples in the image sample set into at least three color channels, and constructing image vectors based on each color channel respectively based on the at least three color channels; Pairing the image vectors of the at least three color channels pairwise to form pairs of image vectors of multiple color channels; Inputting the image vectors of each color channel in the pairs of image vectors of multiple color channels into the twin neural network respectively to generate the similarity between the pairs of image vectors of each color channel; And Adjusting the parameters of the twin neural network based on the loss between the similarity between the pairs of image vectors of each generated color channel and the true similarity value of the image samples.
17. The method according to claim 16, wherein, The image sample set includes: An original image sample set and a reverse image sample set constructed based on the original sample image set with the least significant bits of the pixel values modified.
18. A system for training a twin neural network, including: A sample feature extraction module configured to separate the image samples in the image sample set into at least three color channels, construct image vectors based on each color channel respectively based on the at least three color channels, and pair the image vectors of the at least three color channels pairwise to form pairs of image vectors of multiple color channels; A sample similarity determination module configured to receive the image vectors of each color channel in the pairs of image vectors of multiple color channels as inputs to generate the similarity between the pairs of image vectors of each color channel; and A loss calculation and parameter adjustment module configured to adjust the parameters of the twin neural network based on the loss between the similarity between the pairs of image vectors of each generated color channel and the true similarity value of the image samples.
19. The system according to claim 18, wherein, The image sample set includes: An original image sample set and an inverse image sample set constructed based on the original sample image set in which the least significant bits of pixel values are modified.
20. A non-transitory computer-readable storage medium storing the software or computer program mentioned above, wherein, When the computer program is executed by at least one processor, it causes the at least one processor to execute the method according to any one of claims 1 to 9 and claims 16 to 17.