Cross-modal steganography method, device, equipment and storage medium
By employing a cross-modal steganography method and utilizing a combined architecture of learnable wavelet transform networks and invertible neural networks, the balance between robustness and concealment in existing watermarking steganography techniques is resolved, enabling efficient transmission of encrypted text data in noisy environments.
Patent Information
- Application Number
- CN202411647634.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-11-18
AI Technical Summary
Existing watermarking steganography techniques sacrifice the concealment of the watermark while ensuring robustness, and cannot achieve cross-modal image steganography of encrypted text data, especially when facing noise attacks, the decoding robustness is insufficient.
A combined encoder-decoder architecture of learnable wavelet transform network and invertible neural network is adopted to generate and recover ciphertext cross-modal steganography results data through forward and inverse feature extraction. Combined with self-attention mechanism and multi-frequency domain feature extraction, the feature extraction and recovery capabilities of image and text data are enhanced.
It improves the concealment and robustness of ciphertext steganography, enhances its resistance to various attacks, and achieves more secure and reliable information transmission.
Smart Images

Figure CN119562067B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information hiding, and particularly relates to a cross-modal steganography method and device, equipment and a storage medium. BACKGROUND
[0002] With the rapid development of information technology, information security problems are increasingly prominent. Although traditional cryptography methods can effectively protect the security of data in the transmission process, they are not flexible enough for the secret transmission and authentication of data, especially when information needs to be embedded in public media content. Therefore, information hiding technology, especially image steganography technology, as a covert communication means, has attracted widespread attention. Image steganography technology allows secret information to be hidden in seemingly ordinary images, thereby achieving secret transmission of information. Digital watermarking technology, as an important branch of information hiding, is mainly used to protect the copyright of digital media content, authenticate the source, or embed additional information. Traditional watermark steganography methods usually embed watermarks in the frequency domain of images to improve the invisibility of watermarks. However, their robustness under noise attacks is poor. Recently, deep learning-based methods have achieved better performance than traditional methods because they can better learn the deep features of watermarks and images.
[0003] However, existing watermark steganography techniques sacrifice the concealment of watermarks while ensuring robustness and cannot achieve cross-modal image steganography of ciphertext data. Traditional watermark steganography methods usually embed watermarks in the frequency domain of images, but when subjected to noise attacks, the watermarks will be destroyed and cannot be normally recovered. In recent years, deep learning methods have made some progress in watermark steganography, but these methods often fail to achieve satisfactory results in terms of concealment and robustness. Especially when the coding capacity is too strong, it becomes difficult to guarantee the robustness of decoding, and robustness is the most important for watermark methods. That is, although existing deep learning-based image steganography techniques have played an important role in copyright protection and information security, their balance between robustness and invisibility, and vulnerability to complex environments and malicious attacks, limit their potential for widespread application. Therefore, traditional techniques mainly face two challenges: imperceptibility and robustness, especially when facing various noise attacks. SUMMARY
[0004] In view of this, the embodiments of the present application provide a cross-modal steganography method, device, equipment and storage medium to eliminate or improve one or more defects in the prior art.
[0005] One aspect of the present application provides a cross-modal steganography method, comprising:
[0006] input the carrier image and the ciphertext image corresponding to the ciphertext text data into a preset encoder respectively, so that a learnable wavelet transform network and a reversible neural network in the encoder respectively perform forward feature extraction on the carrier image and the ciphertext image to obtain a target carrier feature vector corresponding to the carrier image and a target ciphertext feature vector corresponding to the ciphertext image;
[0007] generate ciphertext cross-modal steganography result data corresponding to the ciphertext text data based on the target carrier feature vector and the target ciphertext feature vector for network transmission.
[0008] In some embodiments of the present application, further comprising:
[0009] receive a noise image and extract a noise carrier feature vector and a noise ciphertext feature vector corresponding to the noise image, wherein the noise image is formed after network transmission of ciphertext cross-modal steganography result data, and the ciphertext cross-modal steganography result data is generated in advance based on a target carrier feature vector corresponding to a carrier image output by the encoder and a target ciphertext feature vector corresponding to a ciphertext image of a ciphertext text data;
[0010] input the noise carrier feature vector and the noise ciphertext feature vector into a decoder corresponding to the encoder respectively, so that the learnable wavelet transform network and the reversible neural network in the decoder respectively perform inverse feature extraction on the noise carrier feature vector and the noise ciphertext feature vector to obtain a loss carrier image corresponding to the noise carrier feature vector and a restored ciphertext image corresponding to the noise ciphertext feature vector; wherein the loss carrier image corresponding to the noise carrier feature vector is different from the carrier image used to generate the noise image to which the noise carrier feature vector belongs.
[0011] extract information from the ciphertext image to obtain ciphertext text data corresponding to the restored ciphertext image.
[0012] In some embodiments of the present application, the encoder comprises a forward propagation network corresponding to the learnable wavelet transform network, a forward propagation network corresponding to the reversible neural network, and an inverse propagation network corresponding to the learnable wavelet transform network connected in sequence; correspondingly, the inputting of the carrier image and the ciphertext image corresponding to the ciphertext text data into the preset encoder respectively, so that the learnable wavelet transform network and the reversible neural network in the encoder respectively perform forward feature extraction on the carrier image and the ciphertext image to obtain the target carrier feature vector corresponding to the carrier image and the target ciphertext feature vector corresponding to the ciphertext image, comprises:
[0013] The ciphertext image corresponding to the carrier image and the ciphertext text data is respectively input into a preset encoder, so that the forward transmission network corresponding to the learnable wavelet transform network in the encoder respectively performs down-sampling feature extraction on the carrier image and the ciphertext image for multiple different frequency domains, and outputs the first multi-frequency domain feature vector corresponding to the carrier image and the ciphertext image respectively; then the forward transmission network corresponding to the reversible neural network in the encoder respectively performs feature extraction on the first multi-frequency domain feature vector corresponding to the carrier image and the ciphertext image based on a self-attention mechanism for target operation, and outputs the second multi-frequency domain feature vector corresponding to the carrier image and the ciphertext image respectively; and then the inverse transmission network corresponding to the learnable wavelet transform network in the encoder respectively performs up-sampling feature extraction on the second multi-frequency domain feature vector corresponding to the carrier image and the ciphertext image for the same frequency domain, and outputs the target carrier feature vector corresponding to the carrier image and the target ciphertext feature vector corresponding to the ciphertext image.
[0014] In some embodiments of the present application, the decoder comprises: the forward transmission network corresponding to the learnable wavelet transform network, the inverse transmission network corresponding to the reversible neural network, and the inverse transmission network corresponding to the learnable wavelet transform network connected in sequence;
[0015] Correspondingly, the noise carrier feature vector and the noise ciphertext feature vector are respectively input into a decoder corresponding to the encoder, so that the learnable wavelet transform network and the reversible neural network in the decoder respectively perform inverse feature extraction on the noise carrier feature vector and the noise ciphertext feature vector to obtain a loss carrier image corresponding to the noise carrier feature vector and a restored ciphertext image corresponding to the noise ciphertext feature vector, comprising:
[0016] The noise carrier feature vector and the noise ciphertext feature vector are respectively input into the decoder, so that a forward transmission network of the learnable wavelet transform network in the decoder respectively performs down-sampling feature extraction on the noise carrier feature vector and the noise ciphertext feature vector for multiple different frequency domains, and outputs a third multi-frequency domain feature vector corresponding to each of the noise carrier feature vector and the noise ciphertext feature vector; then, a corresponding inverse transmission network of the reversible neural network in the encoder respectively performs feature restoration on the third multi-frequency domain feature vector corresponding to each of the noise carrier feature vector and the noise ciphertext feature vector based on a self-attention mechanism for target operation, and outputs a fourth multi-frequency domain feature vector corresponding to each of the noise carrier feature vector and the noise ciphertext feature vector; and then, a corresponding inverse transmission network of the learnable wavelet transform network in the encoder respectively performs up-sampling feature extraction on the fourth multi-frequency domain feature vector corresponding to each of the noise carrier feature vector and the noise ciphertext feature vector for the same frequency domain, and outputs a loss carrier image corresponding to the noise ciphertext feature vector and a restored ciphertext image corresponding to the noise ciphertext feature vector.
[0017] In some embodiments of the present application, before the ciphertext image corresponding to the carrier image and the ciphertext text data is respectively input into the preset encoder, the method further comprises:
[0018] The learnable wavelet transform network and the reversible neural network are trained based on a preset target loss function and training data, so that the learnable wavelet transform network and the reversible neural network respectively constitute the encoder and the decoder; wherein the training data is composed of each data sample, and each data sample contains a carrier image, a ciphertext image corresponding to ciphertext text data and a noise image with the same size; the ciphertext image corresponding to the ciphertext text data is generated by pre-processing the ciphertext text data by watermark diffusion; the noise image is generated by pre-simulating noise attack on the ciphertext cross-modal steganography result data corresponding to the noise image based on a plurality of preset noise types;
[0019] The learnable wavelet transform network is composed of a plurality of convolutional neural networks for simulating Haar wavelet conversion corresponding to each different frequency domain;
[0020] The reversible neural network is composed of a plurality of deformation convolution layers, and each layer of the deformation convolution layers is associated based on a target operation, and the target operation includes at least one of addition, subtraction, multiplication and division;
[0021] The deformable convolution layer is provided with a deformable attention dense block for performing a preset operation function, the deformable attention dense block comprising a self-attention block and a deformable dense block connected in sequence; wherein the deformable dense block comprises a deformable convolution layer and four convolution layers connected in sequence with the deformable convolution layer.
[0022] In some embodiments of the present application, the target loss function is composed of an encoding loss, an image recovery loss, a message decoding loss and a contrastive learning loss, and weights corresponding to the encoding loss, the image recovery loss, the message decoding loss and the contrastive learning loss; the weights of the encoding loss, the image recovery loss and the contrastive learning loss are positive values, and the weight of the message decoding loss is a negative value.
[0023] The encoding loss is used to represent the pixel-level difference value and the image loss value between the target ciphertext feature vector extracted by the encoder and the ciphertext image, and between the target carrier feature vector extracted by the encoder and the carrier image;
[0024] The image recovery loss is used to represent the pixel-level difference value between the loss carrier image output by the decoder and the carrier image;
[0025] The message decoding loss is used to represent the difference value between the restored ciphertext image output by the decoder and the ciphertext image input to the encoder, which is calculated based on the mean square error;
[0026] The contrastive learning loss is used to represent the contrastive learning loss value between the encoder and the decoder obtained based on a joint contrastive learning unit; wherein the joint contrastive learning unit comprises an encoding contrastive learning unit and a decoding contrastive learning unit connected in sequence.
[0027] In some embodiments of the present application, the information extraction on the ciphertext image to obtain the ciphertext text data corresponding to the restored ciphertext image comprises:
[0028] The restored ciphertext image is input into a preset decoding optimization module based on an attention mechanism and a convolution layer, so that the decoding optimization module outputs the ciphertext text data corresponding to the restored ciphertext image.
[0029] Another aspect of the present application provides a cross-modal steganography device, comprising:
[0030] An encoding module is configured to input a ciphertext image corresponding to a carrier image and ciphertext text data into a preset encoder, so that a learnable wavelet transform network and a reversible neural network in the encoder perform forward feature extraction on the carrier image and the ciphertext image, respectively, to obtain a target carrier feature vector corresponding to the carrier image and a target ciphertext feature vector corresponding to the ciphertext image.
[0031] a steganography module configured to generate steganography result data corresponding to the ciphertext text data based on the target carrier feature vector and the target ciphertext feature vector for network transmission.
[0032] A third aspect of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the cross-modal steganography method when executing the computer program.
[0033] A fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the cross-modal steganography method.
[0034] A fifth aspect of the present application provides a computer program product comprising a computer program and / or instructions, wherein the computer program and / or instructions are executed by a processor to implement the cross-modal steganography method.
[0035] The cross-modal steganography method provided by the present application can embed the ciphertext text data of the text modality into the carrier image of the image modality through the algorithm of deep learning, implement the steganography technology of multi-modal, effectively improve the concealment and robustness of the ciphertext steganography by adopting the combination of the reversible neural network and the end-to-end network, thereby not only improving the concealment of the hidden ciphertext text data, but also enhancing the resistance to various attacks, and realizing more secure and reliable information transmission.
[0036] Additional advantages, objects, and features of the application will be set forth in part by the description that follows, and will become apparent to those skilled in the art upon examination of the following figures and detailed description thereof, or can be learned by practice of the application. The advantages and objects of the application can be realized and attained by the structure particularly pointed out in the written description and claims hereof, as well as the appended drawings.
[0037] It will be appreciated by those skilled in the art that the objectives and advantages of the present application can not be limited to the above specifically described, and the above and other objectives that can be achieved by the present application will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0038] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the principles of the application. The components in the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the application. For purposes of clarity and a consistent approach, portions of the drawings may have been exaggerated from the actual scale, which may render some of the portions larger in size than they would be in actual manufacture. In the drawings:
[0039] Figure 1 A first flowchart of a cross-modal steganography method in an embodiment of the present application.
[0040] Figure 2 A second flowchart of a cross-modal steganography method in an embodiment of the present application.
[0041] Figure 3 An architecture diagram of an encoder in a cross-modal steganography method in an embodiment of the present application.
[0042] Figure 4 A third flowchart of a cross-modal steganography method in an embodiment of the present application.
[0043] Figure 5 An architecture diagram of a decoder in a cross-modal steganography method in an embodiment of the present application.
[0044] Figure 6 A fourth flowchart of a cross-modal steganography method in an embodiment of the present application.
[0045] Figure 7 An architecture diagram of a learnable wavelet transform network in an embodiment of the present application.
[0046] FIG. 8(a) is an example architecture diagram of a forward propagation network corresponding to a network architecture of a deformable attention invertible neural network (DA-INN) in an embodiment of the present application.
[0047] FIG. 8(b) is an example architecture diagram of a reverse propagation network corresponding to the network architecture of the DA-INN in an embodiment of the present application.
[0048] Figure 9 An example architecture diagram of a deformable attention dense block (DA-dense block) in an embodiment of the present application.
[0049] Figure 10 A flowchart of a cross-modal steganography method in an application example of the present application.
[0050] Figure 11A noise result schematic diagram formed after noise simulation on the non-noise attack (Identity) encrypted image of part of the noise types in an application example of the present application.
[0051] Figure 12 An architecture schematic diagram of the joint contrast learning unit in an application example of the present application.
[0052] Figure 13 A first structure schematic diagram of the cross-modal steganography device in an embodiment of the present application.
[0053] Figure 14 A second structure schematic diagram of the cross-modal steganography device in an embodiment of the present application. DETAILED DESCRIPTION
[0054] In order to make the objects, technical solutions and advantages of the present application clearer, further detailed description will be given to the present application in combination with the embodiments and drawings. Herein, the illustrative embodiments of the present application and the description thereof are used to explain the present application, but not as a limitation to the present application.
[0055] It should be noted that, in order to avoid the present application being obscured by unnecessary details, only the structures and / or processing steps closely related to the scheme according to the present application are shown in the drawings, and other details not closely related to the present application are omitted.
[0056] It should be emphasized that the term “comprise / comprising” is used herein to indicate the presence of a feature, element, step or component, but not to exclude the presence or addition of one or more other features, elements, steps or components.
[0057] It should be noted that, if not specifically stated, the term “connection” herein can not only mean direct connection, but also indirect connection with an intermediate.
[0058] Hereinafter, the embodiments of the present application will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0059] The existing deep learning method is divided into an end-to-end method and a method using an invertible neural network (INN).
[0060] Typical end-to-end methods achieve watermark embedding and recovery through end-to-end training. Constraints or optimization measures more suitable for such methods often enable stronger concealment and robustness. Some works use adversarial networks to optimize the encoding and decoding capabilities of the model. In addition, some works focus on watermarking in specific scenarios, such as decoding messages from wild videos, screen capture or multi-source image detection, and embedding watermarks into images of arbitrary resolution. They employ different optimization methods to improve the performance of the model. However, in such methods, the decoding output is obtained through approximate likelihood inference rather than exact calculation, which makes the robustness of watermark extraction vulnerable. Usually, imperceptibility needs to be sacrificed to gain robustness. Therefore, the robustness and concealment of such methods are often unsatisfactory.
[0061] Invertible neural networks (INN) are the key framework of normalizing flow models. Due to the characteristic of preserving information losslessly, INN becomes a valuable choice for watermark steganography tasks. In most neural network-based methods, such as multi-scene multi-task learning model HiNet, CIN which combines reversible and irreversible mechanisms, and IRWArt, a novel invisible robust watermarking framework, both encoding and decoding are done. The process uses a shared INN, which can extract watermarks more accurately than end-to-end methods. For robustness issues, neural network-based methods are effective, as they use a shared reversible neural network for watermark embedding and extraction. Decoding can be seen as the inverse process of encoding, so they can robustly extract watermarks in exact form. However, they also have inherent limitations. First, neural network-based digital watermarking methods struggle to maintain robustness against complex noise such as affine, rotation, etc. In addition, there is a lack of specialized optimization strategies to improve the encoder-decoder capabilities of the model. Therefore, there is still much room for improvement in enhancing robustness and imperceptibility, especially when dealing with complex noise. Among them, INN (Invertible Neural Network) is a special neural network architecture that is unique in that both its forward propagation and backward propagation processes are reversible. This means that the input of the network can be fully recovered from the output through the same network structure and parameters, and vice versa. This feature makes INN excel in tasks that require precise input restoration.
[0062] IRWArt is a novel invisible robust watermarking framework. In this architecture, watermark embedding and recovery are treated as a pair of inverse image transformations, and a reversible neural network INN can be implemented through the inverse and forward processes. To achieve high visual quality, the watermark is embedded in the high-frequency domain with minimal impact on the artwork, and a deep perceptual loss consistent with the human visual system (HVS) is used to supervise image reconstruction. At the same time, a quality enhancement module is constructed to address distortions that may arise due to plagiarism.
[0063] CIN consists of a reversible part and a non-reversible part, where the reversible part achieves high concealment, and the non-reversible part enhances the robustness to strong noise attacks. In the reversible part, diffusion and extraction modules DEM and fusion and split modules FSM are used to embed and extract watermarks in a symmetric manner. In the non-reversible part, attention-based non-reversible modules NIAM and noise-specific selection modules NSM are introduced to solve the asymmetric extraction problem under strong noise attacks.
[0064] However, the existing watermark steganography technology has the disadvantage of sacrificing the concealment of the watermark while ensuring the robustness. Traditional watermark steganography methods usually embed watermarks in the frequency domain of the image, but when subjected to noise attacks, the watermark will be destroyed and cannot be recovered normally. In recent years, deep learning methods have made some progress in watermark steganography, but these methods often have unsatisfactory results in terms of concealment and robustness. Especially when the coding ability is too strong, it becomes difficult to guarantee the robustness of decoding, and robustness is the most important for watermark methods. In addition, although the watermark steganography method based on reversible neural network INN ensures concealment while maintaining high robustness under most noise, it has poor robustness under complex noise (such as affine transformation, etc.), and lacks a special optimization strategy to improve the coding-decoding ability of the model. Therefore, the existing blind watermarking technology still has a lot of room for improvement in improving robustness and concealment.
[0065] Based on this, in order to realize ciphertext cross-modal steganography and improve the concealment and robustness of ciphertext steganography, the embodiments of the present application respectively provide a cross-modal steganography method, a cross-modal steganography device for executing the cross-modal steganography method, an entity device, a computer readable storage and a computer program product, which embeds ciphertext information (i.e. ciphertext text data) into a carrier image through a specific algorithm, while not significantly changing the visual effect of the image, ensures that the embedded ciphertext text can withstand various environmental changes or malicious analysis tests, and maintains its recoverability. Through the combination of image and text, two different modal data, not only can the concealment of information hiding be improved, but also the resistance to various attacks can be enhanced, realizing more secure and high-robustness information transmission.
[0066] The embodiments are specifically described as follows.
[0067] Based on this, the embodiments of the present application provide a cross-modal steganography method that can be implemented by a cross-modal steganography device, referring to Figure 1 , the cross-modal steganography method specifically includes the following contents:
[0068] Step 100: inputting the carrier image and the ciphertext image corresponding to the ciphertext text data into a preset encoder respectively, so that a learnable wavelet transform network and a reversible neural network in the encoder respectively perform forward feature extraction on the carrier image and the ciphertext image to obtain a target carrier feature vector corresponding to the carrier image and a target ciphertext feature vector corresponding to the ciphertext image.
[0069] In one or more embodiments of the present application, the carrier image refers to image data serving as a carrier for hiding ciphertext text data. The ciphertext text data refers to text data containing ciphertext information (i.e., secret information). The ciphertext image refers to image data formed after the ciphertext text data is converted into a picture format. In order to further improve the robustness and reliability of the forward feature extraction performed by the encoder, the image size of the ciphertext image can be set to be the same as that of the carrier image.
[0070] It can be understood that the learnable wavelet transform network LWN (Deform-Attention Invertible Neural Network) is a learnable wavelet network with forward and reverse processes. When the forward process is executed, the model architecture of the learnable wavelet transform network can be used as the corresponding forward propagation network of the learnable wavelet transform network. When the reverse process is executed, the model architecture of the learnable wavelet transform network can be used as the corresponding reverse propagation network of the learnable wavelet transform network. The forward feature extraction in step 100 refers to the execution of the feature extraction process by the corresponding forward propagation network of the learnable wavelet transform network LWN. The learnable wavelet transform network LWN can adaptively embed a watermark in a high-frequency region that is insensitive to the watermark by the human eye, thereby improving the invisibility and noise resistance of the watermark. The Haar transform can be implemented based on a convolutional neural network CNN, which will be described in detail in the following embodiments.
[0071] In one or more embodiments of the present application, the Haar transform (Haar Wavelet) refers to the Haar wavelet transform. The Haar wavelet is a basic form of discrete wavelet transform DWT, which is characterized by simplicity, speed and ease of implementation. It helps to extract different features in the signal by decomposing the signal into high-frequency and low-frequency components, and is widely used in signal processing, image compression and multi-resolution analysis, especially for capturing edges or abrupt changes in images.
[0072] In addition, in order to further improve the effectiveness and reliability of the ciphertext cross-modal steganography, the current received ciphertext text data and the carrier image can be preprocessed before step 100. The watermark diffusion operation can be performed on the ciphertext text data to convert it from the original K-bit character to an image format with the same size as the carrier image, facilitating subsequent steganography operation.
[0073] In a specific example of pre-processing the carrier image, the carrier image can be scaled to a uniform size first. Thus, the picture size is standardized, facilitating subsequent operations; and the carrier image is then batched and divided, so as to effectively utilize the hardware performance of the server.
[0074] In a specific example of pre-processing the ciphertext text data, the ciphertext information is first copied into three parts, representing the R (red), G (green) and B (blue) three channels of the image, and then the length of the ciphertext information is expanded through a fully connected layer. After the fully connected layer, the ciphertext information becomes a string with a length of H*W (H and W represent the height and width of the image, respectively). Then, the expanded ciphertext information representing each channel is reshaped into an image with a size of (H, W, 1) (1 represents a single channel, and a color image usually has 3 channels) by using data reshaping. Finally, the three are spliced in the channel dimension to become a color image with a size of (H, W, 3) (hereinafter referred to as a ciphertext image), which is used for subsequent steganography operations.
[0075] Step 200: generating ciphertext cross-modal steganography result data corresponding to the ciphertext text data based on the target carrier feature vector and the target ciphertext feature vector for network transmission.
[0076] It can be understood that the ciphertext cross-modal steganography result data is an encrypted image corresponding to the ciphertext text data.
[0077] In step 200, the target carrier feature vector and the target ciphertext feature vector can be simply added in value to realize steganography. Specifically, since the target ciphertext feature vector has realized the generation mode with the lowest image influence, the numerical addition can be directly performed. It can be understood that the target carrier feature vector refers to the image feature data obtained by performing forward feature extraction on the carrier image; and the target ciphertext feature vector refers to the image feature data obtained by performing forward feature extraction on the ciphertext image.
[0078] As can be seen from the above description, the cross-modal steganography method provided by the embodiments of the present application can embed the ciphertext text data in the text modality into the carrier image in the image modality through the algorithm of deep learning, realize the steganography technology of multi-modal, and effectively improve the concealment and robustness of the ciphertext steganography by adopting the combination of the reversible neural network and the end-to-end network. Thus, not only the concealment of the hidden ciphertext text data can be improved, but also the resistance to various attacks can be enhanced, realizing more secure and reliable information transmission.
[0079] In order to further realize the cross-modal decoding of the steganography ciphertext and improve the accuracy and reliability of the steganography ciphertext decoding, in the cross-modal steganography method provided by the embodiments of the present application, referring to Figure 2The cross-modal steganography method further specifically comprises the following contents:
[0080] Step 300: receiving a noise image and extracting a noise carrier feature vector and a noise ciphertext feature vector corresponding to the noise image, wherein the noise image is formed after network transmission of ciphertext cross-modal steganography result data, and the ciphertext cross-modal steganography result data is generated based on a target carrier feature vector corresponding to a carrier image output by the encoder and a target ciphertext feature vector corresponding to a ciphertext image of ciphertext text data in advance.
[0081] Step 400: inputting the noise carrier feature vector and the noise ciphertext feature vector into a decoder corresponding to the encoder respectively, so that the learnable wavelet transform network and the reversible neural network in the decoder respectively perform inverse feature extraction on the noise carrier feature vector and the noise ciphertext feature vector to obtain a loss carrier image corresponding to the noise carrier feature vector and a restored ciphertext image corresponding to the noise ciphertext feature vector; wherein the loss carrier image corresponding to the noise carrier feature vector is different from the carrier image used to generate the noise image to which the noise carrier feature vector belongs.
[0082] In one or more embodiments of the present application, the restored ciphertext image corresponds to a ciphertext image of ciphertext text data, is the restored data corresponding to the ciphertext image, and can be referred to as a restored image. The ciphertext image can be referred to as an original image.
[0083] Step 500: performing information extraction on the ciphertext image to obtain ciphertext text data corresponding to the restored ciphertext image.
[0084] It can be understood that the steps 300 to 500 can be executed independently of the steps 100 to 200, or can be executed after the steps 200. Specifically, the steps 300 to 500 can be set according to actual needs.
[0085] For example, in an actual application scenario, a cross-modal steganography device can be set in different client devices A and B. The cross-modal steganography device in the client device A executes the steps 100 to 200, and the ciphertext cross-modal steganography result data obtained after the execution is transmitted in the network and received by the client device B. Then, the cross-modal steganography device in the client device B executes the steps 300 to 500.
[0086] For another example, in a test phase of the cross-modal steganography method, the same server C can be used to execute the steps 100 to 500. In order to improve the effectiveness of the test, noise simulation of network transmission of the ciphertext cross-modal steganography result data needs to be performed between the steps 200 and 300.
[0087] From the above description, the cross-modal steganography method provided by the embodiments of the present application can use a decoder composed of the same model parameters of the learnable wavelet transform network and the reversible neural network in the encoder to realize decoding of the stego ciphertext, can effectively improve the efficiency of cross-modal decoding of the stego ciphertext, and can effectively improve the accuracy and reliability of cross-modal decoding of the stego ciphertext.
[0088] In order to further enhance the imperceptibility and noise resistance of the ciphertext text data corresponding to the watermark, in a cross-modal steganography method provided by an embodiment of the present application, referring to Figure 3 , the encoder comprises a forward propagation network corresponding to the learnable wavelet transform network, a forward propagation network corresponding to the reversible neural network, and an inverse propagation network corresponding to the learnable wavelet transform network, which are connected in sequence.
[0089] Correspondingly, referring to Figure 4 , the step 100 of the cross-modal steganography method specifically comprises the following contents:
[0090] Step 110: inputting the carrier image and the stego ciphertext image corresponding to the ciphertext text data into a preset encoder respectively, so that the forward propagation network corresponding to the learnable wavelet transform network in the encoder respectively performs feature extraction on the carrier image and the stego ciphertext image in multiple different frequency domains, and outputs first multi-frequency domain feature vectors corresponding to the carrier image and the stego ciphertext image respectively; then the forward propagation network corresponding to the reversible neural network in the encoder respectively performs feature extraction on the first multi-frequency domain feature vectors corresponding to the carrier image and the stego ciphertext image based on a self-attention mechanism, and outputs second multi-frequency domain feature vectors corresponding to the carrier image and the stego ciphertext image respectively; and then the inverse propagation network corresponding to the learnable wavelet transform network in the encoder respectively performs feature extraction on the second multi-frequency domain feature vectors corresponding to the carrier image and the stego ciphertext image in the same frequency domain, and outputs a target carrier feature vector corresponding to the carrier image and a target stego feature vector corresponding to the stego ciphertext image.
[0091] In order to further realize cross-modal decoding of the stego ciphertext and improve the accuracy and reliability of decoding of the stego ciphertext, in a cross-modal steganography method provided by an embodiment of the present application, referring to Figure 5 , the decoder comprises: a forward propagation network corresponding to the learnable wavelet transform network, an inverse propagation network corresponding to the reversible neural network, and an inverse propagation network corresponding to the learnable wavelet transform network, which are connected in sequence.
[0092] Correspondingly, referring to Figure 6 , the step 300 of the cross-modal steganography method specifically comprises the following contents:
[0093] Step 310: input the noise carrier feature vector and the noise ciphertext feature vector into the decoder respectively, so that the forward propagation network of the learnable wavelet transform network in the decoder respectively performs down-sampling feature extraction on the noise carrier feature vector and the noise ciphertext feature vector for multiple different frequency domains, and outputs the third multi-frequency domain feature vector corresponding to the noise carrier feature vector and the noise ciphertext feature vector respectively; then the inverse propagation network corresponding to the reversible neural network in the encoder respectively performs feature restoration on the third multi-frequency domain feature vector corresponding to the noise carrier feature vector and the noise ciphertext feature vector based on the self-attention mechanism for target operation, and outputs the fourth multi-frequency domain feature vector corresponding to the noise carrier feature vector and the noise ciphertext feature vector respectively; and then the inverse propagation network corresponding to the learnable wavelet transform network in the encoder respectively performs up-sampling feature extraction on the fourth multi-frequency domain feature vector corresponding to the noise carrier feature vector and the noise ciphertext feature vector for the same frequency domain, and outputs the loss carrier image corresponding to the noise ciphertext feature vector and the restored ciphertext image corresponding to the noise ciphertext feature vector.
[0094] In order to further improve the robustness and steganographic concealment of the encoder and the decoder, in a cross-modal steganography method provided in an embodiment of the present application, referring to Figure 4 , the step 100 in the cross-modal steganography method further specifically comprises the following content:
[0095] Step 010: training the learnable wavelet transform network and the reversible neural network based on a preset target loss function and training data, so that the learnable wavelet transform network and the reversible neural network respectively constitute the encoder and the decoder; wherein the training data is composed of each data sample, and each data sample contains a carrier image, a ciphertext image corresponding to ciphertext text data, and a noise image with the same size; the ciphertext image corresponding to the ciphertext text data is generated by pre-processing the ciphertext text data by watermark diffusion; the noise image is generated by simulating noise attack on the ciphertext cross-modal steganography result data corresponding to the noise image based on a plurality of preset noise types;
[0096] The learnable wavelet transform network is composed of a plurality of convolutional neural networks for simulating Haar wavelet conversion corresponding to each different frequency domain.
[0097] Specifically, the learnable wavelet transform network LWN has both forward and inverse processes. The present application uses a convolutional neural network to simulate the process of Haar transform, realizes a wavelet transform network that can adjust parameters itself, and here the present application selects a typical architecture in wavelet transform, Haar transform. The specific network architecture of LWN is shown in Figure 7 A, V, H and D represent the feature information in the four frequency domains of approximation, horizontal, vertical and diagonal in turn. Here, the present application halves the H and W of the ciphertext image and the carrier image, but splices the four kinds of feature information together to obtain a feature vector with a size of (H / 2, W / 2, 12). The core component of the learnable wavelet transform network LWN is Figure 7 the convolutional neural network for simulating Haar wavelet conversion (denoted as DWT-CNN) in the middle, which is composed of a group of 2D convolution operations and aims to simulate the discrete wavelet transform. The difference between the learnable wavelet transform network LWN and the traditional discrete wavelet transform DWT lies in that the filter kernel of the traditional discrete wavelet transform DWT is fixed. In the learnable wavelet transform network LWN, the DWT is simulated by the CNN, and as the training process of watermark embedding and extraction, the convolution kernel is constantly learned and updated, so as to select more suitable high-frequency features to embed the watermark, thereby improving imperceptibility. When the input ciphertext image and the carrier image are input into the forward process of the DWT-CNN, it will be processed by four different convolutional neural networks DWT-CNN for simulating Haar wavelet conversion, and finally a feature vector composed of four different frequency domains is obtained. Among them, I m represents the image input into the learnable wavelet transform network, h is the height of the image input into the learnable wavelet transform network, and w is the width of the image input into the learnable wavelet transform network; I out represents the first multi-frequency domain feature vector output by the learnable wavelet transform network. K w [i] represents a set of convolution kernels (i.e. wave filter kernels) for simulating discrete wavelet transform (DWT). Kw[i] is learned and updated through a convolutional neural network (CNN) to dynamically optimize feature extraction in different frequency domains. This dynamic learning process is different from the traditional DWT where the filter is fixed.
[0098] The reversible neural network is composed of multiple layers of deformation convolution layers, and each layer of the deformation convolution layers is associated based on a target operation, and the target operation includes at least one of addition, subtraction, multiplication and division.
[0099] The deformable convolution layer is provided with a deformable attention dense block for performing a preset operation function, the deformable attention dense block comprising a self-attention block and a deformable dense block connected in sequence; wherein the deformable dense block comprises a deformable convolution layer and four convolution layers connected in sequence with the deformable convolution layer.
[0100] Specifically, the invertible neural network can adopt a deformable attention invertible neural network DA-INN (Deform-Attention Invertible Neural Network) which combines invertibility and deformable attention mechanism, and can further enhance the robustness of the network in dealing with complex noise attacks. The deformable attention invertible neural network DA-INN is a special invertible neural network architecture, which adopts deformable convolution to solve the low robustness of the invertible neural network when facing deformation attacks, and adopts self-attention mechanism to better extract feature information. The DA-INN is composed of 16 layers of the same network architecture, in order to realize the complete invertibility of the operation, the application needs to adopt addition, subtraction, multiplication and division operations which are easy to restore in the feature extraction process inside each layer. The forward transmission network corresponding to the network architecture of the DA-INN is shown in Figure 8(a), and the reverse transmission network corresponding to the network architecture of the DA-INN is shown in Figure 8(b).
[0101] The operation of each layer in the DA-INN is composed of three feature extraction operations F(), G() and H(), which realize the reversible feature extraction network by using addition, subtraction, multiplication and division. Since the forward operation is parameterized by a reversible function, the tensor can be accurately restored in the reverse process by using shared parameters. Among them, F() represents an operation for processing input data, which is mainly used to adjust some information of the input image or feature; G() represents a nonlinear transformation of the input image or watermark feature, which is used to calculate a weight value, and each point of the input is weighted by an exponential function, and the main purpose is to adjust the scale of a specific feature; H() represents a compensatory operation on the input image or watermark feature, which is an additional feature correction term, and its role is to assist feature recovery and enhance the expression of a specific frequency. In general, F() is used for preliminary processing, G() adjusts the scale distribution of the feature, and H() is responsible for compensating and correcting the feature, and the three work together to form the core functional module in the DA-INN, supporting the embedding and extraction process of the watermark.
[0102] In the present application, and respectively represent the watermark (i.e. the first multi-frequency domain feature vector corresponding to each of the ciphertext images) and the first multi-frequency domain feature vector corresponding to the carrier image input into the invertible neural network; and the corresponding and respectively represent the carrier tensor (i.e., the second multi-frequency domain feature vector corresponding to the carrier image) and the ciphertext information tensor (i.e., the second multi-frequency domain feature vector corresponding to the ciphertext image) after the current layer, respectively.
[0103] Each operation function F(), G() and H() in DA-INN is implemented by a dense block structure, which integrates deformable convolution and self-attention mechanism, called deformable attention dense block DA-dense block. See Figure 9 The deformable attention dense block DA-dense block of the present application consists of two parts: a self-attention block (SE-Block) and a deformable dense block. First, the feature map (i.e., the first multi-frequency domain feature vector) is input into the self-attention block to capture the importance of the channel. Then, the feature map output by the self-attention block is input into the deformable dense block, which contains five convolution layers. The first layer of the five convolution layers is a deformable convolution (DCN, Deformable Convolution), which is densely connected with four regular convolution layers, respectively, the first convolution layer (Convolution Layer-1), the second convolution layer (Convolution Layer-2), the third convolution layer (Convolution Layer-3) and the fourth convolution layer (Convolution Layer-4), and the feature channels are constantly expanding. The last convolution layer reduces the feature map with a large number of channels back to the original number of channels. Among them, the attention module in DA-dense block assigns different attention weights to different channels, effectively processing the feature map distinguished by LWN into high frequency and low frequency. In addition, the adaptive sampling position of the deformable dense block can enhance the robustness of the INN in processing spatial transformation.
[0104] Among them, the deformable convolution DCN (Deformable Convolutional Networks) is a method to enhance the ability of convolutional neural network (CNN). The traditional convolution operation may be limited by fixed sampling positions when processing objects with complex geometric deformation or different shapes, resulting in unsatisfactory feature extraction effect. DCN can effectively improve the performance of CNN in processing geometric deformation and complex scenes by introducing deformable convolution (Deformable Convolution) and deformable RoI pooling (Deformable RoIPooling). SE-Attention (Squeeze-and-Excitation Attention) is a lightweight and efficient attention mechanism introduced in Squeeze-and-Excitation Networks (SE-Net). SE-Attention can significantly improve the performance of convolutional neural network (CNN) by adaptively adjusting the weights of each channel, especially in image classification, object detection and semantic segmentation tasks.
[0105] In order to achieve the purpose of not restoring the original carrier image as much as possible to further prevent the cracking model, in the cross-modal steganography method provided in the embodiment of the application, the target loss function in the cross-modal steganography method specifically comprises the following contents:
[0106] The target loss function is composed of an encoding loss L en , an image recovery loss L re , a message decoding loss L de and a contrastive learning loss L con , and respective weights of the encoding loss L en , the image recovery loss L re , the message decoding loss L de and the contrastive learning loss L con ; the weights of the encoding loss L en , the image recovery loss L re and the contrastive learning loss L con are positive values, and the weight of the message decoding loss L de is a negative value;
[0107] The encoding loss L en is used to represent the respective pixel-level difference values and image loss values between the target ciphertext feature vectors extracted by the encoder and the ciphertext images, and between the target carrier feature vectors extracted by the encoder and the carrier images;
[0108] The image recovery loss L rea pixel-level difference value between the loss carrier image output by the decoder and the carrier image;
[0109] the message decoding loss L de a difference value between the restored ciphertext image output by the decoder and the ciphertext image input to the encoder, calculated based on mean square error;
[0110] the contrast learning loss L con a contrast learning loss value between the encoder and the decoder obtained based on a joint contrast learning unit; wherein the joint contrast learning unit comprises an encoding contrast learning unit and a decoding contrast learning unit connected in sequence.
[0111] The expression of the total loss function is as follows:
[0112] L total = λ1L en + λ2L re + λ3L de + λ4L con
[0113] wherein the encoding loss L en function consists of two parts: L2 loss for evaluating pixel-level differences in images, and image loss (Lpips\cite) measured from the perspective of human visual perception, and λ1 to λ4 are respectively the weights of the encoding loss L en , the image restoration loss L re , the message decoding loss L de and the contrast learning loss L con .
[0114] The joint contrast learning unit is used to perform the method of evaluating the concealment and robustness of steganography in the training stage as a whole. The joint contrast learning JCL (Joint Contrast Learning) method is designed based on SimSiam for the watermarking task. Contrast learning is a self-supervised and label-free training strategy, which is very suitable for the training process of blind watermarking methods. Traditional digital watermarking technology uses L2 loss or mean square error (MSE) loss for evaluation and training, which leads to evaluation indicators based only on low-dimensional features. The JCL method evaluates the encoding and decoding performance of the model in a high-dimensional feature space. The joint contrast learning JCL is a self-supervised contrast learning strategy that can optimize the encoding and decoding performance of the watermark in a high-dimensional feature space. SimSiam (Simple Siamese Network for Representation Learning) is a self-supervised learning method that focuses on representation learning of unlabeled data, aiming to simplify the training process of contrast learning. This method is used to solve the problem of how to effectively train a neural network to obtain useful feature representations without labels.
[0115] To further improve the robustness of the DA-INN, in a cross-modal steganography method provided in an embodiment of the present application, referring to Figure 6 , the step 500 in the cross-modal steganography method specifically includes the following content:
[0116] Step 510: inputting the restored ciphertext image into a decoding optimization module composed of a preset attention mechanism and a convolution layer, so that the decoding optimization module outputs the ciphertext text data corresponding to the restored ciphertext image.
[0117] Specifically, the end-to-end watermarking method uses a decoder to recover information, which makes it very robust to various attacks. However, the INN-based method only relies on the inherent reversible characteristics of the INN itself, resulting in lower robustness when facing complex noise. In order to improve the decoding ability of the INN-based method, the decoding optimization module DOM (Decoding Optimization Module) composed of an attention mechanism and a convolution layer designed by the present application is used to decode the ciphertext information from the restored ciphertext image. The decoding optimization module DOM combines the attention mechanism and the convolution layer, which can improve the decoding ability of the model under complex noise in the decoding process.
[0118] The operation DOM(I RS ) of the decoding optimization module DOM can be represented as follows:
[0119] DOM(I RS )=L M (SE(conv(I RS ))
[0120] where L M is a fully connected layer (FC layer), SE() represents an attention mechanism, and conv() is a convolutional layer. It is worth noting that the DOM established by the present application is applicable to all restored ciphertext images I RS That is, it not only effectively resists complex noise, but also improves the robustness against simple noise attacks. Therefore, the present application combines the advantages of end-to-end methods in resisting noise attacks and the advantages of accurate calculation of reversible neural networks, and through the DOM, the robustness of the model can be greatly improved.
[0121] To further illustrate the above embodiments, the present application also provides an application example of a cross-modal steganography method, specifically a robust image watermark steganography technology (Imperceptibility of the Watermark and Robustness against Noise attacks) based on cross-modal data fusion. Cross-modal data fusion refers to embedding ciphertext text information into a carrier image through a specific algorithm, without significantly changing the visual effect of the image, while ensuring that the embedded ciphertext text can withstand various environmental changes or malicious analysis and maintain its recoverability. Through the combination of image and text, two different modal data, not only can the concealment of information hiding be improved, but also the resistance to various attacks can be enhanced, realizing more secure and high-robustness information transmission. First, in order to enhance the imperceptibility of the watermark, the present application designs a learnable wavelet network (LWN) to adaptively embed the watermark in the high-frequency region that is insensitive to the human eye; secondly, the present application proposes a reversible neural network based on deformation attention (DA-INN), which has the advantage of calculating regression and combines deformation-attention mechanism to enhance the anti-noise ability of the model. In order to further improve the robustness of DA-INN, the present application introduces DOM to make up for its shortcomings in optimization. In addition, the present application establishes a joint contrast learning mechanism (JCL) based on Simsiam, which compares the similarity between the image and the watermark image, as well as the similarity between the watermark and the decoded watermark, thereby further improving the encoding and decoding ability of the model and further enhancing the imperceptibility and robustness of the watermark. Cross-modal refers to the combination of ciphertext with text information and steganography carrier with image through steganography technology. Under the premise that the ciphertext information can still be successfully extracted after facing various noise attacks with the steganography carrier, the concealment of the fusion of the two is improved to achieve the embedding of the ciphertext information without being able to see that the carrier has been written with information.
[0122] Referring to Figure 10 , the application example of the cross-modal steganography method specifically includes the following contents:
[0123] Step 1: Preprocessing the carrier image. The watermark diffusion operation on the ciphertext information converts it from the original K-bit character to an image format with the same size as the carrier image, facilitating subsequent steganography operations.
[0124] Specifically, the watermark diffusion operation on the ciphertext information converts it from the original K-bit character to an image format with the same size as the carrier image, facilitating subsequent steganography operations.
[0125] Among them, the preprocessing operation of the carrier image includes:
[0126] (1) Resize the carrier image to a uniform size. This standardizes the picture size, facilitating subsequent operations;
[0127] (2) Batch divide the carrier image. This effectively utilizes the hardware performance of the server.
[0128] Among them, the watermark diffusion operation on the ciphertext information is specifically:
[0129] First, copy the ciphertext information into three parts, representing the R, G, and B channels of the image, and then expand the length of the ciphertext information through a fully connected layer. After the fully connected layer, the ciphertext information becomes a string with a length of H*W (H and W represent the height and width of the image). Then, reshape the expanded ciphertext information representing each channel into an image with a size of (H, W, 1) (1 represents a single channel, and color images usually have 3 channels). Finally, concatenate the three in the channel dimension to become a color image with a size of (H, W, 3) (hereinafter referred to as the ciphertext image), which is used for subsequent steganography operations.
[0130] Step 2: Pass the ciphertext image and the carrier image through the forward process of the LWN network to obtain ciphertext image tensors and carrier image tensors with different frequency domain information.
[0131] Specifically, the purpose of this step is to downsample the image size to (H / 2, W / 2) but the feature channel becomes 12, where every 3 feature channels form a group, representing LL, HL, LH, and HH 4 image features.
[0132] Among them, the learnable wavelet transform network LWN has both forward and inverse processes. This application uses a convolutional neural network to simulate the process of Haar transform, realizing a wavelet transform network that can adjust parameters itself. Here, the application selects the typical architecture of wavelet transform, Haar transform. For the specific network architecture of LWN, please refer to Figure 7 .
[0133] A, V, H and D refer to the characteristic information in the four frequency domains of approximation, horizontal, vertical and diagonal in turn, where the application halves the H and W of the ciphertext image and the carrier image, but the four characteristic information is spliced together to obtain a feature vector with a size of (H / 2, W / 2, 12). The core component of the learnable wavelet transform network LWN is the convolutional neural network (which can be written as DWT-CNN) for simulating the Haar wavelet transform in Figure 7 The difference between the learnable wavelet transform network LWN and the traditional discrete wavelet transform DWT is that the filter kernel of the traditional discrete wavelet transform DWT is fixed. In the learnable wavelet transform network LWN, the DWT is simulated by the CNN, and as a training process for watermark embedding and extraction, the convolution kernel is constantly learned and updated, so as to select more suitable high-frequency features to embed the watermark, thereby improving imperceptibility. When the input ciphertext image and carrier image are input into the forward process of the DWT-CNN, it will be processed by four different convolutional neural networks DWT-CNN for simulating the Haar wavelet transform, and finally a feature vector composed of four different frequency domains is obtained.
[0134] Step 3: The ciphertext image tensor after LWN is input into the forward process of DA-INN together with the carrier image tensor for further feature extraction, so that the final ciphertext image tensor placed in the carrier image can be more hidden.
[0135] Referring to FIGS. 8(a) and 8(b), the deformation attention reversible neural network DA-INN is a special reversible neural network architecture, which adopts deformation convolution to solve the low robustness of reversible neural networks when facing deformation attacks, and adopts self-attention mechanism to better extract feature information. DA-INN is composed of 16 layers of the same network architecture. In order to realize the complete reversibility of operation, the application needs to adopt addition, subtraction, multiplication and division operations in the feature extraction process of each layer.
[0136] The operation of each layer in DA-INN is composed of three feature extraction operations F(), G(), H(), and the reversible feature extraction network is realized by using addition, subtraction, multiplication and division. Since the forward operation is parameterized by a reversible function, the tensor can be accurately restored in the reverse process by using shared parameters.
[0137] The application uses and to represent the watermark (i.e. the first multi-frequency domain feature vector corresponding to each ciphertext image) and the first multi-frequency domain feature vector corresponding to the carrier image input into the reversible neural network, respectively; and the corresponding and denote the carrier tensor (i.e., the second multi-frequency domain feature vector corresponding to the carrier image) and the ciphertext information tensor (i.e., the second multi-frequency domain feature vector corresponding to the ciphertext image) after the current layer, respectively.
[0138] The forward operation process of the DA-INN is as follows:
[0139]
[0140] where denotes the Hadamard product.
[0141] The inverse operation process of the DA-INN is as follows:
[0142]
[0143] Each operation function F(), G() and H() in the DA-INN is implemented by a dense block structure, which integrates deformable convolution and self-attention mechanism, called deformable attention dense block DA-dense block. See Figure 9 The deformable attention dense block DA-dense block of the present application consists of two parts: a self-attention block (SE-Block) and a deformable dense block. First, the feature map (i.e., the first multi-frequency domain feature vector) is input into the self-attention block to capture the importance of the channel. Then, the feature map output by the self-attention block is input into the deformable dense block, which contains five convolution layers. The first layer of the five convolution layers is a deformable convolution (DCN, Deformable Convolution), which is densely connected with four regular convolution layers, namely the first convolution layer (Convolution Layer-1), the second convolution layer (Convolution Layer-2), the third convolution layer (Convolution Layer-3) and the fourth convolution layer (Convolution Layer-4), and the feature channels are constantly expanded. The last convolution layer reduces the feature map with a large number of channels back to the original number of channels.
[0144] The attention module in the DA-dense block assigns different attention weights to different channels, effectively processing the feature maps distinguished by LWN into high frequency and low frequency. In addition, the adaptive sampling position of the deformable dense block enhances the robustness of the INN when processing spatial transformation.
[0145] Step 4: The reverse process of LWN network is inputted with the ciphertext image tensor and the carrier image tensor after the DA-INN further extracts the image features, and the (H / 2, W / 2, 12) tensor is changed back to the original image size of (H, W, 3) to obtain the final ciphertext image and carrier image.
[0146] The operation adopts the same parameters as the LWN forward process to realize lossless information transformation, and changes the 12-channel tensor back to the same 3-channel image as the input image.
[0147] Step 5: The ciphertext image and the carrier image are simply added to realize steganography. Then it is only a simple value addition operation of corresponding pixel points. Since the ciphertext image has realized the generation mode with the lowest impact on the image, direct numerical addition can be realized.
[0148] Step 6: The image is passed through the noise pool designed to simulate the noise (i.e. noise) attack that may occur in reality to improve the robustness of steganography technology when facing complex real-world environments.
[0149] In order to comprehensively simulate the noise situation in the real world, finally, the encryption image (i.e. ciphertext cross-modal steganography result data) of the no-noise attack (Identity) and the following 12 noise types are selected by the present application, such as set N ALL :
[0150] N All = {Identity, GaussianNoise, GaussianBlur, Saltpepper,
[0151] Resize, Saturation, Hue, Contrast, Brightness,
[0152] Rotation, Affine, Dropout, Cropout}
[0153] GaussianNoise represents adding random Gaussian distributed noise to simulate random errors in the signal; GaussianBlur represents blurring the image using a Gaussian kernel to reduce the sharpness of the image; Saltpepper represents adding random black and white noise points to simulate interference in the image transmission or storage process; Resize represents reducing or enlarging the image to test the robustness of the watermark to resolution changes; Saturation represents adjusting the color saturation of the image to simulate the impact of image color changes on the watermark; Hue represents modifying the color tone of the image to test the robustness of the watermark to color shifts; Contrast represents adjusting the light and dark difference of the image to test the impact of contrast changes on the watermark; Brightness represents adjusting the overall brightness of the image to test the impact of brightness changes on the watermark; Rotation represents rotating the image at an angle to simulate the robustness of the watermark to rotation interference; Affine represents a comprehensive attack including translation, rotation, scaling, and shearing linear transformations to test the performance of the watermark under complex geometric transformations; Dropout represents randomly discarding part of the pixels in the watermark image to simulate image data loss or occlusion; Cropout represents randomly cropping part of the area of the watermark image to simulate the situation where the image is partially intercepted.
[0154] The noise results formed after some noise types are simulated on the encrypted image without noise attack (Identity) are shown in Figure 11 Cropout and Dropout are not shown in Figure 11 because these attacks involve replacing part of the watermark image with the corresponding part of the original image, which has a great impact on the decoding of the watermark information, but it is difficult to detect visually.
[0155] Step 7: Copy the image subjected to noise attack as redundant information to input into the decoding process.
[0156] Since the forward process of DA-INN has two input information, and the forward and inverse processes share parameters, the application uses the copied image subjected to noise attack as the second input to input into the inverse process in the inverse process.
[0157] Step 8: Both are sequentially subjected to the forward process of LWN, the inverse process of DA-INN, and the inverse process of LWN to obtain the restored ciphertext image and carrier image.
[0158] The parameters of LWN and DA-INN in the inverse process are exactly the same as those in the forward process, and finally the ciphertext image and carrier image are restored.
[0159] Step 9: The reduced ciphertext image is passed through the DOM to extract the ciphertext information, obtaining the final ciphertext information. To make the model more difficult to crack, the reduced carrier image will be made as different as possible from the original carrier image.
[0160] The end-to-end watermarking method uses a decoder to recover the information, which makes it very robust to various attacks. However, the INN-based method only relies on the inherent reversible characteristics of the INN itself, resulting in lower robustness when facing complex noise. To improve the decoding ability of the INN-based method, the DOM module designed by the present application is used to decode the ciphertext information from the reduced ciphertext image. The operation of DOM can be represented as follows:
[0161] DOM(I RS )=L M (SE(conv(I RS ))
[0162] where L M is a fully connected layer (FC layer), SE() represents the attention mechanism, and conv() is the convolution layer. It is worth noting that the DOM established by the present application is applicable to all reduced ciphertext images I RS , that is, it not only effectively resists complex noise, but also improves the robustness to simple noise attacks. Therefore, the present application combines the advantages of the end-to-end method in resisting noise attacks and the advantages of the reversible neural network in accurate calculation, greatly improving the robustness of the model through DOM.
[0163] As for the reduced carrier image, due to the reversible nature of the reversible neural network, to prevent the model from being cracked by the thief and obtaining the model of the reverse process, so that the original image is restored, the present application sets the weight of the image restoration loss function (image recovery loss) to a negative number during the training process of the method, that is, the model will try to not restore the original image, and the final image will be far away from the original image.
[0164] In particular, the present application adopts a model optimization strategy combined with contrastive learning in the optimization strategy of model training, and designs a rich loss function.
[0165] The overall loss function is composed of four parts: the encoding loss L en , the image recovery loss L re , the message decoding loss L de , and the contrastive learning loss L con . These losses aim to ensure that the encrypted image is close to the carrier image, the ciphertext information matches the reduced information, and the reduced image is significantly different from the carrier image to prevent unauthorized reconstruction of the original image.
[0166] The expression of the total loss function is as follows:
[0167] L total = λ1L en + λ2L re + λ3L de + λ4L con
[0168] Coding loss L en The function consists of two parts: L2 loss, which is used to evaluate the pixel-level difference in the image, and Lpips\cite, which measures the image loss from the perspective of human visual perception. re The L2 loss between the carrier image and the restored image is used to evaluate the image reconstruction quality. de The function uses mean square error (MSE) to calculate the difference between the ciphertext information and the restored information. con The loss function uses the evaluation function in JCL during the encoding and decoding process.
[0169] The joint contrast learning unit is used to perform the method of evaluating the concealment and robustness of steganography in the training phase. It is a joint contrast learning JCL method designed for watermarking tasks based on SimSiam. Contrast learning is a self-supervised and label-free training strategy, which is very suitable for the training process of blind watermarking methods. Traditional digital watermarking technology uses L2 loss or mean square error (MSE) loss for evaluation and training, which results in evaluation indicators based only on low-dimensional features. The JCL method evaluates the encoding and decoding performance of the model in a high-dimensional feature space. As shown in Figure 12 , JCL contains two parts: encoding contrast learning (JCL: encoding contrast) and decoding contrast learning (JCL: decoding contrast). The structure of the joint contrast learning unit JCL can be expressed as follows:
[0170] Figure 12 As shown in the JCL, it contains two parts: encoding contrast learning and decoding process, which are used to compare the quality of encoding and decoding, respectively. In Figure 12 , Ip represents the image to be protected (The image to be protected
[0171] ), which is the input image that needs to be protected; Iw represents the watermarked image (Watermarked Image), which is the image data containing steganographic watermark information; P-Feature represents the features extracted from Ip; W-Feature represents the features extracted from Iw; Z Ip represents the projection result of P-Feature; Z Iwdenotes the projection result of W-Feature; P Ip denotes the prediction result of P-Feature; P Iw denotes the prediction result of W-Feature; NCS(Z Ip , P Iw ) denotes the negative cosine similarity between Z Ip and P Iw ; NCS(Z Iw , P Ip ) denotes the negative cosine similarity between Z Iw and P Ip ; E(Ip, Iw) denotes the joint contrastive loss of Ip and Iw, which evaluates the similarity of the projection and prediction results in combination with the negative cosine similarity; Projection-Head denotes a fully connected layer for mapping features to a shared feature space to enhance the representation capability; Joint training denotes joint training; Ms denotes the original watermark (Secret Message), i.e., the input watermark information without steganographic processing; M RS denotes the recovered watermark (Recovered Secret Message) after embedding and recovery, i.e., the watermark data extracted from the watermark image Iw; S-Feature denotes the feature extracted from Ms; RS-Feature denotes the feature extracted from M RS ; Z Ms denotes the projection result of S-Feature; Z MRS denotes the projection result of RS-Feature; P Ms denotes the prediction result of S-Feature; P MRS denotes the prediction result of RS-Feature; NCS(Ms, P Ms ) denotes the negative cosine similarity between Ms and P Ms ; NCS(Z Ms , P MRS ) denotes the negative cosine similarity between Z Ms and P MRS ; E(Ms, M RS ) denotes the joint contrastive loss of Ms and M RS , which evaluates the similarity of the watermark projection and prediction results in combination with the negative cosine similarity.
[0172] The structure of the joint contrastive learning unit can be as follows:
[0173] Z T = ρ(T feature )
[0174] PT = γ(Z T )
[0175] where T denotes the corresponding tensor in the model of the present application, T feature is the feature extracted from T. ρ() and γ() are the projection head and the prediction head, respectively. Z T and P T denote the projection and the prediction of T feature , respectively. The role of ρ() is to map the features to a shared feature space, thereby enhancing the representation ability of the model. The role of γ() is to obtain the predicted representation of the image using the projected features. To prevent the collapse solution, the projection Z T does not participate in the gradient update. The i-th tensor for contrast enhancement must be input into the network to obtain the corresponding Z i and P i . Z i is the projection result obtained by processing the feature T feature of the i-th tensor through the projection head ρ(); P i is the predicted representation generated by processing Zi through the prediction head γ().
[0176] It should be noted that the robustness to affine transformation noise can be achieved by using other targeted network design structures such as STN network, but compared with the DCN network, it does not solve the problem of CNN itself and cannot cope with more cases. In the optimization strategy, other contrast learning design frameworks can also be used, but the purpose is not changed compared with the present application, which is to use contrast learning to obtain a more comprehensive model evaluation. In the implementation of INN, a variety of different designs can be used, but the purpose is to achieve complete recoverability to ensure that the watermark can be effectively stego and restored in the encoding and decoding of the network.
[0177] That is, compared with existing image steganography technology, the technical implementation of the application example realizes cross-modal data fusion, combines text secret information with image carriers, and is significantly different from some existing technologies (such as IRWArt). And it realizes high concealment while maintaining robustness. The INN-based method such as (IRWArt and CIN) is difficult to achieve robustness when facing complex noise such as affine transformation, and the method of the application example realizes robustness when facing more powerful and more comprehensive noise by combining DCN and designing DOM. And it can still maintain high robustness even when the noise intensity exceeds twice the preset noise parameter. The application example adopts an optimization strategy based on contrast learning and designs JCL, which introduces the contrast learning method into the watermark steganography technology, effectively realizing a more effective watermark steganography model optimization strategy. As a more practical technology, the computing resources and coding and decoding speed required by digital watermark technology are also important indicators. By combining traditional watermark algorithms with deep learning watermark algorithms, the method of the application example is superior to other INN-based methods in terms of computing speed and required resources.
[0178] In summary, the present application can embed text modal ciphertext information into image modal carrier images through deep learning algorithms, realize multi-modal steganography technology, and because of the combination of reversible neural networks and end-to-end networks, the concealment and robustness of the model have reached an excellent level. Compared with other steganography technologies, existing deep learning-based watermark steganography technologies can be divided into end-to-end, based on generative adversarial networks, and based on INN. Among them, the latter two are to improve the concealment of the watermark, and the INN-based method has been proven to have excellent watermark concealment in recent years, but it is difficult to achieve sufficient robustness when facing complex noise. Therefore, the present application combines the excellent robustness performance of the end-to-end algorithm with the INN-based algorithm, designs a decoding optimization module (DOM), changes the ordinary convolution in INN to deformation convolution and adds a self-attention mechanism to improve the robustness of INN when facing multiple noises while preserving the information embedding concealment of INN. The present application combines traditional digital watermark steganography algorithms (based on Haar transform) with deep learning digital watermark algorithms, designs an LWN network, and combines the effective separation of high and low frequencies in traditional watermark steganography algorithms into deep learning watermark steganography algorithms, which not only improves the invisibility of steganography technology, but also effectively reduces the demand for computing performance of the model.
[0179] From the software level, the present application also provides a cross-modal steganography device for executing all or part of the cross-modal steganography method, see Figure 13 , which specifically includes the following content:
[0180] The encoding module 10 is configured to input the carrier image and the ciphertext image corresponding to the ciphertext text data into a preset encoder respectively, so that a learnable wavelet transform network and a reversible neural network in the encoder perform forward feature extraction on the carrier image and the ciphertext image respectively to obtain a target carrier feature vector corresponding to the carrier image and a target ciphertext feature vector corresponding to the ciphertext image.
[0181] The steganography module 20 is configured to generate ciphertext cross-modal steganography result data corresponding to the ciphertext text data based on the target carrier feature vector and the target ciphertext feature vector for network transmission.
[0182] The embodiments of the cross-modal steganography device provided in the present application can be specifically used to execute the processing flow of the embodiments of the cross-modal steganography method described above, and the functions thereof will not be repeated here. Please refer to the detailed description of the above-mentioned embodiments of the cross-modal steganography method.
[0183] The part of the cross-modal steganography device performing ciphertext cross-modal steganography can be completed in a server or in a client device. Specifically, it can be selected according to the processing capability of the client device and the limitation of the user's use scenario. The present application does not limit this. If all operations are completed in the client device, the client device can further include a processor for specific processing of ciphertext cross-modal steganography.
[0184] The above-mentioned client device can have a communication module (i.e., a communication unit) and can be communicatively connected with a remote server to realize data transmission with the server. The server can include a server of a task scheduling center side, and can also include a server of an intermediate platform in other implementation scenarios, such as a server of a third-party server platform communicatively connected with the server of the task scheduling center. The server can include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.
[0185] The server and the client device can communicate with each other using any suitable network protocol, including a network protocol that has not been developed at the filing date of the present application. The network protocol can include, for example, a TCP / IP protocol, a UDP / IP protocol, an HTTP protocol, an HTTPS protocol, etc. Of course, the network protocol can also include, for example, a RPC protocol (Remote Procedure Call Protocol) and a REST protocol (Representational State Transfer) used on the above-mentioned protocols.
[0186] As can be known from the above description, the cross-modal steganography device provided by the embodiments of the present application can embed the ciphertext text data in the text mode into the carrier image in the image mode through the algorithm of deep learning, implement the steganography technology in multiple modes, and effectively improve the concealment and robustness of the ciphertext steganography by adopting the combination of the reversible neural network and the end-to-end network, so as to not only improve the concealment of the hidden ciphertext text data, but also enhance the resistance to various attacks, and implement more secure and reliable information transmission.
[0187] On this basis, in order to further realize the cross-modal decoding of the steganography ciphertext and improve the accuracy and reliability of the decoding of the steganography ciphertext, in the cross-modal steganography device provided by the embodiments of the present application, referring to Figure 14 , the cross-modal steganography device further specifically includes the following contents:
[0188] The receiving module 30 is configured to receive a noise image and extract a noise carrier feature vector and a noise ciphertext feature vector corresponding to the noise image, wherein the noise image is formed after the ciphertext cross-modal steganography result data is transmitted through a network, and the ciphertext cross-modal steganography result data is generated based on a target carrier feature vector corresponding to a carrier image output by the encoder and a target ciphertext feature vector corresponding to a ciphertext image of ciphertext text data in advance.
[0189] The decoding module 40 is configured to input the noise carrier feature vector and the noise ciphertext feature vector into a decoder corresponding to the encoder respectively, so that the learnable wavelet transform network and the reversible neural network in the decoder respectively perform inverse feature extraction on the noise carrier feature vector and the noise ciphertext feature vector to obtain a loss carrier image corresponding to the noise carrier feature vector and a restored ciphertext image corresponding to the noise ciphertext feature vector; wherein the loss carrier image corresponding to the noise carrier feature vector is different from the carrier image used to generate the noise image to which the noise carrier feature vector belongs.
[0190] The recovery module 50 is configured to perform information extraction on the ciphertext image to obtain ciphertext text data corresponding to the restored ciphertext image.
[0191] The embodiments of the present application also provide an electronic device, which can include a processor, a memory, a receiver and a transmitter, the processor being configured to execute the cross-modal steganography method mentioned in the above embodiments, wherein the processor and the memory can be connected through a bus or other means to be connected through the bus. The receiver can be connected with the processor and the memory through wired or wireless means.
[0192] The processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, or a combination of the above.
[0193] The memory, as a non-transitory computer readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as program instructions / modules corresponding to the cross-modal steganography method in the embodiments of the present application. The processor executes various functions and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory, that is, implements the cross-modal steganography method in the above method embodiments.
[0194] The memory can include a program storage area and a data storage area. The program storage area can store an operating system and application programs required by at least one function; the data storage area can store data created by the processor and the like. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0195] The one or more modules are stored in the memory and, when executed by the processor, perform the cross-modal steganography method in the embodiments.
[0196] In some embodiments of the present application, a user equipment can include a processor, a memory and a transceiver unit which can include a receiver and a transmitter, the processor, the memory, the receiver and the transmitter can be connected through a bus system, the memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to transceive signals.
[0197] As an implementation manner, the functions of the receiver and the transmitter in the present application can be realized by a transceiving circuit or a dedicated chip for transceiving, and the processor can be realized by a dedicated processing chip, a processing circuit or a general-purpose chip.
[0198] As another implementation manner, the server provided by the embodiment of the present application can be implemented by using a general computer. That is, program codes for implementing the functions of the processor, the receiver and the transmitter are stored in the memory, and the general processor implements the functions of the processor, the receiver and the transmitter by executing the codes in the memory.
[0199] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the steps of the aforementioned cross-modal steganography method. The computer readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0200] The embodiment of the present application further provides a computer program product, which includes computer programs / instructions. The computer programs / instructions are executed by a processor to implement the steps of the aforementioned cross-modal steganography method.
[0201] Those skilled in the art should understand that the exemplary components, systems and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software or a combination thereof. The actual implementation depends on the specific application and design constraints imposed on the particular implementation. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are program or code segments used to perform the required tasks. The program or code segments can be stored in a machine readable medium or transmitted through a data signal carried in a carrier wave in a transmission medium or communication link.
[0202] It should be noted that the present application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of brevity, detailed descriptions of well-known methods are omitted. In the above embodiments, several specific steps are described and shown as examples. However, the method processes of the present application are not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications and additions, or change the order of the steps, after understanding the spirit of the present application.
[0203] In the present application, the features described and / or exemplified for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or in combination with or instead of the features of other embodiments
[0204] The above descriptions are only the preferred embodiments of the present application, and are not intended to limit the present application. The embodiments of the present application can be variously changed and modified by those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the scope of protection of the present application.
Claims
1. A cross-modal steganography method, characterized in that, include: The carrier image and the ciphertext image corresponding to the ciphertext data are respectively input into a preset encoder, so that the learnable wavelet transform network and the invertible neural network in the encoder perform forward feature extraction on the carrier image and the ciphertext image respectively to obtain the target carrier feature vector corresponding to the carrier image and the target ciphertext feature vector corresponding to the ciphertext image. The learnable wavelet transform network is composed of four convolutional neural networks that correspond one-to-one with the approximate frequency domain, horizontal frequency domain, vertical frequency domain and diagonal frequency domain respectively to simulate Haar wavelet transform, so that the learnable wavelet transform network is used to extract the characteristic information corresponding to the approximate frequency domain, horizontal frequency domain, vertical frequency domain and diagonal frequency domain of the carrier image and the ciphertext image respectively. After halving the height and width of the carrier image and the ciphertext image, the characteristic information corresponding to the approximate frequency domain, horizontal frequency domain, vertical frequency domain and diagonal frequency domain of the carrier image and the ciphertext image respectively are concatenated together to obtain the target carrier feature vector corresponding to the carrier image and the target ciphertext feature vector corresponding to the ciphertext image. Based on the target carrier feature vector and the target ciphertext feature vector, generate the ciphertext cross-modal steganography result data corresponding to the ciphertext text data for network transmission; Receive a noisy image and extract the noise carrier feature vector and noise ciphertext feature vector corresponding to the noisy image, wherein the noisy image is formed after the ciphertext cross-modal steganography result data is transmitted over the network, and the ciphertext cross-modal steganography result data is generated in advance based on the target carrier feature vector corresponding to a carrier image output by the encoder and the target ciphertext feature vector corresponding to the ciphertext image of a ciphertext data. The noise carrier feature vector and the noise ciphertext feature vector are respectively input into the decoder corresponding to the encoder, so that the learnable wavelet transform network and the invertible neural network in the decoder perform inverse feature extraction on the noise carrier feature vector and the noise ciphertext feature vector respectively, to obtain the lost carrier image corresponding to the noise carrier feature vector and the restored ciphertext image corresponding to the noise ciphertext feature vector; wherein, the lost carrier image corresponding to the noise carrier feature vector is different from the carrier image used to generate the noise image to which the noise carrier feature vector belongs; The restored ciphertext image is input into a preset decoding optimization module based on an attention mechanism and convolutional layers, so that the decoding optimization module outputs the ciphertext text data corresponding to the restored ciphertext image.
2. The cross-modal steganography method according to claim 1, characterized in that, The encoder includes a forward propagation network corresponding to a learnable wavelet transform network, a forward propagation network corresponding to a reversible neural network, and an inverse propagation network corresponding to the learnable wavelet transform network, which are connected in sequence. Correspondingly, the step of inputting the carrier image and the ciphertext image corresponding to the ciphertext data into a preset encoder, such that the learnable wavelet transform network and the invertible neural network in the encoder perform positive feature extraction on the carrier image and the ciphertext image respectively, to obtain the target carrier feature vector corresponding to the carrier image and the target ciphertext feature vector corresponding to the ciphertext image, includes: The carrier image and the ciphertext image corresponding to the ciphertext data are respectively input into a preset encoder. The forward propagation network corresponding to the learnable wavelet transform network in the encoder performs downsampling feature extraction on the carrier image and the ciphertext image for multiple different frequency domains, and outputs the first multi-frequency domain feature vector corresponding to the carrier image and the ciphertext image respectively. Then, the forward propagation network corresponding to the invertible neural network in the encoder performs target-oriented feature extraction on the first multi-frequency domain feature vectors corresponding to the carrier image and the ciphertext image respectively based on the self-attention mechanism, and outputs the second multi-frequency domain feature vectors corresponding to the carrier image and the ciphertext image respectively. Then, the inverse propagation network corresponding to the learnable wavelet transform network in the encoder performs upsampling feature extraction on the second multi-frequency domain feature vectors corresponding to the carrier image and the ciphertext image respectively for the same frequency domain, and outputs the target carrier feature vector corresponding to the carrier image and the target ciphertext feature vector corresponding to the ciphertext image.
3. The cross-modal steganography method according to claim 1, characterized in that, The decoder includes: a forward propagation network corresponding to the learnable wavelet transform network, an inverse propagation network corresponding to the reversible neural network, and an inverse propagation network corresponding to the learnable wavelet transform network, connected in sequence. Correspondingly, the step of inputting the noise carrier feature vector and the noise ciphertext feature vector into the decoder corresponding to the encoder, so that the learnable wavelet transform network and the invertible neural network in the decoder perform inverse feature extraction on the noise carrier feature vector and the noise ciphertext feature vector, respectively, to obtain the lost carrier image corresponding to the noise carrier feature vector and the restored ciphertext image corresponding to the noise ciphertext feature vector, includes: The noise carrier feature vector and the noise ciphertext feature vector are respectively input into the decoder, so that the forward propagation network of the learnable wavelet transform network in the decoder performs downsampling feature extraction on the noise carrier feature vector and the noise ciphertext feature vector for multiple different frequency domains, and outputs the third multi-frequency domain feature vector corresponding to each of the noise carrier feature vector and the noise ciphertext feature vector; then, the inverse propagation network corresponding to the reversible neural network in the encoder performs feature restoration based on the self-attention mechanism on the third multi-frequency domain feature vectors corresponding to each of the noise carrier feature vector and the noise ciphertext feature vector for target operation, and outputs the fourth multi-frequency domain feature vector corresponding to each of the noise carrier feature vector and the noise ciphertext feature vector; then, the inverse propagation network corresponding to the learnable wavelet transform network in the encoder performs upsampling feature extraction on the fourth multi-frequency domain feature vectors corresponding to each of the noise carrier feature vector and the noise ciphertext feature vector for the same frequency domain, and outputs the lost carrier image corresponding to the noise ciphertext feature vector and the restored ciphertext image corresponding to the noise ciphertext feature vector.
4. The cross-modal steganography method according to claim 1, characterized in that, Before inputting the carrier image and the ciphertext image corresponding to the ciphertext text data into the preset encoder respectively, the method further includes: The learnable wavelet transform network and the invertible neural network are trained based on a preset target loss function and training data, so that the learnable wavelet transform network and the invertible neural network constitute the encoder and the decoder, respectively. The training data consists of various data samples, and each data sample contains a carrier image of the same size, a ciphertext image corresponding to the ciphertext text data, and a noise image. The ciphertext image corresponding to the ciphertext text data is generated after pre-processing the ciphertext text data with watermark diffusion. The noise image is generated after pre-simulating a noise attack on the ciphertext cross-modal steganography result data of the ciphertext text data corresponding to the noise image based on multiple preset noise types. The reversible neural network is composed of multiple deformable convolutional layers, and the deformable convolutional layers are associated with each other based on a target operation, which includes at least one of addition, subtraction, multiplication and division. The deformable convolutional layer is provided with a deformable attention dense block for executing a preset operation function. The deformable attention dense block includes an adjoining self-attention block and a deformable dense block. The deformable dense block includes a deformable convolutional layer and four convolutional layers that are densely connected to the deformable convolutional layer in sequence.
5. The cross-modal steganography method according to claim 4, characterized in that, The target loss function consists of encoding loss, image restoration loss, message decoding loss, and contrastive learning loss, as well as the weights corresponding to each of the encoding loss, image restoration loss, message decoding loss, and contrastive learning loss; the weights of the encoding loss, image restoration loss, and contrastive learning loss are positive, and the weight of the message decoding loss is negative. The encoding loss is used to represent the pixel-level difference value and image loss value between the target ciphertext feature vector extracted by the encoder and the ciphertext image, and between the target carrier feature vector extracted by the encoder and the carrier image, respectively. The image restoration loss is used to represent the pixel-level difference between the loss carrier image output by the decoder and the carrier image; The message decoding loss is used to represent the difference between the restored ciphertext image output by the decoder and the ciphertext image input to the encoder, calculated based on the mean square error. The contrastive learning loss is used to represent the contrastive learning loss value between the encoder and the decoder obtained based on the joint contrastive learning unit; wherein, the joint contrastive learning unit includes an encoding contrastive learning unit and a decoding contrastive learning unit connected together.
6. A cross-modal steganography device, characterized in that, include: An encoding module is used to input the carrier image and the ciphertext image corresponding to the ciphertext data into a preset encoder, so that the learnable wavelet transform network and the invertible neural network in the encoder perform forward feature extraction on the carrier image and the ciphertext image respectively, to obtain the target carrier feature vector corresponding to the carrier image and the target ciphertext feature vector corresponding to the ciphertext image; the learnable wavelet transform network is composed of four convolutional neural networks that are one-to-one with the approximate frequency domain, horizontal frequency domain, vertical frequency domain and diagonal frequency domain respectively for simulating Haar wavelet transform, so that the learnable wavelet transform network is used to extract the characteristic information corresponding to the approximate frequency domain, horizontal frequency domain, vertical frequency domain and diagonal frequency domain of the carrier image and the ciphertext image respectively, and after halving the height and width of the carrier image and the ciphertext image, the characteristic information corresponding to the approximate frequency domain, horizontal frequency domain, vertical frequency domain and diagonal frequency domain of the carrier image and the ciphertext image respectively are concatenated together to obtain the target carrier feature vector corresponding to the carrier image and the target ciphertext feature vector corresponding to the ciphertext image; The steganography module is used to generate ciphertext cross-modal steganography result data corresponding to the ciphertext text data based on the target carrier feature vector and the target ciphertext feature vector for network transmission; The cross-modal steganography device is also used to perform the following: Receive a noisy image and extract the noise carrier feature vector and noise ciphertext feature vector corresponding to the noisy image, wherein the noisy image is formed after the ciphertext cross-modal steganography result data is transmitted over the network, and the ciphertext cross-modal steganography result data is generated in advance based on the target carrier feature vector corresponding to a carrier image output by the encoder and the target ciphertext feature vector corresponding to the ciphertext image of a ciphertext data. The noise carrier feature vector and the noise ciphertext feature vector are respectively input into the decoder corresponding to the encoder, so that the learnable wavelet transform network and the invertible neural network in the decoder perform inverse feature extraction on the noise carrier feature vector and the noise ciphertext feature vector respectively, to obtain the lost carrier image corresponding to the noise carrier feature vector and the restored ciphertext image corresponding to the noise ciphertext feature vector; wherein, the lost carrier image corresponding to the noise carrier feature vector is different from the carrier image used to generate the noise image to which the noise carrier feature vector belongs; The restored ciphertext image is input into a preset decoding optimization module based on an attention mechanism and convolutional layers, so that the decoding optimization module outputs the ciphertext text data corresponding to the restored ciphertext image.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the cross-modal steganography method as described in any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the cross-modal steganography method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Lithology image data digital watermark processing method and system
CN118283195A