An adaptive multi-modal generative hidden method and system

CN122223168BActive Publication Date: 2026-08-21JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610686481.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-08-21
Estimated Expiration
2046-05-19

AI Technical Summary

Technical Problem

然而,在高端安全应用中,同步传输文字指令与侦察影像的双模态需求对现有技术提出了严峻挑战

Benefits of technology

1、本发明基于预训练Stable Diffusion主干网络构建生成式隐藏通道,直接实现文本与图像秘密信息同步嵌入,降低部署成本与复杂度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223168B_ABST
    Figure CN122223168B_ABST
Patent Text Reader

Abstract

The application provides a self-adaptive multi-modal generation hiding method and system, which comprises the following steps: performing latent space feature coding and semantic extraction processing on original secret text and original secret image to obtain an enhanced text prompt condition; embedding multi-modal information in the process of reverse diffusion based on the enhanced text prompt condition and a carrier image to generate a joint feature vector fused with two modalities of text and image; suppressing multi-modal interaction interference and compensating generation error by using the joint feature vector fused with the two modalities of text and image and the enhanced text prompt condition to obtain a high-fidelity secret-containing image; and recovering and obtaining secret text and secret image from the high-fidelity secret-containing image. The application suppresses embedding distortion through a self-adaptive weight mapping, multi-scale difference analysis and noise compensation mechanism, the generated secret-containing image is natural and realistic, and the recovered secret image has excellent visual quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an adaptive multimodal generation and hiding method and system. Background Technology

[0002] Information hiding technology, also known as steganography, aims to seamlessly embed confidential data into non-sensitive data carriers. By ensuring that unauthorized observers cannot detect the hidden information, it enables secure communication where only authorized individuals can access the confidential data. This technology can be widely applied to various media such as images, audio, video, and text, preserving the perceptual characteristics of the original content while hiding the data. Therefore, it plays an irreplaceable role in information security fields such as privacy communication, covert transmission, and data protection. With the continuous evolution of steganography techniques, its application value in copyright protection and secure transmission is becoming increasingly prominent.

[0003] In terms of technical approaches, traditional carrier steganography primarily embeds information by fine-tuning existing media. While the least significant bit modification method is classic, it is easily detected by modern statistical analysis tools because it directly alters pixel values. To reduce this perceptual impact and maintain the statistical consistency of the carrier, researchers subsequently developed an adaptive method combining integrated lattice coding, polar coordinate steganography, and a distortion cost function, aiming to enhance security while reducing embedding distortion. With the advent of deep learning, the introduction of end-to-end learning frameworks and adversarial training strategies has further improved steganography performance. However, these carrier-modification-based methods still face a core bottleneck: any intervention in the mask image inevitably introduces visual or statistical distortion, leaving them vulnerable to detection by high-precision steganalysis tools.

[0004] To address this limitation, generative steganography has emerged as a paradigm innovation. Unlike modifying a cover image, generative steganography directly utilizes secret data as synthetic elements to construct entirely new steganographic images. This direct generation method eliminates dependence on a pre-set carrier, fundamentally improving the data's concealment and resistance to detection. For example, models based on generative adversarial networks can already integrate watermarks or encrypted instructions into the texture and structural features of images, achieving high-fidelity synthesis of single-modal data. However, in high-end security applications, the dual-modal requirement of simultaneously transmitting text instructions and reconnaissance images poses a significant challenge to existing technologies. Summary of the Invention

[0005] In view of the above, the main objective of this invention is to propose an adaptive multimodal generation and hiding method and system to solve the above-mentioned technical problems.

[0006] This invention proposes an adaptive multimodal generation and hiding method, which includes the following steps: Step 1: Perform latent spatial feature encoding and semantic extraction on the original secret text and original secret image to obtain enhanced text prompt conditions; Step 2: Based on the enhanced text prompt conditions and carrier image, embed multimodal information during the reverse diffusion process and generate a joint feature vector that integrates text and image modalities; Step 3: Use the joint feature vector that integrates text and image modalities and the enhanced text prompt conditions to suppress multimodal interaction interference and compensate for generation errors in order to obtain a high-fidelity dense image; Step 4: Recover and obtain the secret text and secret image from the high-fidelity encrypted image.

[0007] This invention also proposes an adaptive multimodal generative hiding system, the system comprising: The multimodal feature encoding and semantic extraction module is used for: The original secret text and original secret image are processed by latent spatial feature encoding and semantic extraction to obtain enhanced text prompt conditions; The adaptive weight mapping and dynamic capacity allocation module is used for: Based on enhanced text prompts and carrier images, multimodal information is embedded during the back diffusion process, and a joint feature vector that integrates text and image modalities is generated. The multi-scale interactive adjustment and dense image generation module is used for: By utilizing the joint feature vector that integrates text and image modalities and the enhanced text cue conditions, multimodal interaction interference is suppressed and generation errors are compensated to obtain high-fidelity dense images; The multi-scale feature fusion decoding and archive restoration module is used for: Recover secret text and secret image from high-fidelity encrypted image.

[0008] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention constructs generative hidden channels based on a pre-trained Stable Diffusion backbone network, directly achieving synchronous embedding of secret information in text and images, reducing deployment costs and complexity.

[0009] 2. This invention suppresses embedding distortion through adaptive weight mapping, multi-scale difference analysis and noise compensation mechanisms, resulting in natural and realistic dense images, and the restored secret images have excellent visual quality.

[0010] 3. This invention utilizes multi-scale feature fusion decoding technology and noise tracing and directional compensation strategies to resolve attribute conflicts in cross-modal embedding, ensuring the accuracy (ACC reaches 100%) and robustness of secret text and image recovery.

[0011] 4. This invention breaks away from the traditional carrier modification logic and autonomously generates dense images that conform to natural distribution, avoiding steganalysis from the bottom layer and greatly improving the security of privacy communication. Attached Figure Description

[0012] Figure 1 This is a flowchart of the adaptive multimodal generation and hiding method proposed in this invention; Figure 2 This is a schematic diagram showing the visual fidelity results of steganalysts generated under different cueing conditions using the adaptive multimodal generation and hiding method proposed in this invention. Figure 3 This is a schematic diagram of the adaptive multimodal generation and hiding system proposed in this invention. Detailed Implementation

[0013] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0014] These and other aspects of the embodiments of the present invention will become clear from the following description and accompanying drawings. In these descriptions and drawings, some specific embodiments of the present invention are specifically disclosed to illustrate some ways of implementing the principles of the embodiments of the present invention; however, it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0015] Please see Figure 1 and Figure 2 This invention proposes an adaptive multimodal generation and hiding method, which includes the following steps: Step 1: Perform latent spatial feature encoding and semantic extraction on the original secret text and original secret image to obtain enhanced text prompt conditions; In step 1, latent spatial feature encoding and semantic extraction are performed on the original secret text and the original secret image to obtain enhanced text hints. The specific steps are as follows: The original secret text is processed using configurable parameter segmentation encoding to obtain a binary sequence; A compact feature vector is extracted from the binary sequence using a semantic hashing algorithm, and the cosine similarity value between the compact feature vector and the original prompt word is calculated. When the cosine similarity value is less than the preset semantic threshold, the embedding strength coefficient of the original secret text information is re-encoded based on the cosine similarity value to obtain the text feature vector; otherwise, the compact feature vector is used as the text feature vector. The original secret image is normalized to obtain a normalized secret image. The encoder module of the variational autoencoder is used to transform the size-normalized secret image into a low-dimensional feature vector in the latent space, thus obtaining the image feature vector. The text feature vector and the image feature vector are fused by weight modulation to establish an enhancement constraint term for the diffusion model; The enhanced text prompt conditions are obtained by superimposing the enhanced constraint terms of the diffusion model with the original text prompt conditions; In the process of processing the original secret text into a binary sequence based on configurable parameter segmentation encoding, the following relationship exists: ; in, Represents a binary sequence. This represents the total number of segments in the original secret text message. Indicates the first Original secret text information, This indicates that the data will be converted to... Functions with binary bits; In the process of extracting compact feature vectors from binary sequences using semantic hashing algorithms and calculating the cosine similarity value between the compact feature vectors and the original prompt words, the following relationship exists: ; in, This represents the cosine similarity value between the compact feature vector and the original prompt word. Indicates the original prompt word, Represents a compact feature vector. Indicates L2 norm processing; In this step, semantic coherence is verified by calculating the cosine similarity between the compact feature vector and the original prompt word; When the cosine similarity value is less than a preset semantic threshold, the process of re-encoding the embedding strength coefficient of the original secret text information based on the cosine similarity value to obtain the text feature vector has the following relationship: ; in, This represents the embedding strength coefficient of the original secret text information; This represents the semantic threshold, with a default value of 0.6. This represents the strength coefficient of the re-encoded text information; In the process of using the encoder module of a variational autoencoder to transform the size-normalized secret image into a low-dimensional feature vector in the latent space, and obtaining the image feature vector, the following relationship exists: ; in, Represents the image feature vector. This represents the secret image after size normalization. This indicates that the encoding module of the variational autoencoder is used for processing. In the process of fusing text feature vectors and image feature vectors through weight modulation to establish the enhancement constraint term of the diffusion model, the following relationship exists: ; in, This represents the enhancement constraint term in the diffusion model. Represents text weight mapping, Represents image weight mapping, It represents the Hadamardi (or Hadama) stack; In the process of superimposing the enhancement constraints of the diffusion model with the original text prompt conditions to obtain the enhanced text prompt conditions, the following relationship exists: ; in, Indicates the original text prompt conditions. This indicates conditions for enhanced text suggestions; The enhanced text prompts are used to guide the diffusion model to generate steganalogs that are both natural and covert.

[0016] Step 2: Based on the enhanced text prompt conditions and carrier image, embed multimodal information during the reverse diffusion process and generate a joint feature vector that integrates text and image modalities; In step 2, based on the enhanced text prompt conditions and the carrier image, multimodal information is embedded during the back-diffusion process, and a joint feature vector that integrates text and image modalities is generated. The specific steps are as follows: The gradient magnitude of the carrier image is calculated using the Sobel operator; A comprehensive edge intensity map is constructed using gradient magnitude and the Canny algorithm. At each time step of the diffusion process, the regional redundancy of the integrated edge intensity map is calculated; The weight mapping is calculated by integrating the edge intensity map; By using the weight mapping and the preset capacity allocation ratio, the adjusted text information weight mapping is obtained through calculation; By using the text information weight mapping adjusted in the previous time step and the regional redundancy of the comprehensive edge intensity map, the text sub-weights are dynamically calculated. Constraints on the rate of change between adjacent time steps are constructed using text sub-weights; By using the weight mapping and the preset capacity allocation ratio, the adjusted image information weight mapping is calculated. Under the constraint of the rate of change between adjacent time steps, the adjusted image information weight mapping is used to perform element-level fine embedding processing in the latent space to obtain a joint feature vector that integrates text and image modalities. In the process of calculating the gradient magnitude of the carrier image using the Sobel operator, the following relationship exists: ; in, This represents the gradient magnitude of the carrier image. Represents a carrier image. Indicates the horizontal pixel direction. Indicates the vertical pixel direction. Indicates the sign of the partial derivative. Represents a tiny increment in the numerical values ​​of the carrier image. This represents the coordinate increment of the carrier image in the horizontal spatial dimension. This represents the coordinate increment of the carrier image in the vertical spatial dimension. and These represent the horizontal and vertical gradients of the image, respectively. In each time step of the diffusion process, during the calculation of the regional redundancy of the comprehensive edge intensity map, the following relationship exists: ; in, Indicates at time step Position coordinates during diffusion Regional redundancy at that location Indicates position coordinates Edge strength at the location, This represents the overall edge intensity map. This represents the maximum value of the edge intensity map; In the process of calculating the weight mapping through the comprehensive edge intensity map, the following relationship exists: ; in, Represents weight mapping, This indicates processing via an exponential function; This indicates the adjustment parameter, with a default value of 0.5; In the process of calculating the adjusted text information weight mapping by utilizing the weight mapping and the preset capacity allocation ratio, the following relationship exists: ; in, This represents the adjusted text information weight mapping. Indicates the capacity allocation ratio. Indicates the length of the binary sequence of text. This represents the length of the binary sequence of the secret image. Represents a binary sequence of text. Represents a secret image binary sequence; In the process of dynamically calculating the text sub-weights using the adjusted text information weight mapping from the previous time step and the regional redundancy of the comprehensive edge strength map, the following relationship exists: ; in, Represents the smoothing coefficient. This represents the redundancy adjustment parameter. Indicates time step Text sub-weights, Indicates time step Text sub-weights; In constructing constraints on the rate of change between adjacent time steps using text sub-weights, the following relationship exists: ; Wherein, 0.1 is the coefficient for the maximum rate of change of weights, to ensure smooth weight updates; In the process of using the adjusted image information weight mapping to perform element-wise refined embedding in the latent space to obtain a joint feature vector that integrates text and image modalities, the following relationship exists: ; in, This represents the adjusted image information weight mapping. This indicates that the adjusted image information weights are mapped to position coordinates. The weight value at that point, This represents the joint feature vector that combines text and image modalities in terms of position coordinates. eigenvalues ​​at that location Represents the original latent feature vector at position coordinates The eigenvalue at that location is 0.01, which is the basic embedding strength coefficient. The same principle applies to the embedding of textual information. The basic embedding strength coefficient is carefully designed to ensure reliable extraction of embedded information while minimizing interference with latent features, ultimately achieving accurate matching between bimodal information and image region characteristics.

[0017] It should be noted that in this step, in the relational expression contained in the process of obtaining the joint feature vector that integrates the two modalities of text and image, the symbol "+" corresponds to the embedding of information "1", and the symbol "-" corresponds to the embedding of information "0".

[0018] Furthermore, in this step, during the iterative process of backdiffusion, to achieve optimal adaptation between bimodal information and image region characteristics, a collaborative optimization mechanism between dynamic weight adjustment and edge feature extraction needs to be established. Specifically, the core of dynamic weight adjustment lies in constructing a closed-loop system that integrates redundancy assessment and adaptive update: firstly, a redundancy assessment model is established through region feature analysis; then, the weight parameters are adaptively adjusted based on the assessment results; and finally, a dynamic weight allocation strategy matching the image region characteristics is generated.

[0019] At each time step of the reverse diffusion process The system is based on latent features Information entropy analysis establishes a regional adaptive mechanism, which specifically includes three key steps: First, by quantifying latent features... First, spatial information entropy is used to construct a regional information capacity assessment matrix; second, this is combined with a comprehensive edge intensity map. Calculate the redundancy index for each pixel; finally, dynamically adjust the weight mapping based on this index. and The allocation strategy is based on the following: image edge regions (high edge intensity) have complex structures, high information entropy, and low redundancy, so the amount of information embedded needs to be reduced to avoid distortion; while smooth regions (low edge intensity) are more suitable for carrying secret information due to their high redundancy, thus achieving an optimized balance between information concealment and visual quality.

[0020] It should be noted that, in comparison, the Canny algorithm achieves more accurate edge localization through Gaussian smoothing and non-maximum suppression, effectively reducing noise and false detections. Subsequently, non-maximum suppression is applied to refine the edges, making them sharper and more continuous. Then, high and low dual threshold processing is used to distinguish between strong and weak edges, thereby generating an edge intensity map based on Canny. By integrating the detection responses of the Canny and Sobel operators—the former contributing accurate edge localization and the latter providing gradient intensity information in the horizontal and vertical directions—a comprehensive edge intensity representation is established. This fused map effectively captures fine-grained and structured edge information, providing a robust foundation for subsequent redundancy assessment and adaptive weight allocation. This comprehensive edge intensity map accurately reflects the edge structure of the image.

[0021] Step 3: Use the joint feature vector that integrates text and image modalities and the enhanced text prompt conditions to suppress multimodal interaction interference and compensate for generation errors in order to obtain a high-fidelity dense image; In step 3, the joint feature vector that integrates text and image modalities and the enhanced text cue conditions are used to suppress multimodal interaction interference and compensate for generation errors, so as to obtain a high-fidelity dense image. The specific steps are as follows: The Gaussian pyramid downsampling operator is used to capture features from the joint feature vector that combines text and image modalities and the baseline pure text feature vector to obtain the multi-scale feature interaction bias. In the process of feature capture of feature vectors by the Gaussian pyramid downsampling operator, an interaction interference index is constructed based on the multi-scale feature interaction bias. When the interaction interference index exceeds the preset interference threshold, the enhanced text prompt conditions are adaptively corrected using the multi-scale feature interaction deviation to obtain the corrected text prompt conditions; otherwise, the enhanced text prompt conditions are used as the corrected text prompt conditions. Controlled noise is injected during the noise prediction stage, and the added Gaussian noise is combined with the noise predicted by the model to obtain balanced regularized noise. Using the corrected text prompt conditions and balanced regularized noise, the joint feature vector that integrates text and image modalities is updated according to the reverse process of the diffusion model to obtain the joint feature vector that integrates text and image modalities in the previous time step. After completing all diffusion time steps, the final fused feature is obtained. The final fused features are decoded by the decoder of the variational autoencoder to obtain a high-fidelity dense image; In the process of using the Gaussian pyramid downsampling operator to capture features from the joint feature vector that integrates text and image modalities, as well as the baseline pure text feature vector, to obtain the multi-scale feature interaction bias, the following relationship exists: ; in, This represents the amount of cross-sectional feature bias. This represents a joint feature vector that integrates both text and image modalities. The plain text feature vector representing the baseline. Indicates the downsampling scale; This indicates that the Gaussian pyramid downsampling operator is used to capture features at the original, half, and quarter resolutions. The The acquisition method is as follows: under the same reverse diffusion environment and random seed, only the original text semantic guide words are input, and secret information embedding modulation is not performed. The standard latent feature response is calculated by the pre-trained generation model. In the process of constructing the interaction interference index based on multi-scale feature interaction bias, the following relationship exists: ; in, Indicates the interaction interference index. This represents the downsampled scale tensor component generated after embedding multimodal secret information. This represents the downsampled scale tensor component generated without information embedding and solely guided by the original text. This represents the reference scale tensor component generated after embedding multimodal secret information. This represents a reference scale component generated without information embedding and solely guided by the original text. Indicates the reference standard scale; In the process of adaptively correcting the enhanced text prompt conditions using multi-scale feature interaction bias to obtain the corrected text prompt conditions, the following relationship exists: ; in, This indicates the corrected text prompt conditions. Indicates the iteration correction number at the current time step. This represents the correction factor.

[0022] It should be noted that in this step, the preset interference threshold follows the principle of visual perception limit and semantic stability constraint. Specifically, by quantifying the structural similarity (SSIM) change of dense images under different embedding intensities through pre-experiments, the preset interference threshold is set within the critical threshold range that can maintain the semantic consistency of the generated image and not produce sensory artifacts (for example, the value range is [1.2, 1.5]). When the real-time calculated interference index exceeds this threshold, the system will forcibly suppress the drift in the potential space by reducing the step size factor, thereby ensuring the high stealth of the multimodal hiding task.

[0023] It should be noted that in this step, a controlled noise injection mechanism is introduced in the noise prediction stage. By adding Gaussian noise, the balance regularization in the multimodal fusion process is ensured.

[0024] In this step, the embedding of multimodal information is susceptible to feature interaction interference. This negative impact stems from the distribution differences in the semantic space and the heterogeneity of feature dimensions among different modalities. This interference usually manifests as distortion of embedded information and a decrease in algorithm stability. To address this issue, a multi-scale denoising mechanism is used to integrate the framework for dynamic monitoring and real-time correction, which can accurately suppress feature interaction interference.

[0025] Step 4: Recover and obtain the secret text and secret image from the high-fidelity encrypted image; In step 4, the secret text and secret image are recovered and obtained from the high-fidelity encrypted image. The specific steps are as follows: By utilizing the differences in high-fidelity dense images at different scales, a noise intensity sequence is obtained by constructing a noise accumulation curve; When the noise intensity sequence reaches the preset condition (showing an upward trend for three consecutive time steps and the current noise intensity sequence exceeds the preset threshold), the noise source scale is located using the noise intensity sequence. By utilizing the noise source scale, the joint feature vector that integrates text and image modalities is compensated to correct noise interference and obtain the compensated latent spatial features. The compensated latent spatial features are subjected to a 3×3 neighborhood weighted average local smoothing process to obtain smoothed latent spatial features. The smoothed latent spatial features are decoded through parallel decoding branches to obtain secret text and secret images; In the process of obtaining a noise intensity sequence by constructing a noise accumulation curve using the difference values ​​of high-fidelity dense images at different scales, the following relationship exists: ; in, Indicates time step Accumulated amount of diffused noise over time Indicates the downsampling scale The corresponding weighting coefficients, Indicates at time step Current sampling scale The corresponding difference value; In the process of locating the scale of noise source using noise intensity sequences, the following relationship exists: ; in, Indicates the scale of the noise source; This represents the marginal maxima operator, used to find and output the marginal maximum value within a predefined scale set {1.0, 0.5, 0.25}. The operator determines the specific scale at which the maximum value is obtained, and is used to accurately locate the dominant scale that contributes the most to the current noise offset in the multi-scale feature space. In the process of compensating the joint feature vector that integrates text and image modalities by utilizing the noise source scale to correct noise interference and obtain the compensated latent spatial features, the following relationship exists: ; in, Indicates the total number of time steps; This represents the compensated potential spatial characteristics, used in subsequent diffusion steps; Indicates time step The compensation coefficient at that time is designed to make the compensation strength decrease with the diffusion process (the initial step has low noise tolerance and strong compensation; the later step is close to the completion of generation and weak compensation), avoiding over-correction that leads to feature distortion. Indicated at the noise source scale Next step The dense feature components at time step t refer to the specific resolution level at that time step t. Latent tensor slices containing secret information features; Indicated at the noise source scale The baseline plain text feature components refer to the scale at the same time step. Reference feature slices guided only by the original text and without embedded information; To verify the compensation effect, a new calculation is performed after compensation. If the value is less than 0.9 (i.e., the cumulative noise reduction is more than 10%), the compensated characteristics are retained; otherwise, the value is increased appropriately. After increasing by 0.01, compensation is repeated until the effect requirements are met. This feedback mechanism ensures the effectiveness of targeted compensation and avoids wasting computational resources on ineffective corrections. In the process of performing a 3×3 neighborhood weighted average local smoothing process on the compensated latent spatial features to obtain the smoothed latent spatial features, the following relationship exists: ; in, Indicates position coordinates The smoothed latent feature space eigenvalues ​​at the given location. Indicates position coordinates The compensated latent spatial eigenvalues ​​at the location, Represented by position coordinates The set of 8 pixel locations excluding the center within a 3×3 neighborhood of the center. Represented by pixel position coordinates The feature values ​​of the neighboring pixels in the 3×3 neighborhood of the center point, excluding the center point.

[0026] Furthermore, the default value of the preset threshold mentioned in this step is 0.12.

[0027] Furthermore, in this step, the embedding of bimodal information during the diffusion process will cause slight perturbations to the latent spatial features. If not controlled in time, these perturbations will lead to the accumulation of noise prediction errors as the number of diffusion steps increases, ultimately affecting the visual fidelity and information extraction accuracy of the stegographic image. To address this, a noise source tracing and directional compensation mechanism is designed based on multiscale difference analysis (MSDA). Through a three-step process of real-time monitoring, precise positioning, and directional repair, dynamic suppression of noise accumulation is achieved.

[0028] Please see Figure 3 This invention also provides an adaptive multimodal generative hiding system, the system comprising: The multimodal feature encoding and semantic extraction module is used for: The original secret text and original secret image are processed by latent spatial feature encoding and semantic extraction to obtain enhanced text prompt conditions; The adaptive weight mapping and dynamic capacity allocation module is used for: Based on enhanced text prompts and carrier images, multimodal information is embedded during the back diffusion process, and a joint feature vector that integrates text and image modalities is generated. The multi-scale interactive adjustment and dense image generation module is used for: By utilizing the joint feature vector that integrates text and image modalities and the enhanced text cue conditions, multimodal interaction interference is suppressed and generation errors are compensated to obtain high-fidelity dense images; The multi-scale feature fusion decoding and archive restoration module is used for: Recover secret text and secret image from high-fidelity encrypted image.

[0029] For further details, please refer to Figure 2 The diagram illustrates the visual realism of steganalyst images generated under different cue conditions using the adaptive multimodal generation and hiding method proposed in this invention. The visual realism of steganalyst images generated by the method of this invention is shown. The same ciphertext information is used to generate steganalyst images with different text cue. These example images cover a variety of scenes such as kites, birds, animals and landscapes. The realistic details and rich content in the images fully demonstrate the excellent generation performance of this method. Natural Image Quality Evaluator (NIQE), Structural Similarity (SSIM), and Peak Signal-to-Noise Ratio (PSNR) are two commonly used metrics for evaluating the perceptual quality of images, while Extraction Accuracy (ACC) is one of the most critical quantitative metrics for measuring the performance of a steganalysis system. As shown in Table 1, the steganographic images generated by different methods are used when only a secret image is hidden, and the method of the present invention hides both a secret image and text information. Table 2 shows the steganographic images generated by different methods when only text information is hidden, and the method of the present invention hides both a secret image and text information. The NIQE, ACC, PSNR and SSIM values ​​are compared. The first column is the method and the first row is the evaluation index. As can be clearly seen from Tables 1 and 2, the method proposed in this invention is superior to other methods in all four indicators. This shows that the present invention can not only hide multimodal information, but also effectively reduce distortion in the information hiding process, further verifying its superior performance in steganography. Among them, the existing technologies in Table 1 are as follows: Existing technology 1 (Jing J, Deng X, Xu M, et al. Hinet: Deep image hiding by invertible network[C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2021: 4733-4742.), Existing technology 2 (Yu J, Zhang X, Xu Y, et al. Cross: Diffusion model makes controllable, robust secure image steganography[J]. Advances in Neural Information Processing Systems, 2023, 36: 80730-80743.), Existing technology 3 (Yang Y, Liu Z, Jia J, et al. DiffStega: towards universal training-free coverless imagesteganography with diffusion models[J]. arXiv preprint arXiv:2407.10459,2024.); The existing technologies listed in Table 2 are as follows: Existing technology 1 (Liu Q, Xiang X, Qin J, et al. A robust coverless steganography scheme using camouflage image[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2021, 32(6): 4038-4051.), Existing technology 2 (Guo B, Ping P, Xu F. Highly robust and diverse coverless imagesteganography against passive and active steganalysis[J]. IEEE Transactions on Dependable and Secure Computing, 2024, 22(3): 2771-2787.), and Existing technology 3 (Zhou Q, Wei P, Qian Z, et al. Improved generative steganography based on diffusion model[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2025.).

[0030] Table 1. Performance comparison of different hiding methods in terms of NIQE, PSNR, and SSIM.

[0031] Table 2. Performance comparison of different hiding methods in NIQE and ACC

[0032] In Tables 1 and 2, "↑" indicates that the larger the value, the better the indicator; and "↓" indicates that the smaller the value, the better the indicator.

[0033] In summary, the adaptive multimodal generative hiding method proposed in this invention can generate steganalytic images that simultaneously embed text and image information without additional training or fine-tuning. This method uses a Stable Diffusion backbone network to construct generative hidden channels and dynamically allocates cross-modal embedding capacity through an adaptive weight mapping mechanism. Furthermore, by integrating multi-scale difference analysis and noise control into a visual quality optimization strategy, error accumulation is effectively suppressed and the fidelity of secret recovery is improved. Experimental results show that the method successfully embeds and restores text and image information while maintaining image quality, verifying its effectiveness and feasibility in generative multimodal information hiding tasks.

[0034] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0035] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0036] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. An adaptive multimodal generation and hiding method, characterized in that, The method includes the following steps: Step 1: Perform latent spatial feature encoding and semantic extraction on the original secret text and original secret image to obtain enhanced text hints. The specific steps are as follows: The original secret text is processed using configurable parameter segmentation encoding to obtain a binary sequence; A compact feature vector is extracted from the binary sequence using a semantic hashing algorithm, and the cosine similarity value between the compact feature vector and the original prompt word is calculated. When the cosine similarity value is less than the preset semantic threshold, the embedding strength coefficient of the original secret text information is re-encoded based on the cosine similarity value to obtain the text feature vector; otherwise, the compact feature vector is used as the text feature vector. The original secret image is normalized to obtain a normalized secret image. The encoder module of the variational autoencoder is used to transform the size-normalized secret image into a low-dimensional feature vector in the latent space, thus obtaining the image feature vector. The text feature vector and the image feature vector are fused by weight modulation to establish an enhancement constraint term for the diffusion model; The enhanced text prompt conditions are obtained by superimposing the enhanced constraint terms of the diffusion model with the original text prompt conditions; Step 2: Based on the enhanced text prompts and the carrier image, multimodal information is embedded during the back-diffusion process to generate a joint feature vector that integrates both text and image modalities. The specific steps are as follows: The gradient magnitude of the carrier image is calculated using the Sobel operator; A comprehensive edge intensity map is constructed using gradient magnitude and the Canny algorithm. At each time step of the diffusion process, the regional redundancy of the integrated edge intensity map is calculated; The weight mapping is calculated by integrating the edge intensity map; By using the weight mapping and the preset capacity allocation ratio, the adjusted text information weight mapping is obtained through calculation; By using the text information weight mapping adjusted in the previous time step and the regional redundancy of the comprehensive edge intensity map, the text sub-weights are dynamically calculated. Constraints on the rate of change between adjacent time steps are constructed using text sub-weights; By using the weight mapping and the preset capacity allocation ratio, the adjusted image information weight mapping is calculated. Under the constraint of the rate of change between adjacent time steps, the adjusted image information weight mapping is used to perform element-level fine embedding processing in the latent space to obtain a joint feature vector that integrates text and image modalities. Step 3: Utilize the joint feature vector that integrates text and image modalities and the enhanced text prompt conditions to suppress multimodal interaction interference and compensate for generation errors, thereby obtaining a high-fidelity dense image. The specific steps are as follows: The Gaussian pyramid downsampling operator is used to capture features from the joint feature vector that combines text and image modalities and the baseline pure text feature vector to obtain the multi-scale feature interaction bias. In the process of feature capture of feature vectors by the Gaussian pyramid downsampling operator, an interaction interference index is constructed based on the multi-scale feature interaction bias. When the interaction interference index exceeds the preset interference threshold, the enhanced text prompt conditions are adaptively corrected using the multi-scale feature interaction deviation to obtain the corrected text prompt conditions. Conversely, the enhanced text prompt conditions will be used as the corrected text prompt conditions; Controlled noise is injected during the noise prediction stage, and the added Gaussian noise is combined with the noise predicted by the model to obtain balanced regularized noise. Using the corrected text prompt conditions and balanced regularized noise, the joint feature vector that integrates text and image modalities is updated according to the reverse process of the diffusion model to obtain the joint feature vector that integrates text and image modalities in the previous time step. After completing all diffusion time steps, the final fused feature is obtained. The final fused features are decoded by the decoder of the variational autoencoder to obtain a high-fidelity dense image; Step 4: Recover and obtain the secret text and secret image from the high-fidelity encrypted image.

2. The adaptive multimodal generation and hiding method according to claim 1, characterized in that, In the process of processing the original secret text into a binary sequence based on configurable parameter segmentation encoding, the following relationship exists: ; in, Represents a binary sequence. This represents the total number of segments in the original secret text message. Indicates the first Original secret text information, This indicates that the data will be converted to... Functions with binary bits; In the process of extracting compact feature vectors from binary sequences using semantic hashing algorithms and calculating the cosine similarity value between the compact feature vectors and the original prompt words, the following relationship exists: ; in, This represents the cosine similarity value between the compact feature vector and the original prompt word. Indicates the original prompt word, Represents a compact feature vector. Indicates L2 norm processing; When the cosine similarity value is less than a preset semantic threshold, the process of re-encoding the embedding strength coefficient of the original secret text information based on the cosine similarity value to obtain the text feature vector has the following relationship: ; in, This represents the embedding strength coefficient of the original secret text information. Indicates semantic threshold, This represents the strength coefficient of the re-encoded text information; In the process of using the encoder module of a variational autoencoder to transform the size-normalized secret image into a low-dimensional feature vector in the latent space, and obtaining the image feature vector, the following relationship exists: ; in, Represents the image feature vector. This represents the secret image after size normalization. This indicates that the encoding module of the variational autoencoder is used for processing. In the process of fusing text feature vectors and image feature vectors through weight modulation to establish the enhancement constraint term of the diffusion model, the following relationship exists: ; in, This represents the enhancement constraint term in the diffusion model. Represents text weight mapping, Represents image weight mapping, It represents the Hadamardi (or Hadama) stack; In the process of superimposing the enhancement constraints of the diffusion model with the original text prompt conditions to obtain the enhanced text prompt conditions, the following relationship exists: ; in, Indicates the original text prompt conditions. This indicates conditions for enhancing text prompts.

3. The adaptive multimodal generation and hiding method according to claim 1, characterized in that, In the process of calculating the gradient magnitude of the carrier image using the Sobel operator, the following relationship exists: ; in, This represents the gradient magnitude of the carrier image. Represents a carrier image. Indicates the horizontal pixel direction. Indicates the vertical pixel direction. Indicates the sign of the partial derivative. Represents a tiny increment in the numerical values ​​of the carrier image. This represents the coordinate increment of the carrier image in the horizontal spatial dimension. This represents the coordinate increment of the carrier image in the vertical spatial dimension; In each time step of the diffusion process, during the calculation of the regional redundancy of the comprehensive edge intensity map, the following relationship exists: ; in, Indicates at time step Position coordinates during diffusion Regional redundancy at that location Indicates position coordinates Edge strength at the location, This represents the overall edge intensity map. This represents the maximum value of the edge intensity map; In the process of calculating the weight mapping through the comprehensive edge intensity map, the following relationship exists: ; in, Represents weight mapping, This indicates processing via an exponential function. Indicates the adjustment parameter; In the process of calculating the adjusted text information weight mapping by utilizing the weight mapping and the preset capacity allocation ratio, the following relationship exists: ; in, This represents the adjusted text information weight mapping. Indicates the capacity allocation ratio. Indicates the length of the binary sequence of text. This represents the length of the binary sequence of the secret image. Represents a binary sequence of text. Represents a secret image binary sequence; In the process of dynamically calculating the text sub-weights using the adjusted text information weight mapping from the previous time step and the regional redundancy of the comprehensive edge strength map, the following relationship exists: ; in, Represents the smoothing coefficient. This represents the redundancy adjustment parameter. Indicates time step Text sub-weights, Indicates time step Text sub-weights; In constructing constraints on the rate of change between adjacent time steps using text sub-weights, the following relationship exists: ; In the process of using the adjusted image information weight mapping to perform element-wise refined embedding in the latent space to obtain a joint feature vector that integrates text and image modalities, the following relationship exists: ; in, This represents the adjusted image information weight mapping. This indicates that the adjusted image information weights are mapped to position coordinates. The weight value at that location, This represents the joint feature vector that combines text and image modalities in terms of position coordinates. eigenvalues ​​at that location Represents the original latent feature vector at position coordinates The eigenvalue at that location.

4. The adaptive multimodal generation and hiding method according to claim 1, characterized in that, In the process of using the Gaussian pyramid downsampling operator to capture features from the joint feature vector that integrates text and image modalities, as well as the baseline pure text feature vector, to obtain the multi-scale feature interaction bias, the following relationship exists: ; in, This represents the amount of cross-sectional feature bias. This represents a joint feature vector that integrates both text and image modalities. The plain text feature vector representing the baseline. Indicates the downsampling scale. This indicates processing using the Gaussian pyramid downsampling operator; In the process of constructing the interaction interference index based on multi-scale feature interaction bias, the following relationship exists: ; in, Indicates the interaction interference index. This represents the downsampled scale tensor component generated after embedding multimodal secret information. This represents the downsampled scale tensor component generated without information embedding and solely guided by the original text. This represents the reference scale tensor component generated after embedding multimodal secret information. This represents a reference scale component generated without information embedding and solely guided by the original text. Indicates the reference standard scale; In the process of adaptively correcting the enhanced text prompt conditions using multi-scale feature interaction bias to obtain the corrected text prompt conditions, the following relationship exists: ; in, This indicates the corrected text prompt conditions. Indicates the iteration correction number at the current time step. This represents the correction factor.

5. The adaptive multimodal generation and hiding method according to claim 4, characterized in that, In step 4, the secret text and secret image are recovered and obtained from the high-fidelity encrypted image. The specific steps are as follows: By utilizing the differences in high-fidelity dense images at different scales, a noise intensity sequence is obtained by constructing a noise accumulation curve; When the noise intensity sequence reaches a preset condition, the noise source scale is located using the noise intensity sequence. By utilizing the noise source scale, the joint feature vector that integrates text and image modalities is compensated to correct noise interference and obtain the compensated latent spatial features. The compensated latent spatial features are subjected to a 3×3 neighborhood weighted average local smoothing process to obtain smoothed latent spatial features. The smoothed latent spatial features are decoded through parallel decoding branches to obtain secret text and secret images.

6. The adaptive multimodal generation and hiding method according to claim 5, characterized in that, In the process of obtaining a noise intensity sequence by constructing a noise accumulation curve using the difference values ​​of high-fidelity dense images at different scales, the following relationship exists: ; in, Indicates time step Accumulated amount of diffused noise over time Indicates the downsampling scale The corresponding weighting coefficients, Indicates at time step Current sampling scale The corresponding difference value; In the process of locating the scale of noise source using noise intensity sequences, the following relationship exists: ; in, Indicates the scale of the noise source. Represents the marginal maxima operator; In the process of compensating the joint feature vector that integrates text and image modalities by utilizing the noise source scale to correct noise interference and obtain the compensated latent spatial features, the following relationship exists: ; in, Indicates the total number of time steps. This represents the latent spatial characteristics after compensation. Indicates time step The compensation coefficient at that time, Indicated at the noise source scale Next step The dense characteristic components of time, Indicated at the noise source scale The baseline plain text feature components; In the process of performing a 3×3 neighborhood weighted average local smoothing process on the compensated latent spatial features to obtain the smoothed latent spatial features, the following relationship exists: ; in, Indicates position coordinates The smoothed latent feature space eigenvalues ​​at the given location. Indicates position coordinates The compensated latent spatial eigenvalues ​​at the location, Represented by position coordinates The set of 8 pixel locations excluding the center within a 3×3 neighborhood of the center. Represented by pixel position coordinates The feature values ​​of the neighboring pixels in the 3×3 neighborhood of the center point, excluding the center point.

7. An adaptive multimodal generative hiding system, characterized in that, The system employs the adaptive multimodal generation and hiding method according to any one of claims 1 to 6, and the system comprises: The multimodal feature encoding and semantic extraction module is used for: The original secret text and original secret image are processed by latent spatial feature encoding and semantic extraction to obtain enhanced text prompt conditions; The adaptive weight mapping and dynamic capacity allocation module is used for: Based on enhanced text prompts and carrier images, multimodal information is embedded during the back diffusion process, and a joint feature vector that integrates text and image modalities is generated. The multi-scale interactive adjustment and dense image generation module is used for: By utilizing the joint feature vector that integrates text and image modalities and the enhanced text cue conditions, multimodal interaction interference is suppressed and generation errors are compensated to obtain high-fidelity dense images; The multi-scale feature fusion decoding and archive restoration module is used for: Recover secret text and secret image from high-fidelity encrypted image.

Citation Information

Patent Citations

  • Deep steganography image secret information blind extraction method based on self-supervised learning

    CN120543358A

  • Screen shooting resistant robust text image watermarking method fusing edge attention gating and multi-scale cavity convolution

    CN120707366A