Visual illusion hidden image creation method based on text graph large model
Through the potential diffusion model based on the literary and biographical image big model and combined with the hidden space phase migration mechanism, the visual structure of the reference image is fused into the scene described by the target text, solving the problem that the existing technology is difficult to create high-quality visual illusion hidden images, and realizing the deep fusion and flexible control of visual structure and semantic information.
Patent Information
- Application Number
- CN202510056722.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-14
AI Technical Summary
The prior art is difficult to create high-quality visual illusion hidden images, especially when using real and complex images, and cannot achieve effective visual illusion effects.
Based on the literary and biographical image big model, especially the latent diffusion model (LDM), by introducing cross-modal text guidance and constructing massive graphic and text training data, combined with the hidden spatial phase migration mechanism of the latent diffusion model, the visual structural clues of the reference image are fused into the scene described by the target text to generate high-quality visual illusion hidden art.
The visual illusion hidden image creation is realized that deeply integrates the structured information of the reference image with the target text semantic information. The generated images are both faithful to the target text description and reflect the visual structure of the reference image, and have high visual quality and flexibility.
Smart Images

Figure CN120088354A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence image creation, focusing on digital art and media creation empowered by artificial intelligence, and specifically relates to a method and system for creating an optical illusion hidden image based on a text-to-image large model. Background Art
[0002] With the continuous progress of multimodal representation learning and vision transformers, text-to-image large models have achieved great success. Driven by them, fields such as AI painting, digital art, and intelligent media have developed rapidly, and intelligent content creation communities have flourished unprecedentedly. Currently, generative vision large models represented by text-to-image large models have become the focus of industry research, and they have achieved remarkable results in the generation and creation of digital media such as images, videos, and 3D scenes. In addition to direct applications in content generation, a large amount of research work has explored derivative applications of vision large models, such as image and video editing, virtual digital humans, virtual fitting, layout design, artistic word design, image and video stylization, and so on.
[0003] Although vision large models have greatly improved the efficiency of content creation, their potential in the field of digital art creative design remains to be explored. Optical illusion hidden images are a form of digital art that integrates artistry, creativity, fun, and imagination. They cleverly dissolve or hide one image into the scene details of another image, creating an optical illusion effect where different image contents are observed from different angles. The creation of such images often relies on artificial creative inspiration and aesthetic wisdom, with relatively high creation difficulty and cost. Existing research on creating image optical illusions based on computation can be roughly divided into two categories: (1) generating camouflage images based on texture transfer; (2) generating optical illusions based on modeling human visual perception. The first category of methods aims to study texture transfer methods for natural scene images, transferring the texture style of the background image to specific targets of the foreground image block to achieve the "camouflage" of the foreground content in the environment. Due to overemphasizing visual concealment, the perceptibility of the hidden content in the generated camouflage images is relatively weak, making it difficult to be applied to digital art creative design. The second category of methods starts from the perspective of the human brain's visual perception mechanism and simulates simple optical illusion effects, such as geometric optical illusions, color optical illusions, motion optical illusions, etc., by explicitly modeling the human visual stimulus response process. Such methods can usually only produce simple images composed of simple geometric elements (such as straight lines, curves, color blocks, etc.), and cannot construct optical illusion effects based on real and complex images. Summary of the Invention
[0004] Aiming at the technical problems existing in the prior art, the purpose of the present invention is to provide a method for creating an optical illusion hidden image based on a text-to-image large model. Based on the text-to-image large model, the present invention realizes the intelligent creation of creative optical illusion hidden images and expands the application of the vision large model in the fields of digital art and digital creativity. In terms of technical implementation, the present invention relies on the text-to-image large model in the diffusion model paradigm. The denoising diffusion probabilistic model (DDPM), as the current mainstream generation model paradigm, maps an image to the Gaussian noise space through a forward diffusion process of gradually adding noise, and learns a reverse denoising process (sampling process) from Gaussian noise to the image to generate an image. The denoising diffusion implicit model (DDIM) improves the reverse sampling process of the diffusion model and significantly shortens the sampling time. On this basis, the text-to-image large model realizes text-guided open-domain image creation by increasing the number of parameters of the denoising network, introducing cross-modal text guidance, and constructing a large amount of paired text-image training data. Furthermore, the latent diffusion model (LDM) significantly reduces the training and inference calculation overhead of the text-to-image large model by migrating the training of the diffusion model from the high-dimensional pixel space to the low-dimensional feature space. Based on the LDM, the present invention proposes a plug-and-play latent diffusion model hidden space phase migration mechanism to realize the harmonious integration of the visual structure clues of any reference image into any scene described by the target text, and generate high-quality optical illusion hidden art.
[0005] The present invention relies on the large-scale text-to-image diffusion model to convert any input image into an optical illusion hidden image. This task can be described as inputting an arbitrary reference image x and an arbitrary target text y to the model to obtain the converted optical illusion hidden image. So that when observing at a close distance, the detailed scene content described by the text y can be observed, and when observing at a long distance, the structured visual clues of the input reference image x can be observed, that is, the harmonious integration of the structured information of x and the content semantic information of y is realized.
[0006] For the above task, the present invention realizes the generation and creation of an optical illusion hidden image from the perspective of a text-driven image translation method for the first time, and extends the powerful content creation ability of a vision large model to a new application field. The present invention is also the first method to model and control the spatial structure of an image generated by a large model by using the phase spectrum of the intermediate features of the diffusion model sampling trajectory, and realizes an application breakthrough by means of traditional signal processing technology to assist the cutting-edge AIGC technology. Based on a pre-trained large-scale text-to-image diffusion model, without additional model training and model fine-tuning, the present invention cleverly realizes the dissolution and penetration of the visual structure clues of a reference image into the scene image generated under the guidance of a target text by dynamically regulating the phase spectrum of the intermediate features of the diffusion model sampling trajectory. In addition, the present invention proposes an asynchronous phase migration method to control the intensity of the penetration of the reference image structure, so as to flexibly control the visual saliency of the hidden content of the generated optical illusion hidden image.
[0007] The technical solution adopted by the present invention is as follows:
[0008] An optical illusion hidden image creation method based on a text-to-image large model, comprising the following steps:
[0009] 1) Construct a pre-trained large-scale text-to-image latent diffusion model LDM, load the pre-trained model parameters, input a reference image x, and extract the feature z encoded by the VAE encoder of the LDM.
[0010] 2) Construct an inversion trajectory, apply DDIM inversion to the encoded feature z to project it step by step into the Gaussian noise space to obtain the inverted noise feature, and use an empty text as the guidance condition of the LDM for the inversion trajectory.
[0011] 3) Construct a reconstruction trajectory, apply DDIM sampling to the inverted Gaussian noise to gradually obtain the reconstructed feature of the initial encoded feature z The reconstruction trajectory also uses an empty text guidance condition to ensure the reconstructibility of the sampling result, that is
[0012] 4) Construct a generation trajectory of the same length as the reconstruction trajectory, which gradually denoises a randomly initialized Gaussian noise feature through a DDIM sampling process with the same number of steps, using the target prompt text y as the guidance condition of the LDM to obtain the finally generated latent space feature Then use the VAE decoder of the LDM to decode the sampled latent space feature to obtain the final generated image
[0013] 5) Embed the phase migration module proposed in the present invention between the reconstructed trajectory and the generated trajectory. This module transfers the phase spectrum of the intermediate denoised features in the reconstructed trajectory to the denoised features at the corresponding time step in the generated trajectory in a plug-and-play manner without training, thereby embedding the structural visual cues of the input reference image x into the sampling trajectory, and realizing the deep fusion of the structural information of the reference image x and the semantic information of the target text y in the feature latent space of the diffusion model, so that the generated image decoded not only conforms to the content semantics of y but also reflects the visual structure of x, thus obtaining the visual illusion art effect of harmoniously hiding and dissolving the reference image x into the scene described by any prompt text y.
[0014] As the technical core of the present invention, the specific technical solution of the dynamic phase migration mechanism is elaborated as follows:
[0015] First, for a pair of features at the same time step of the reconstructed trajectory and the generated trajectory (each sampling step is a time step), which may be called the reconstructed trajectory feature and the generated trajectory feature respectively, use the two-dimensional discrete Fourier transform 2D-FFT to transform them from the spatial domain to the frequency domain, and obtain the real part and the imaginary part of them in the frequency domain respectively. Then calculate the amplitude spectrum and the phase spectrum of the reconstructed trajectory feature and the generated trajectory feature from the real part and the imaginary part respectively. The amplitude spectrum is the square root of the sum of the squares of the real part and the imaginary part, and the phase spectrum is the arctangent of the ratio of the imaginary part to the real part. Further, replace the phase spectrum of the generated trajectory feature with the phase spectrum of the reconstructed trajectory feature. Specifically, use the phase spectrum of the reconstructed trajectory feature and the amplitude spectrum of the generated trajectory feature to synthesize a new frequency domain feature, and then use the two-dimensional inverse discrete Fourier transform 2D-IFFT to transform the newly synthesized frequency domain feature back to the spatial domain, and use the obtained spatial domain feature as the new generated trajectory feature at this time step. And so on, repeat this phase migration process at each time step of the generated trajectory to realize the deep fusion of the structural information of the reference image and the semantic information of the target text in the generated trajectory.
[0016] To ensure the high visual quality of the generated image, the text-guided generated trajectory is divided into two parts: the migration stage and the non-migration stage. The migration stage is the early part of the denoising process of the generated trajectory and has a decisive impact on the structural composition of the final generated image. The non-migration stage is the later part of the generated trajectory, and its main role is to improve the detail quality of the generated image. Therefore, the present invention only applies gradual phase migration in the early migration stage to inject the structured information of the reference image, while removing the phase migration module in the later non-migration stage to make full use of the native denoising process of the text-to-image large model to ensure the high visual quality of the generated image.
[0017] Considering that the structural information of the intermediate features of the reconstructed trajectory becomes increasingly prominent as the denoising process progresses, the present invention proposes an attenuation strategy for latent space phase migration to avoid overly strong structured embedding caused by directly migrating the feature phase spectrum in the later stage of the denoising process, thereby affecting the naturalness of the generated image. Specifically, direct phase replacement is adopted in the early part of the migration stage, while in the later part, a linear fusion of the phase spectra of the reconstructed trajectory features and the generated trajectory features is used to replace the phase spectrum of the generated trajectory features, and the fusion weight of the phase spectrum of the reconstructed trajectory features is gradually reduced to avoid overly strong structured penetration.
[0018] To better fuse the structural information of the reference image and the semantic information of the prompt text in the latent space of the diffusion model, the present invention proposes to append a self-correction module after each phase migration module. This module takes the generated trajectory features after phase migration as input and predicts the generated trajectory features themselves at the current time step again under the guidance of the prompt text, which to a certain extent promotes the fusion of the input reference image structural information and the target prompt text semantic information in the feature space, thereby improving the naturalness and visual quality of the generated optical illusion image.
[0019] To flexibly control the visual saliency of the hidden content (input image x) in the generated optical illusion image, the present invention further generalizes the above phase migration method from the input features at the same time step of the reconstructed trajectory and the generated trajectory to different time steps, and proposes asynchronous phase migration. Specifically, for the generated trajectory features at a certain time step, migrating the phase with the reconstructed trajectory features with weaker denoising at an earlier time step will weaken the intensity of the reference image structure penetration, and conversely, migrating the phase with the reconstructed trajectory features with higher denoising at a later time step will enhance the embedding of the structural information. Therefore, by changing the time step interval d (which can be negative) between the reconstructed trajectory features and the generated trajectory features, flexible control of the saliency of the hidden content in the generated optical illusion image can be achieved. The above asynchronous phase migration requires the reconstructed trajectory and the generated trajectory to be denoised in a non-parallel form, and the misalignment distance is adjustable, which brings difficulties to engineering implementation. The present invention proposes an efficient implementation method, which realizes non-parallel asynchronous phase migration in the form of parallel denoising of the reconstructed trajectory and the generated trajectory. Specifically, for the reconstructed trajectory features at each time step, first predict the approximate value of the reconstructed trajectory features after d-step denoising according to itself, and then migrate the phase spectrum of the predicted approximate value to the generated trajectory features corresponding to the original time step in the generated trajectory.
[0020] The advantages of the present invention are as follows:
[0021] 1. The present invention realizes the creation of an optical illusion hidden image for the first time from the perspective of a text-driven image translation method, expanding the application of the text-to-image large model in the fields of digital art and digital creativity. Compared with the existing advanced text-driven image translation methods, the present invention is the only technical method that realizes the deep fusion of the structural information of the reference image and the semantic content of the prompt text, enabling the generated image to be faithful to the description of the prompt text in terms of content scene details, while presenting the hidden content of the reference image from a farther perspective.
[0022] 2. The present invention relies on a large-scale text-to-image diffusion model, without model training, model fine-tuning, and any online optimization process, and has high inference efficiency, providing a plug-and-play method for the creation of optical illusion hidden art, greatly reducing the cost and overhead of manual design and creation.
[0023] 3. The present invention has high flexibility, allowing users to create optical illusion hidden images with arbitrary scenes and arbitrary hidden contents, and allowing users to flexibly control the visual saliency of the hidden contents in the generated images. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is the overall architecture diagram of this embodiment.
[0025] Figure 2 It is a schematic diagram of the detailed implementation of the core phase migration module of this embodiment.
[0026] Figure 3 It is a schematic diagram of the asynchronous phase migration method proposed in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] To make the technical solutions and implementation manners of the present invention more obvious and understandable, the model architecture and technical details of the present invention are described in detail below in conjunction with the accompanying drawings.
[0028] This embodiment discloses a method for creating an optical illusion hidden image based on a large text-to-image model, and the overall architecture is as Figure 1 shown, and the specific steps are described as follows:
[0029] Step 1: Construct and load a pre-trained large text-to-image model. The large text-to-image model in this embodiment uses the latent diffusion model (LDM) of StableDiffusion v1.5 version.
[0030] Step 2: Input the reference image x, extract features from it using the encoder E of the LDM, and obtain the initial feature z 0 = E(x). Based on z 0 construct the DDIM inversion trajectory of, and project z 0 step by step into the Gaussian noise space to obtain the corresponding noise representation where Tinv Denote the length of the inversion trajectory. In this embodiment, set T inv = 1000. At each time step of the inversion trajectory, use empty text as the guiding condition, and its expression is:
[0031]
[0032] where is the noise diffusion coefficient predefined by DDPM, ∈ θ is the noise estimation network in the LDM model, f θ (z t , t, v φ ) is the approximate z estimated from the z at the current step t . 0 .
[0033] Step 3: Based on the noise representation obtained by inversion Construct the T-step reconstruction trajectory by DDIM sampling , where The reconstruction trajectory uses the same empty text as in the inversion trajectory as the guiding condition to ensure that the final reconstructed feature is approximately consistent with the initial feature z of the reference image, that is 0 . In this embodiment, set T = 100. Specifically, the expression for each time step in the reconstruction trajectory is:
[0034]
[0035] Step 4: Construct the T-step generation trajectory parallel to the reconstruction trajectory by DDIM sampling , where is a noise signal randomly sampled from a standard Gaussian distribution, that is is the final denoising result of the generation trajectory, that is, the finally generated latent space feature. To make the scene content of the generated image determined by the prompt text, use the target prompt text v as the guiding condition in the generation trajectory. Further, to enhance the influence of the prompt text on the semantic content of the generated image, adopt the classic classifier-free guidance technique of the diffusion model, and use the linear combination of the noise estimation guided by the target prompt text and the noise estimation guided by the empty text as the final noise prediction for each time step in the generation trajectory:
[0036]
[0037] where ω is the guidance scale in the classifier-free guidance technique. In this embodiment, set it to ω = 7.5.
[0038] Step 5: Based on the constructed reconstruction trajectory and generation trajectory, embed a phase migration module between them to achieve the gradual fusion of the reference image structure information and the target text semantic information in the latent space of the generation trajectory. The implementation details of the phase migration module are as Figure 2 shown. Taking time step t as an example, the phase migration module (PTM) receives the reconstruction trajectory features at time step t and the generation trajectory features and migrates the phase spectrum of to to output the updated generation trajectory features Then, based on the updated generation trajectory features in step 4, use the denoising diffusion implicit model DDIM to gradually sample the noise signal randomly sampled from the standard Gaussian distribution to obtain the latent space features Specifically, apply the two-dimensional discrete Fourier transform to the reconstruction features and generation features respectively to transform them from the spatial domain to the frequency domain:
[0039]
[0040] where and are the real and imaginary parts of the reconstruction trajectory features respectively, and are the real and imaginary parts of the generation trajectory features respectively, and FFT is the two-dimensional fast Fourier transform. According to the results of the discrete Fourier transform, the amplitude spectrum and phase spectrum of the features can be further calculated:
[0041]
[0042] where and are the amplitude spectrum and phase spectrum of the reconstruction trajectory features respectively, and are the amplitude spectrum and phase spectrum of the generation trajectory features respectively. Combine the phase spectra of the reconstruction features and generation features with the fusion coefficient b t at time step t, combine the fused phase spectrum with the amplitude spectrum of the generation features to obtain the frequency domain features after phase migration, and then transform it back to the spatial domain through the two-dimensional Fourier inverse transform to obtain the updated generation trajectory features The above process is formulated as follows:
[0043]
[0044] where IFFT is the two-dimensional fast Fourier inverse transform. Finally, a self-correction module is added to the tail of the phase migration module, which is based on the generation features after phase migration Re - predict guided by the target prompt text itself, promoting the fusion of the embedded reference image structure information and the semantic information of the target text. The formula of the self - correction module is as follows:
[0045]
[0046] The setting of the fusion coefficient {b t} of the phase migration is elaborated below. To ensure the text fidelity and visual quality of the generated image, the generation trajectory is divided into an early migration stage and a late non - migration stage by the time step λT, and the phase migration is only applied to the early migration stage. In addition, as the denoising process progresses, the structural information of the reconstructed trajectory features is continuously enhanced. To avoid the reduction of the naturalness of the generated image caused by over - strong structural penetration, in this embodiment, the migration stage of the generation trajectory is further divided into two sub - stages: direct migration and decay migration by the time step τT. In the direct migration sub - stage, direct phase replacement is adopted, that is, set b t = 1; in the decay migration sub - stage, the phase fusion coefficient b t gradually decays to zero, continuously weakening the intensity of the phase migration. The above process can be expressed by the following formula:
[0047]
[0048] Among them, in this embodiment, λ = 0.4 and τ = 0.6 are set.
[0049] To achieve flexible control of the saliency of the hidden content in the generated visual illusion image, in this embodiment, the proposed phase migration module (PTM) is further extended from the migration at the same time step to the migration at different time steps, obtaining an asynchronous phase migration module (APTM). The schematic diagram of the method is as Figure 3 shown. Based on the PTM, the APTM first predicts the corresponding feature after d - step denoising according to the reconstructed trajectory feature at each time step, and then migrates the phase spectrum of to the generation trajectory feature at the current time step. When d is a positive number, the saliency of the hidden content in the generated visual illusion image (i.e., the generated image ) increases with the increase of d; when d is a negative number, the saliency of the hidden content weakens with the decrease of d. This process can be expressed by the following formula:
[0050]
[0051] This embodiment allows for flexible regulation of the saliency of the hidden visual content by adjusting the time - step interval d, and d ∈[-10, 10] is a more recommended value range.
[0052] Model inference:
[0053] Based on the text-to-image large model, the present invention proposes a plug-and-play phase migration technology, which realizes the creation of an optical illusion hidden image without any model training or model fine-tuning. The above steps 1 to 5 are both the model construction process and the model inference process.
[0054] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Those skilled in the art can modify or equivalently replace the technical solutions of the present invention without departing from the spirit and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.
Claims
1. A method for creating a visual illusion hidden image based on a large model of a Vincent graph, the steps comprising: 1) Use the pre-trained Wenshengtu model to extract the initial feature z0 of the reference image x; 2) Based on the constructed inversion trajectory, the denoising diffusion implicit model DDIM is used to invert the initial feature z0 to obtain the inverted Gaussian noise Each time step of the inversion trajectory uses an empty text As a guiding condition; 3) Based on the constructed reconstruction trajectory, the Gaussian noise after the inversion is denoised using the denoising diffusion implicit model DDIM Sampling is performed to obtain the reconstruction feature corresponding to the coding feature z0 Each time step of the reconstructed trajectory uses an empty text As a guiding condition; 4) Build The generated trajectory is based on the noise signal randomly sampled by the denoising diffusion implicit model DDIM obeying the standard Gaussian distribution. Sampling is performed to obtain latent space features The generated trajectory is equal in length to the reconstructed trajectory, and each time step of the generated trajectory is guided by the target prompt text v; 5) Using the Wensheng graph model to analyze the latent space features Decode to obtain a generated image that meets the semantics of the target prompt text v and has the visual structure of the reference image x 2. The method according to claim 1, characterized in that A phase migration module is embedded between the reconstruction trajectory and the generation trajectory to gradually fuse the structural information of the reference image x and the semantic information of the target prompt text v in the latent space of the generation trajectory.
3. The method according to claim 2, characterized in that In step 4), at the tth time step, the phase migration module receives the reconstructed trajectory feature of the tth time step and generate trajectory features The trajectory features will be rebuilt The phase spectrum is transferred to the generated trajectory characteristics In the above example, update the generated trajectory features Then based on the updated generated trajectory features The denoising diffusion implicit model DDIM is used to randomly sample the noise signal that follows the standard Gaussian distribution. Perform step-by-step sampling to obtain latent space features 4. The method according to claim 3, characterized in that Update the generated trajectory features The method is: to reconstruct the trajectory features and generate trajectory features Perform two-dimensional discrete Fourier transform respectively: in, and Reconstructed trajectory features are The real and imaginary parts of and Generate trajectory features The real and imaginary parts of the signal are obtained by using the two-dimensional fast Fourier transform (FFT). Then the amplitude spectrum and phase spectrum of the signal are calculated: in, and Reconstructed trajectory features are The amplitude spectrum and phase spectrum of and Generate trajectory features The amplitude spectrum and phase spectrum of the tth time step are then used as the fusion coefficient b t Reconstruction trajectory features Phase spectrum and generate trajectory features Phase spectrum Fusion is performed to obtain the fused phase spectrum The phase spectrum and the amplitude spectrum of the generated trajectory characteristics The frequency domain features after phase shift are obtained by combining them, and then the frequency domain features after phase shift are transformed back to the spatial domain through two-dimensional Fourier inverse transform to obtain the updated generated trajectory features 5. The method according to claim 4, characterized in that The generated trajectory is divided into a migration phase and a non-migration phase, and the migration phase is divided into a direct migration sub-phase and an attenuated migration sub-phase; that is, Wherein, T is the total number of time steps of the generated trajectory, λ and τ are proportional coefficients, and λ+τ=1.
6. The method according to claim 3, 4 or 5, characterized in that: A self-correction module is added at the end of the phase shift module; the self-correction module generates trajectory features based on the updated Re-predict the generated trajectory features based on the target prompt text v as the guide condition Promote the fusion of the image structure information of the reference image x and the semantic information of the target prompt text v.
7. The method according to claim 1, characterized in that An asynchronous phase migration module is embedded between the reconstructed trajectory and the generated trajectory to gradually fuse the structural information of the reference image x and the semantic information of the target prompt text v in the generated trajectory latent space; at the tth time step, the asynchronous phase migration module firstly generates a new image based on the reconstructed trajectory feature of the tth time step. Predict the corresponding features after d time steps of denoising Then The phase spectrum of the current t-th time step is transferred to the generated trajectory characteristics When d is a positive number, the generated image The hidden content saliency of increases with the increase of d; when d is a negative number, the generated image The saliency of the hidden content weakens as d decreases.
8. The method according to claim 1, characterized in that In step 4), the noise estimation guided by the target prompt text v and the empty text The linear combination of the guided noise estimates is used as the final noise prediction at each time step in the generated trajectory.
9. The method according to claim 1, characterized in that: The reconstructed trajectory is based on the inverted Gaussian noise Constructed The T-step reconstruction trajectory is constructed based on the initial feature z0. DDIM inversion trajectory.
10. The method according to claim 1, characterized in that The Vincent diagram model is a latent diffusion model (LDM).
Citation Information
Patent Citations
Generating method for digital disguise image
CN102779326A
Method and system for realizing visual invisibility
CN113701564A
Large model image-text generation method based on multi-modal information fusion
CN117271816A
Text-guided single-target object track mask video generation method and system
CN118612525A
Method and system for optimizing large model of text graph based on multi-objective optimization
CN118865394A