Digital culture creative content generation method based on artificial intelligence

By combining semantic analysis, latent variable optimization, random diffusion modeling, frequency modulation and partial differential equation control methods, the problems of semantic deviation and local control in text-driven image generation in the prior art are solved, and cultural creative content generation with high accuracy and controllability are achieved.

CN120070636AActive Publication Date: 2025-05-30SHENZHEN BAIXUN CULTURE MEDIA CO LTD

Patent Information

Application Number
CN202510141544.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

Existing text-driven image generation methods are prone to semantic deviations during the generation process, making it difficult to accurately capture the details in text descriptions, especially in cultural and creative applications, which lack precise control of local area textures, colors and light and shadow characteristics.

Method used

Using a digital cultural creative content generation method based on artificial intelligence, combining semantic analysis, latent variable optimization, random diffusion modeling, frequency modulation and partial differential equation control, we will intelligently generate from text description to high-quality image content. This method improves semantic consistency by constructing mathematical constraints, uses geometric optimization to ensure the stability of latent variables, and combines frequency domain modulation and Fourier transform to enhance the controllability of local textures.

Benefits of technology

It effectively improves the mapping accuracy of text semantics to visual content, enhances the model's adaptability to complex text structures, realizes precise control of local details, meets the needs of high-end artistic creation, and presents subtle but reasonable changes between different instances, improving the richness of cultural and creative content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070636A_ABST
    Figure CN120070636A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and particularly relates to a digital culture creative content generation method based on artificial intelligence. The method comprises the following steps of: 1, calculating global text embedding; step 2, obtaining a potential variable after manifold transformation; 3, using the transformed potential variable as an initial state, constructing a potential energy function including text semantic constraint and a global regularization item in a potential space, and realizing diffusion update of the potential variable through a discretized stochastic differential equation; generating a local texture map by adopting an inverse Fourier transform method; and 4, using the generated local texture map as a local driving factor, constructing a partial differential equation based on a diffusion and reaction mechanism, and performing spatio-temporal evolution on image brightness distribution by the partial differential equation under a preset initial condition to realize self-adaptive synthesis of image structure layout and generate image content. According to the invention, intelligent generation from text description to high-quality image content is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method for generating digital cultural and creative content based on artificial intelligence. Background Art

[0002] With the continuous development of artificial intelligence, computer vision, and generative modeling technologies, the method of automatically generating digital cultural and creative content based on text descriptions has become one of the important directions in the cultural and creative industry. Currently, many digital art, film and television special effects, game design, and cultural heritage restoration tasks rely on artificial intelligence technologies to automatically generate high-quality visual content. In particular, text-driven content generation technologies enable creators to generate images that conform to specific artistic styles or cultural themes through natural language descriptions. This method not only improves the production efficiency of creative content but also reduces the dependence on professional graphic designers, providing new possibilities for the automation and intelligent development of the cultural and creative industry.

[0003] Currently, most text-driven image generation methods are based on GANs, VAEs, or diffusion models. The core idea of these methods is to learn large-scale image-text paired data so that the model can find the mapping relationship between text descriptions and image content in a high-dimensional latent space. For example, StyleGAN performs excellently in high-resolution image generation, while CLIP-based generative models such as DALL·E can better understand text semantics and convert them into corresponding visual content. However, these methods generally have the following problems: Although methods such as CLIP enhance the correlation between text and image through cross-modal contrast learning, semantic deviation may still occur during the generation process. For example, the text description input by the user may contain multiple elements (such as "an ancient castle against the sunset background"), but the generative model may not be able to accurately capture all the details and tend to generate an approximate scene rather than content that exactly matches the text description. Traditional GAN or diffusion models often rely on an end-to-end training process, and it is difficult to precisely adjust features such as texture, color, and light and shadow in local areas when generating images. Especially in cultural and creative applications such as digital painting and art style transfer, creators usually hope to control the details of specific areas while maintaining the overall style consistency. However, most existing methods lack controllability for local features, resulting in the generated images being difficult to meet the requirements of high-end artistic creation in terms of fineness. Summary of the Invention

[0004] The main objective of the present invention is to provide a method for generating digital cultural and creative content based on artificial intelligence. By combining semantic parsing, latent variable optimization, stochastic diffusion modeling, frequency modulation, and partial differential equation control, it realizes the intelligent generation of high-quality image content from text descriptions. This method enhances semantic consistency by constructing mathematical constraints, ensures the stability of latent variables through geometric optimization, and combines frequency domain modulation and Fourier transform to enhance the controllability of local textures. In addition, a diffusion-reaction mechanism is introduced to make the generated content exhibit richer variations between different instances while maintaining the stability of the artistic style.

[0005] To solve the above technical problems, the present invention provides a method for generating digital cultural and creative content based on artificial intelligence, and the method includes:

[0006] Step 1: Divide the input text description into several semantic segments according to predefined syntactic and semantic rules. For each semantic segment, calculate the local semantic score using predefined word vector mapping and activation functions, and calculate the global text embedding using the local semantic scores of each semantic segment and a preset weight.

[0007] Step 2: Add the global text embedding to the noise generated using the standard normal distribution to construct an initial latent variable. Consider the initial latent variable as being located on an implicit manifold, construct a local loss function with the global text embedding as a reference, calculate the deviation between the initial latent variable and the text embedding using a preset Riemannian metric, and use the Riemannian gradient descent method to perform multiple iterative updates on the initial latent variable to obtain the latent variable after manifold transformation.

[0008] Step 3: Use the transformed latent variable as the initial state, construct a potential energy function including text semantic constraints and a global regularization term in the latent space, and realize the diffusion update of the latent variable through a discretized stochastic differential equation. Use the latent variable after diffusion update as the input, determine the texture amplitude, center frequency, and frequency diffusion parameters according to a preset frequency modulation function, and use the inverse Fourier transform method to generate a local texture map.

[0009] Step 4: Use the generated local texture map as a local driving factor, construct a partial differential equation based on the diffusion and reaction mechanism, and perform spatio-temporal evolution on the image brightness distribution under preset initial conditions to realize the adaptive synthesis of the image structure layout and generate image content.

[0010] Further, in Step 1, let the input text description be T = {w 1 , w 2 ,..., w j ,..., w M}; where w jDenote the j-th word, where j is an integer subscript index with a value range from 1 to M, and M is the total number of words; predefined word vector mapping Map the word w j to a real vector; Denote the vocabulary, which represents the set of all possible words that can appear; is the dimension of the word vector, indicating that the word w j has feature values in the word vector space; map each word to the vector space to obtain the word vectors φ(w j ); using predefined syntactic and semantic rules, decompose the text into N semantic fragments, and the i-th semantic fragment is is the j-th word in the i-th semantic fragment; the word index n i,j represents the position of the word in the entire text; L i is the number of words in the i-th semantic fragment; N is the total number of semantic fragments after the text is partitioned.

[0011] Furthermore, in step 1, define a semantic scoring function for each semantic fragment:

[0012]

[0013] where ||·|| 1 represents the L 1 norm operation; ψ(T i ) is the local semantic score; and use the Sigmoid activation function to construct the fragment semantic vector:

[0014] s i = σ(ψ(T i ) - θ)v 0 ;

[0015] where s i is the fragment semantic vector of the i-th semantic fragment; σ(·) is the Sigmoid activation function; θ is the activation threshold, which is a preset scalar used to translate the output of ψ(T i ) to control the response interval of the Sigmoid activation function; v 0 is the basic semantic direction vector, which is a preset vector in ; d s is the dimension of the semantic vector, which determines the size of the feature space of the global text embedding; the global text embedding is obtained by weighted averaging of the fragment semantic vectors:

[0016]

[0017] where represents the global text embedding; ||·|| 2Denotes the L2 norm.

[0018] Further, in step 2, the initial latent variable is constructed through the following process: for k = 1,..., d s , define:

[0019]

[0020] U k , V k ~Uniform(0, 1);

[0021] where U k and V k are random variables independently sampled from the uniform distribution Uniform(0, 1) used for generating the k-th noise component δ k ; all the noise components form a noise vector denotes the d s -dimensional identity matrix to ensure no correlation between the dimensions of the noise; the initial latent variable ξ 0 is constructed using the noise vector:

[0022] ξ 0 = e + λδ;

[0023] where λ adjusts the noise amplitude and its value range is from 0.1 to 0.5.

[0024] Further, in step 2, assume the initial latent variable ξ 0 is distributed on the implicit manifold , and its local geometry is described by a metric constructed based on a global text embedding, and the local loss function is defined as:

[0025] Further, in step 2, the deviation between the latent variable and the text embedding is calculated through a preset Riemannian metric, and the process of performing multiple-step iterative updates on the latent variable using the Riemannian gradient descent method to obtain the latent variable after manifold transformation includes:

[0026] The preset Riemannian metric is g:

[0027]

[0028] where ξm is the m-th component of ξ 0 ; ξl is the l-th component of ξ 0 ; em is the m-th component of e m ; el is the l-th component of e l ; glm is the element in the m-th row and l-th column of the Riemannian metric g; ml Update the latent variables using the Riemannian gradient descent through the following formula:

[0029]

[0030] where g -1 represents the inverse matrix of the Riemannian metric; the finally obtained latent variables after the manifold transformation are T 1 is the total number of iterations; is the gradient operator; is the local loss function of the latent variable ξ t-1 obtained in the (t - 1)-th iteration; ξ t is the latent variable obtained in the t-th iteration.

[0031] Furthermore, in step 3, based on the obtained introduce a diffusion process to further balance the text semantic constraints and generation exploration, and construct the following potential energy function with a global regularization term

[0032] where δ 1 > 0, which is a preset value and provides a global contraction effect; introduce the stochastic differential equation of the diffusion process:

[0033]

[0034] where D > 0, which is a preset diffusion coefficient, and W(t) is a standard Wiener process; let, define the discretization step Δt, and the iteration formula is:

[0035]

[0036] where After another T 1 iterations, finally obtain the latent variables after diffusion update

[0037] Furthermore, in step 3, the formula of the frequency modulation function is:

[0038]

[0039] where ω is the frequency variable and belongs to the set of real numbers a 0 is the constant base value of amplitude modulation; a 1 is the proportional coefficient related to in amplitude modulation and is a real constant; the center frequency ω baseis the reference center frequency, which is a preset constant; <·> represents the inner product operation; b is the proportionality coefficient; the frequency standard deviation σ 0 is the reference standard deviation, which is a positive preset constant; the local texture map is obtained through the inverse Fourier transform

[0040] where x is the texture map position coordinate; represents the operation of taking the real part; s is the imaginary unit; Ω x is the texture map spatial domain; x ∈ Ω x ; the phase function where c 0 and c 1 are constants.

[0041] Furthermore, in step 4, the obtained local texture map is used to generate the structural layout of the two-dimensional image; let the image be I(x, y, v), and its evolution is driven by the partial differential equation of the diffusion and reaction mechanism:

[0042]

[0043] where the local diffusion coefficient where κ 0 > 0, is a preset value; τ > 0, is a preset value, ensuring that the diffusion weakens at the texture mutation; v is the number of iterations; the decay rate where μ 0 > 0, is a preset value; the texture-driven weight λ T >0, is a preset value; the initial condition is selected as I(x, y, 0) = I 0 (x, y), which is a zero-mean noise image, and after T 3 iterations, the image content I init (x, y) = I(x, y, T 3 ).

[0044] The artificial intelligence-based digital cultural and creative content generation method of the present invention has the following beneficial effects: The present invention effectively improves the mapping accuracy from text semantics to visual content. Existing text-to-image generation methods generally rely on large-scale training data for end-to-end learning. However, in complex semantic description scenarios, it is often difficult to ensure that the generated content can accurately reflect every detail described in the text. The present invention decomposes the input text into multiple semantic segments by presetting grammar and semantic rules, and constructs a semantic scoring function based on mathematical modeling to ensure that the text semantics can maintain a high degree of integrity during the conversion into latent variables. This method not only improves the accuracy of text parsing, but also enhances the model's adaptability to complex text structures, enabling the generated cultural and creative content to more accurately correspond to the description of the input text. Secondly, the present invention optimizes the latent variables, achieving a balance between the diversity and stability of the generated content. In existing methods, the construction of latent variables usually relies on random noise or standard Gaussian distribution, which to a certain extent limits the controllability of the generated content. The present invention combines semantic embedding and geometric optimization methods, introducing specific mathematical constraints during the construction of latent variables, enabling them to be distributed on an implicit geometric structure. This optimization strategy not only ensures the stability of the generated results, but also enhances the creativity of the generated content, allowing the same text input to exhibit subtle but reasonable variations between different instances, thereby improving the richness of cultural and creative content. Another great beneficial effect of the present invention is the improvement of the local detail control ability. In application scenarios such as art creation and cultural heritage restoration, it is often necessary to precisely control local regions of an image to ensure that the texture, color, and structure of specific regions meet expectations. However, existing neural network-based methods have significant limitations in local control, mainly because the end-to-end black-box training mode is difficult to provide resolvable control variables. The present invention introduces a frequency domain modulation method during the generation process, enabling local texture features to be dynamically adjusted according to the state of the latent variables, and achieving more precise local structure control in combination with the Fourier transform. Compared with traditional methods, this method can adjust the details of local regions without affecting the overall style consistency, making the generated content more in line with specific artistic requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0046] Figure 1Schematic flowchart of the method for generating digital cultural and creative content based on artificial intelligence provided by an embodiment of the present invention. Detailed implementation manners

[0047] The method of the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments of the present invention.

[0048] Example 1, refer to Figure 1 : A method for generating digital cultural and creative content based on artificial intelligence, the method comprising:

[0049] Step 1: Divide the input text description into several semantic segments according to predefined syntactic and semantic rules, calculate the local semantic scores for each semantic segment using predefined word vector mapping and activation functions, and calculate the global text embedding using the local semantic scores of each semantic segment and a preset weight;

[0050] This step adopts a text semantic modeling method based on artificial intelligence. By performing syntactic analysis and semantic parsing on the text description, it is divided into several independent but interrelated semantic segments. This division method not only considers the logical levels in the text but also retains the context relationship between words to ensure that the deep features of cultural and creative content can be comprehensively captured when calculating the semantic embedding later. After completing the text semantic decomposition, it is necessary to vectorize each semantic segment for high-dimensional semantic calculation by the computer. For this purpose, the present invention adopts a pre-trained word vector mapping method to project each word into a high-dimensional space so that its semantic information can be quantitatively expressed. In this way, each semantic segment of the text can form a corresponding vector representation, which can not only capture the meaning of the word itself but also implicitly contain its context relationship. After obtaining the vector representation of the semantic segment, it is necessary to further evaluate the semantic importance of each segment to ensure that the contributions of different segments can be fully considered when constructing the global text embedding. For this purpose, the present invention adopts a semantic score calculation method based on non-linear transformation. By introducing exponential mapping and activation functions, the semantic scores can be dynamically adjusted among different segments. This calculation method can effectively avoid the semantic dilution problem that may be caused by the traditional weighted average method, so that when constructing the global text embedding, the segments with higher semantic relevance can obtain higher weights, thereby enhancing the semantic consistency of the generated image.

[0051] After obtaining the scores of all semantic segments, the next step is to construct the global text embedding. Different from the prior art, the present invention does not simply average all segments, but introduces a global semantic fusion method based on weighted similarity. This method first calculates the mean of all segment semantic vectors as the benchmark representation of the overall text semantics, then calculates the Euclidean distance between each segment and this benchmark, and defines a weight function based on this distance to determine the contribution degree of this segment in the global embedding. This weight assignment mechanism ensures that segments closer to the global semantic center can play a leading role in the final embedding, while segments with larger deviations are given lower weights to avoid semantic drift during image generation. In addition, the present invention further normalizes the global text embedding to ensure its numerical consistency with random noise and manifold transformation during the latent variable construction stage. This normalization method can not only improve the computational stability, but also reduce the problem of uneven semantic scales caused by different text input lengths and complexities to a certain extent, enabling the finally generated creative content to maintain a stable semantic expression ability. Through the above method, the text semantic decomposition and global embedding construction process described in the present invention can ensure that the semantic information of the input text is fully utilized and form a stable and analyzable semantic representation in the high-dimensional space. This semantic representation can not only provide high-quality initialization conditions for the subsequent latent variable construction, but also play a global constraint role throughout the generation process, enabling the finally generated digital cultural creative content to accurately reflect the core idea described in the text. In the prior art, text semantic processing usually relies on simple word vector averaging or sentence vector encoding methods, while the present invention significantly improves the accuracy and consistency of semantic expression by introducing mechanisms such as non-linear activation, weight adaptive assignment, and global normalization

[0052] Step 2: Add the global text embedding to the noise generated using the standard normal distribution to construct the initial latent variable; consider the initial latent variable as located on a hidden manifold, construct a local loss function with reference to the global text embedding, calculate the deviation between the initial latent variable and the text embedding through a preset Riemannian metric, and use the Riemannian gradient descent method to perform multiple iterative updates on the initial latent variable to obtain the latent variable after manifold transformation;

[0053] In the first stage of this step, random noise is generated using the standard normal distribution and combined with the global text embedding to construct the initial latent variable. The essence of this process is to introduce a certain degree of randomness, enabling the same text input to exhibit different artistic styles in different generation instances while keeping the core of the text semantic information unchanged. Compared with traditional deterministic generation methods, the present invention uses the method of random perturbation to make the latent variable distributed in different semantic neighborhoods, thus forming a richer creative expression space. The theoretical basis of this method lies in the variational auto-encoding idea in the probabilistic generation model, that is, expanding the latent space of data through noise perturbation so that the generation results can cover a wider range of possibilities. The use of the standard normal distribution ensures the uniformity and controllability of the noise, enabling the random perturbation to provide sufficient creativity without causing too much deviation from the text semantic expression. In addition, an adjustable scaling factor is introduced during the noise injection process to control the influence degree of the noise on the latent variable. By adjusting this factor, the generation results can be flexibly adjusted between being faithful to the text description and showing creative variations, enabling this method to adapt to the generation requirements of different cultural and creative contents.

[0054] In the second stage of this step, considering that the direct addition of the text semantic embedding and the random noise still belongs to the linear transformation in the Euclidean space, a manifold transformation is further introduced on this basis, enabling the latent variable to be more naturally embedded in the high-dimensional latent space and maintaining a reasonable semantic topological relationship in the local structure. The core idea of the manifold transformation is to define a metric space using Riemannian geometry theory, enabling the update of the latent variable to follow certain geometric rules rather than simply moving along the gradient direction in the Euclidean space. Specifically, the present invention first defines a local loss function based on the global text embedding. This loss function measures the distance between the latent variable and the text embedding, and a Riemannian metric tensor is constructed based on this loss to define the local geometric structure. The introduction of the metric tensor enables the curvature of the space to be dynamically adjusted, thus guiding the latent variable to evolve along a more appropriate path during the optimization process. The advantage of this method is that it can ensure that the latent variable still maintains a close connection with the text semantics during the change process, while allowing a certain degree of deformation within the reasonable semantic range to enhance the creativity of the generated content.

[0055] In the third stage of this step, a method based on Riemannian gradient descent is used to optimize the latent variables, making them gradually converge to a better position in the latent space. In traditional gradient descent methods, the update direction of the optimized variables only depends on the gradient of the loss function. However, in this method, through the gradient calculation on the manifold, the update direction is not only affected by the gradient but also regulated by the local geometric structure. Specifically, in each iterative update, first, the deviation between the current state of the latent variables and the text semantic embedding is calculated, and the inverse matrix of the local metric tensor is calculated based on this deviation. Subsequently, by calculating the gradient on the manifold and normalizing it with the Riemannian metric, an optimization direction that better conforms to the topological structure of the latent space is obtained. This process ensures that the evolution of the latent variables is not just a simple numerical optimization process but a dynamic adjustment process based on geometric constraints, enabling the generated content to maintain a high degree of coherence and controllability between different styles and text descriptions.

[0056] Compared with existing text-to-image generation methods based on GAN or VAE, the present invention constructs latent variables using manifold transformation and combines methods of random perturbation and geometric optimization, enabling the latent variables to not only cover a richer creative space but also maintain high accuracy and stability in semantic expression. Traditional methods often adopt a fixed latent distribution, resulting in limited diversity in the generated results. In contrast, by introducing manifold optimization under global text semantic constraints, this method endows the generation process with both artistic freedom and the ability to not completely deviate from the semantic intent of the text description. Additionally, existing methods usually use Euclidean gradient descent when optimizing latent variables. Although this method is computationally simple, it cannot adapt to the complex changes in high-dimensional non-linear spaces. The present invention, by introducing Riemannian geometry methods, enables the optimization process to perform smoother and more efficient searches in higher-dimensional latent spaces, thereby improving the quality and diversity of the generated results.

[0057] Step 3: Use the transformed latent variables as the initial state, construct a potential energy function including text semantic constraints and a global regularization term in the latent space, and realize the diffusion update of the latent variables through a discretized stochastic differential equation; use the diffused and updated latent variables as the input, determine the texture amplitude, center frequency, and frequency diffusion parameters according to a preset frequency modulation function, and generate a local texture map using the inverse Fourier transform method;

[0058] Based on the diffusion model of stochastic differential equations, this step dynamically evolves the latent variables so that their distribution in the latent space can be more uniform and stable. In physical phenomena in nature, the diffusion process usually describes the random movement of particles in a fluid. In the generation system of the present invention, the introduction of the diffusion process aims to enable the latent variables to adaptively adjust within a semantically reasonable range, avoiding being trapped in local extreme points due to over-optimization, which may lead to a lack of diversity in the generated results. The diffusion process combines the text semantic information with random perturbations by constructing a specific potential energy function, enabling the latent variables to change along a path that conforms to the logic of artistic expression. This potential energy function consists of two key parts: one is the global text semantic constraint term, which ensures that the latent variables do not deviate from the original semantic information, making the generated images still faithful to the text description; the other is the regularization term, which is used to control the range of change of the latent variables, avoiding problems such as excessive divergence or too fast convergence during the update process. By continuously adjusting the weight ratio of the potential energy function during the diffusion process, the evolution of the latent variables has both strong stability and can explore more possibilities to a certain extent, providing richer basic information for subsequent image synthesis.

[0059] After completing the diffusion update of the latent variables, this step further adopts a frequency-domain modulation method to ensure the controllability of the generated content in terms of visual details. Traditional generation methods based on the pixel space often have difficulty accurately controlling the local texture distribution of images, while the frequency-domain method can directly adjust the style, detail level, and local texture of the image by adjusting the weights of different frequency components, making the generated content more in line with specific artistic expression requirements. In the method of the present invention, the core idea of frequency-domain modulation is to construct a set of modulation functions based on the Fourier transform, and the amplitude and phase of this modulation function are controlled by the latent variables, enabling the finally generated texture map to be naturally embedded into the overall cultural and creative image. First, by calculating the norm of the latent variables, the basic amplitude of the modulation function is determined, and this amplitude determines the overall intensity of the texture, enabling the texture details under different semantic descriptions to be adaptively adjusted; second, according to the relationship between the latent variables and the text semantic embedding, the central frequency of the texture is dynamically calculated, enabling different text inputs to correspond to different main texture structures and achieving style consistency; finally, by adjusting the frequency diffusion parameter, the details of the texture can be finely changed according to the text description. For example, when the input text describes soft cultural elements, the range of frequency diffusion is small to generate a more uniform and smooth texture, while when describing content with strong artistic impact, the range of frequency diffusion is large to enhance the visual contrast and dynamic change effect.

[0060] After completing the frequency-domain modulation, this step converts the modulated frequency-domain information back to the spatial domain through inverse Fourier transform to generate a local texture map. The role of the inverse Fourier transform is that it can remap the frequency components to the pixel space, enabling the modulated features to be directly reflected in the final image generation process. The introduction of a phase modulation function is also involved in this process to ensure that the spatial distribution of the texture map can match the semantic information of the original text. The core idea of phase modulation is to calculate the Euclidean distance between the latent variable and the text semantic embedding and map it as a phase offset, thereby fine-tuning the local features of the texture during the reconstruction process, so that the finally generated image not only conforms to the overall semantic expression but also has rich local detail variations. In this way, the method of the present invention can enhance the artistic style of the image while maintaining the overall visual beauty, making it more in line with the needs of the cultural and creative industries.

[0061] Step 4: Use the generated local texture map as a local driving factor to construct a partial differential equation based on the diffusion and reaction mechanism. This partial differential equation evolves the spatial and temporal distribution of the image brightness under preset initial conditions to achieve the adaptive synthesis of the image structure layout and generate the image content.

[0062] Specifically, during the execution of this step, first, the local texture map generated in the previous stage is used as a driving factor to provide preliminary guidance for the structural layout of the entire image. Since the local texture map is generated after the latent variables undergo diffusion updates and frequency-domain modulation, it already contains a large number of visual features corresponding to the text semantics, such as color variations, boundary features, and texture densities in different regions. Based on this texture map, a set of partial differential equations is constructed to describe the evolution process of the image brightness distribution, enabling the local structure of the image to be dynamically adjusted over continuous time, thereby achieving the gradual construction of the image from the initial state to the final complete image. The introduction of partial differential equations endows the entire generation process with a property similar to physical simulation, ensuring that the features in different regions do not change abruptly but evolve gradually in a smooth manner, thus forming a more natural composition. This method is different from traditional image generation networks, which usually directly generate the final image through the end-to-end mapping of neural networks. Instead, this method introduces a physical constraint during the generation process, making the layout and detail evolution of the content more in line with natural laws. During the construction of the partial differential equations, a combination of a diffusion term and a reaction term is adopted to simultaneously control the global smoothness and local detail changes of the image. Among them, the role of the diffusion term is to enable adjacent pixels to interact through the calculation of local gradients, thereby ensuring a smooth transition of the image; the role of the reaction term is to enhance or suppress certain specific structural features in specific regions according to the driving information provided by the local texture map. This combination method enables the generation process to ensure global consistency while being able to flexibly adjust in local regions, thus providing adaptability for different types of cultural and creative content. For example, when generating an image in the style of traditional Chinese ink painting, the diffusion term can control the natural rendering of the ink, while the reaction term can adjust the color concentration and brushstroke patterns in different regions to make the final image more in line with the visual characteristics of ink painting; when generating a future science fiction style work, the reaction term can enhance the high-contrast structures in local regions, making the finally generated image more technological and futuristic. Therefore, by controlling the different weight ratios of the diffusion term and the reaction term, the adaptive generation of images with different styles can be achieved, making this method more adaptable in terms of the diversity of cultural and creative content generation. During the numerical solution process of the partial differential equations, a discretization method is used to approximately calculate the continuous time evolution process to ensure that the generation process can converge to a stable image structure within a finite time. The key to the numerical solution lies in how to design reasonable boundary conditions so that there are no discontinuities or abrupt transitions at the edge part of the image. For this purpose, this method adopts a strategy based on Neumann boundaries when setting the boundary conditions, that is, the gradient is forced to be zero in the boundary region of the image, thereby ensuring the consistency of the content at the edge part without obvious truncation or jump.In addition, to further enhance the level of detail in the generated results, this method also introduces a dynamic adjustment mechanism based on local feature comparison in the reaction term. That is, when calculating the reaction intensity, not only the information provided by the local texture map is considered, but also the local contrast of the current state of the image is additionally added to ensure that high-contrast regions can remain clear during the evolution process, while low-contrast regions can transition more naturally. Through this dynamic adjustment method, the finally generated image not only has a strong visual hierarchy but also ensures that the overall artistic style conforms to the semantic intention described in the text.

[0063] Embodiment 2: In step 1, let the input text description be T = {w 1 , w 2 ,..., w j ,..., w M}; where w j represents the j-th word, j is an integer subscript index, and its value range is from 1 to M, where M is the total number of words; the predefined word vector mapping maps the word w j to a real number vector; represents the vocabulary, which is the set of all possible words; is the dimension of the word vector, indicating that the word w j has feature values in the word vector space; each word is mapped to the vector space to obtain the word vectors φ(w j ); using the predefined syntactic and semantic rules, the text is decomposed into N semantic fragments, and the i-th semantic fragment is is the j-th word in the i-th semantic fragment; the word index n i,j represents the position of this word in the entire text; L i is the number of words in the i-th semantic fragment; N is the total number of semantic fragments after the text is partitioned.

[0064] Specifically, the input text description is represented as a word sequence T = {w 1 , w 2 ,..., w M}, where w j represents the j-th word in the text, M is the total number of words, and j is an integer index indicating the order of the word in the text. Since words in natural language text have different meanings and contexts, it is necessary to use the predefined word vector mapping function φ(w j ) to map each word w jis mapped to a high-dimensional vector space so that it can express its semantic information numerically. This mapping function is usually trained by deep learning models (such as the embedding layer of Word2Vec, GloVe, or Transformer language models) so that the semantic relationships between words can be preserved in the high-dimensional vector space. Specifically, φ(w j ) has a domain of the vocabulary V, which is the set of all possible words that can appear, and a range that is a -dimensional vector space, where represents the dimension of the word in the word vector space. Under this mapping, each word w j can be represented as a real-valued vector with feature values, such that the semantic information of the text can be measured by calculating the similarity metric between vectors. After obtaining the vector representations of all words, it is necessary to further structure the text so as to construct a semantic embedding suitable for calculation in subsequent steps. Since the semantics of the text are usually not determined by a single word, but rather by combinations of multiple words expressing more complex concepts, simply representing words independently is not sufficient to support high-quality content generation. Therefore, in the second stage of this step, we split the text into N independent semantic fragments T i , where each semantic fragment T i is a separate phrase or sentence that can express complete semantic information. Each semantic fragment T i consists of a group of words , where n i,j represents the position index of the word in the entire text, and L iRepresents the number of words in this semantic segment. The total number of semantic segments into which the text is divided is denoted by N, which is determined by the length, complexity of the text, and the semantic segmentation algorithm. During the text segmentation process, multiple factors need to be comprehensively considered to ensure that the division of semantic segments can accurately reflect the hierarchical structure and semantic logic of the text. On the one hand, syntax-based segmentation rules can ensure that the text is reasonably split according to sentences or phrases. For example, punctuation marks such as full stops, commas, and semicolons can be used as the basis for division. On the other hand, semantics-based segmentation rules can group words with closely related semantics into the same segment according to the semantic similarity, co-occurrence relationship, or dependency relationship between words. For example, in natural language processing, common methods include autoregressive segmentation based on the attention mechanism, clustering segmentation based on word embedding similarity, and tree-structured segmentation based on syntactic dependencies. These methods can effectively capture the hierarchical relationships in the text, enabling the final semantic segments to more accurately reflect the core meaning of the text. Compared with traditional text processing methods, the method of the present invention has stronger adaptability and intelligence in semantic decomposition. Traditional methods usually adopt a fixed window size or rule-based sentence segmentation methods, while the method of the present invention combines artificial intelligence technology and can dynamically adjust the segmentation strategy according to the specific content of the text, making the division of semantic segments more flexible. For example, when describing an abstract painting, it may be necessary to integrate the descriptions of multiple visual elements into a longer semantic segment, while when describing an art work with more details, the text needs to be split into smaller sub-units to more accurately express the characteristics of each part. This ability of dynamic adjustment enables this method to adapt to different expression requirements when processing different types of cultural and creative texts, thus ensuring a higher semantic consistency between the generated image content and the input text.

[0065] Embodiment 3: In step 1, define a semantic scoring function for each semantic segment:

[0066]

[0067] where, ||·|| 1 represents the L 1 norm operation; ψ(T i ) is the local semantic score; and use the Sigmoid activation function to construct the segment semantic vector:

[0068] s i = σ(ψ(T i ) - θ)v 0 ;

[0069] where, s i is the segment semantic vector of the i-th semantic segment; σ(·) is the Sigmoid activation function; θ is the activation threshold, which is a preset scalar used to shift ψ(T i) output to control the response interval of the Sigmoid activation function; v 0 is the basic semantic direction vector, which is a preset vector in ; d s is the dimension of the semantic vector, which determines the size of the feature space of the global text embedding; the global text embedding is obtained by weighted averaging the semantic vectors of each segment:

[0070]

[0071] wherein, represents the global text embedding; ||·|| 2 represents the L2 norm.

[0072] Specifically, in order to measure the importance of different text segments, this step defines a semantic scoring function ψ(T i ). This scoring function calculates the L 1 norm of all word vectors in the segment and normalizes them using an exponential transformation, so that the semantic score can reflect the overall semantic activity of the segment. The use of the L 1 norm ensures that the contributions of word vectors in all dimensions can be considered, while the exponential transformation can emphasize words with high semantic intensity and avoid the unnecessary influence of words with low semantic intensity on the overall score. In addition, in order to ensure the numerical stability of the semantic score, the function also introduces a logarithmic transformation, so that the scoring function can maintain good numerical distribution characteristics under different text inputs. The essence of this scoring method is to perform a non-linear mapping on the features of different semantic segments in the word vector space to ensure that the subsequent semantic vectors can better reflect the hierarchical structure of the text. After calculating the semantic score of the text segment, this step further uses the Sigmoid activation function to transform the score to obtain the semantic vector s i of the segment. The introduction of the Sigmoid function normalizes the numerical range of the score to between (0, 1), and by controlling the activation threshold θ, different intensities of semantic scores can obtain different responses during the mapping process. The core function of this process is that through non-linear mapping, segments with low semantic intensity can be weakened, while segments with high semantic intensity can be more prominent, thus providing more discriminative inputs when constructing the global text embedding. In addition, the basic semantic direction vector v 0, this vector provides a reference direction in the entire embedding space, enabling the semantic representations of all segments to be more stable in the high-dimensional space and reducing the instability caused by numerical differences between semantic segments. In this way, this step ensures that the representation of each semantic segment in the embedding space has strong distinctiveness and stability, laying the foundation for the calculation of subsequent global embeddings. After obtaining the semantic vectors of all segments, this step adopts a global text embedding method based on weighted average to ensure that the global semantics can integrate information from different segments and have stronger robustness. Traditional text embedding methods usually adopt a simple average strategy, while the method of the present invention introduces a weighted average method based on Gaussian weights. By calculating the Euclidean distance between each segment vector and the mean of all segment vectors and normalizing the weights using the Gaussian kernel function, segments closer to the global center have higher weights, while segments far from the center contribute less to the final embedding. The advantage of this method is that it can adaptively adjust the influence of different segments, ensuring that the final global text embedding can more accurately reflect the main semantic information of the text and avoiding the deviation of the global embedding caused by the existence of certain abnormal segments. In addition, the use of Gaussian weights can effectively reduce the influence of noise, making the calculation of the global embedding smoother and ensuring the stability of the generated results.

[0073] Example 4: In step 2, the initial latent variable is constructed through the following process: For k = 1,..., d s , define:

[0074]

[0075] U k , V k ~ Uniform(0, 1);

[0076] where U k and V k are random variables independently sampled from the uniform distribution Uniform(0, 1) used for generating the k-th noise component δ k ; all the noise components form a noise vector denotes the d s -dimensional identity matrix to ensure no correlation between the dimensions of the noise; the initial latent variable ξ 0 is constructed using the noise vector:

[0077] ξ 0 = e + λδ;

[0078] where λ adjusts the noise amplitude, and its value range is from 0.1 to 0.5.

[0079] Specifically, the core idea of the Box-Muller transform is to construct random variables with a standard normal distribution from uniformly distributed random variables to ensure that in a high-dimensional space, the noise components of each dimension conform to the properties of the normal distribution. Specifically, for each dimension k of the latent variable, two independent random variables U k and V k are first sampled from the uniform distribution in the interval [0, 1], and then the noise component δ k that follows the standard normal distribution is calculated using the transformation relationship. In this process, the operation of lnU k is used to transform the uniform distribution into a variable that conforms to the exponential distribution, and cos(2πV k ) maps this value to the angular space of the standard normal distribution, so that the finally obtained noise follows a normal distribution with a mean of zero and a variance of one. The advantage of this method is that it can efficiently generate standard normal distribution variables from independent uniformly distributed variables through simple mathematical transformations, and the generated noise is independent in different dimensions, so that the finally constructed latent variable can maintain good distribution properties. This method ensures the uniformity of the noise in the entire latent space, so that the generated cultural and creative content can show rich diversity between different instances, rather than being limited to a specific pattern or style. After generating the high-dimensional standard normal noise vector, the next step is to combine this noise vector with the text semantic embedding vector to construct the initial latent variable. The text semantic embedding vector e is obtained from the previous steps, which contains the core semantic information of the input text and is the basis of the entire generation system. However, if only the semantic embedding vector is used for generation, then the final cultural and creative image will lack diversity, because the same text input will always be mapped to the same point in the embedding space, resulting in a fixed pattern of the generated content. Therefore, in this step, we introduce a noise vector and control its influence with a scaling factor λ to construct a latent variable ξ 0Among them, the role of λ is particularly crucial as it determines the trade-off between the semantic embedding vector and the random noise. If λ takes a smaller value, the initial latent variable is almost dominated by the text semantic embedding, and the generated cultural and creative content will mainly rely on the input text with lower randomness, resulting in relatively stable generated content for the same text input. If λ takes a larger value, the influence of the noise increases, the diversity of the generated content increases, but it may lead to a deviation from the text semantics. Therefore, λ needs to be adjusted within a reasonable range to ensure that the generated content is both faithful to the text description and has sufficient artistic innovation. In the present invention, the value range of λ is set to 0.1 to 0.5 to ensure that in different types of cultural and creative content generation tasks, the generation style can be flexibly adjusted to be applicable to different application scenarios, such as virtual art creation, stylized painting generation, and cultural visual design. Compared with traditional text-to-image generation methods, the present invention introduces a more rigorous mathematical method in the initialization process of the latent variable to ensure that the generated noise conforms to the normal distribution characteristics and can provide reasonable perturbations in the latent space, enabling the generated content to show stronger creative variations between different instances. Traditional latent variable initialization methods usually adopt simple Gaussian noise superposition or random sampling methods. Although such methods can provide a certain degree of randomness, they often lead to uneven distribution of noise in the high-dimensional space, thereby affecting the quality of the final generated content. The present invention generates standard normal noise through the Box-Muller transform and ensures the independence between the dimensions of the noise, so that the initial latent variable has a more reasonable distribution in the entire latent space, making the subsequent generation process more stable. At the same time, the adaptive adjustment mechanism of λ ensures the organic combination of the text semantic embedding and the random noise, enabling the generated content to maintain semantic consistency and show more style variations at the detail level. The advantage of this method is that it can provide more flexible control capabilities in the artistic stylized image generation task, enabling different text inputs to show different degrees of changes in the generated content by adjusting the randomness weight, so as to meet different creative needs.

[0080] Embodiment 5: In step 2, let the initial latent variable be ξ 0 be distributed on the implicit manifold and its local geometry is described by a metric constructed based on a constraint of the global text embedding, and define the local loss function as:

[0081] Specifically, in traditional text-to-image generation methods, the initialization of latent variables usually adopts a random distribution or directly samples from Gaussian noise. Although this approach can provide a certain degree of randomness, it often lacks fine control over text semantics and easily leads to instability in the style and structure of the generated results. In the method of the present invention, we assume that the latent variables are not uniformly distributed throughout the high-dimensional space but are constrained by text semantics and distributed on a specific manifold M. Therefore, in the process of optimizing the latent variables, a reasonable loss function needs to be defined to enable adjustment under the local geometric structure of the manifold. The role of this loss function is to minimize the distance between the initial latent variable ξ 0 and the global text semantic embedding e, so that the latent variables converge stably on the manifold without excessive deviation or random drift. At the same time, cos(2πV k ) as a periodic modulation term, its main role is to introduce controlled fluctuations in the latent variables, so that their distribution can have a certain style change between different instances without falling into a fixed pattern. The core mathematical idea of this design comes from the modulation method in signal processing, that is, by applying a periodic factor to a variable, making it change regularly within a specific range, thereby enhancing the artistic style diversity of the generated content. In this step, the introduction of this periodic term enables the latent variables not to fall into local extrema during the optimization process, but to reasonably explore on the implicit manifold under the premise of conforming to text semantics, so that the generated cultural and creative content has both stable semantic expression and controllable changes in artistic style.

[0082] Compared with the traditional L2 loss function, the loss function of the present invention makes the optimization process of latent variables more in line with the geometric characteristics of the high-dimensional manifold by combining periodic modulation. During the training process of the generation model, if the optimization process of latent variables overly relies on the fixed L2 loss, it is easy to cause all latent variables to tend to a static mean point, resulting in limited diversity of the generated results. In this method, cos(2πV k)This makes the loss term have a certain variability during different optimization iterations, thereby avoiding potential variables falling into local minima and promoting natural variations in the generated content among different styles. For example, in the task of generating cultural and creative content, if the input text describes a visual art work with an oriental ink painting style, during the optimization process of the potential variables, the periodic modulation term can make the generated brushstroke details vary among different instances, making the texture of some parts show a stronger brush touch of a writing brush, while other parts may be more delicate and soft. If the input text describes an art work with a digital fantasy style, this modulation term can be used to enhance the modulation of different frequencies in the local area of the potential variables, so that the finally generated work can show a more dreamy and futuristic light and shadow effect. The introduction of this controllability enables the generation method of the present invention not only to be consistent with the text description in terms of the overall style, but also to achieve highly flexible style adjustment at the detail level to meet the requirements of different application scenarios in the digital cultural and creative industry. From the perspective of geometric optimization, the loss function of this step can be understood as a constrained optimization problem, that is, on the high-dimensional manifold M, the potential variables need to maintain the smoothness and coherence on the manifold as much as possible while satisfying certain text semantic constraints. Since the manifold itself is a non-linear subspace, during the optimization process, the gradient descent direction of the potential variables is not only restricted by the L2 norm, but also affected by the local geometry. The introduction of the periodic modulation term enables the update direction of the potential variables to show a certain change among different instances to enhance the dynamics of the finally generated content. The application of this method in the art creation task is particularly important because art works often have certain style characteristics, but at the same time need to show unique detail variations among different instances, and this method can achieve this goal through the geometric optimization of the potential variables, making the generated works have both unified style characteristics and rich variations in specific details. For example, in the task of digital restoration of cultural heritage, different input texts may describe different parts of the same ancient mural, and this method can ensure local adjustment of the art styles of different regions on the premise of consistent overall style to more accurately restore the detail characteristics of cultural heritage. In addition, in the task of immersive art generation, the optimization process of this method can ensure that the generated content still meets the specific visual style requirements under different angles and different lighting conditions, thereby improving the stability and adaptability of the generation system.

[0083] Example 6: In step 2, the process of calculating the deviation between the potential variable and the text embedding through a preset Riemannian metric and performing multiple-step iterative update on the potential variable using the Riemannian gradient descent method to obtain the potential variable after manifold transformation includes:

[0084] The preset Riemannian metric is g:

[0085]

[0086] where ξ m is the m-th component of ξ 0 ; ξl is the l-th component of ξ 0 ; e m is the m-th component of e; e l is the l-th component of e; g ml is the element in the m-th row and l-th column of the Riemannian metric g; Update the latent variable using Riemannian gradient descent through the following formula:

[0087]

[0088] where g -1 represents the inverse matrix of the Riemannian metric; the finally obtained latent variable after the manifold transformation is T 1 is the total number of iterations; is the gradient operator; is the local loss function of the latent variable ξ t-1 obtained in the (t - 1)-th iteration; ξ t is the latent variable obtained in the t-th iteration.

[0089] Specifically, in the optimization process of this step, it is first necessary to define the geometric structure of the space where the latent variable is located. Assume that the manifold M where the latent variable ξ 0 is located is not a uniform Euclidean space, but a space with a certain curvature affected by the text semantic embedding e. To describe the local geometric properties of this space, the Riemannian metric g is introduced. Its definition method ensures that when the latent variable is close to the text embedding, the influence of the metric is small, making the optimization process closer to the gradient descent of the standard Euclidean space. When the latent variable is far from the text embedding, the influence of the metric gradually increases, thereby guiding the latent variable to return to the semantic center. The core idea of this metric is to adjust the local geometric structure of the space so that the latent variable can change along a more reasonable path during the optimization process, rather than simply updating along the gradient direction, thus avoiding the problem of semantic drift caused by over-optimization. The construction method of the Riemannian metric fully considers the relative relationship between the latent variable and the text embedding. The key correction term (ξ m -e m )(ξ l -e l ) reflects how the changes of the latent variable in different dimensions are aligned with the text semantic embedding. At the same time, through the normalization factor Ensure the numerical stability of the optimization process so that even when the latent variables are very close to the text embeddings, the optimization can still converge smoothly. This approach is actually a weighted method based on local structure, that is, when the latent variables are close to the text embeddings, the optimization process is subject to less geometric constraints, allowing for more free exploration. While when the latent variables are far from the text embeddings, the geometric constraints are enhanced, making the optimization direction more controlled, thus ensuring that the finally generated content still conforms to the text semantic description.

[0090] After defining the Riemannian metric, the core of the optimization process is how to perform gradient updates based on this metric. In traditional gradient descent methods, the update of variables only depends on the calculation of local gradients, that is, taking steps along the direction of gradient descent. While in the Riemannian space, due to the influence of the metric, the gradient needs to be corrected through the inverse transformation of the metric tensor so that the update direction can better conform to the geometric structure of the manifold. This means that in this method, the update formula of gradient descent not only involves the gradient of the loss function, but also needs to use the inverse matrix of the metric g to transform the gradient to ensure that the optimization path can be adjusted along the optimal manifold direction. The advantage of this method is that it can make the optimization process of latent variables more conform to the structure of the data manifold, thus ensuring the stability and semantic consistency of the generated content. For example, in the task of stylized art generation, traditional methods often have difficulty achieving smooth transitions between different styles. While this method makes the changes between different styles more natural and can maintain a certain coherence between different instances by geometrically constraining the optimization path of latent variables. In addition, during the optimization process of this method, the setting of the learning rate has also been finely adjusted. Since the step size in the gradient descent process directly affects the convergence speed and stability of the optimization, a fixed step size of 0.05 is used for update to ensure that the optimization process can converge to a stable state within a reasonable time. Compared with the adaptive learning rate method, this method using a fixed step size makes the update process of latent variables more stable, avoiding the optimization oscillation problem caused by excessive changes in the learning rate. In addition, since the optimization process is based on manifold geometry, the setting of the step size needs to match the numerical range of the metric to ensure that the optimization process can maintain a smooth update throughout the latent space. Through this optimization method, the finally obtained latent variables Through multiple steps of iteration, it converges to a stable point on the manifold, enabling the generated content to achieve a balance between artistic style and semantic consistency. Compared with traditional text-to-image generation methods, the method of the present invention introduces the concept of Riemannian geometry in the process of optimizing latent variables, making the generated content not only more stable in the overall style, but also able to show richer variations in local details. For example, in digital art creation, works of certain specific styles may require different brushstroke features to be shown in different local areas, and this method can, through the geometric constraints in the optimization process, make these style features consistent among different instances while showing local diversity. In addition, in the personalized customization task of cultural and creative content, this method can adaptively adjust the optimization path in the latent space according to the characteristics of different text inputs to ensure that the generated results can meet the semantic requirements of specific text descriptions while showing rich variations in local details.

[0091] Example 7: In step 3, on the basis of the obtained , a diffusion process is introduced to further balance the text semantic constraint and generation exploration, and the following potential energy function with a global regularization term is constructed

[0092] where δ 1 > 0, which is a preset value, providing a global contraction effect; the stochastic differential equation introducing the diffusion process is:

[0093]

[0094] where D > 0, which is a preset diffusion coefficient, and W(t) is a standard Wiener process; let, define the discretization step size Δt, and the iteration formula is:

[0095]

[0096] where After another T 1 iterations, the finally obtained latent variables after diffusion update are

[0097] Specifically, in this step, a potential energy function is first defined, which contains two main parts. One is the constraint term that ensures that the latent variables do not deviate too far from the text semantic embedding, that is This part ensures that the latent variables are always adjusted around the text semantic embedding e, so that the finally generated cultural and creative content always maintains a close association with the original text description. This constraint avoids semantic drift of the generated content due to increased randomness during the optimization process, making the finally generated image not only highly artistic but also able to accurately reflect the core meaning of the text. The second is this regularization term, whose role is to introduce a global contraction effect to ensure that the latent variables do not expand without limit but maintain a certain convergence trend as a whole. This global regularization mechanism ensures that the evolution of the latent variables will not diverge excessively due to the complexity of the high-dimensional space, thus enhancing the stability of the optimization process. In addition, the parameter δ 1 as a controllable global contraction parameter, the setting of its value directly affects the stability and exploration ability of the finally generated content. When δ 1 takes a small value, the latent variables have a higher degree of freedom and can explore a larger range in the semantic space, making the generated content show richer variations; while when δ 1 takes a large value, the distribution of the latent variables will be more concentrated, making the style of the generated content more consistent. Therefore, the choice of δ 1 needs to be adapted according to the specific cultural and creative application scenarios. For example, in digital art drawing tasks, it may be desired that the generated results have higher variability, while in cultural heritage digital restoration tasks, more stable style consistency may be required. After constructing the potential energy function, this method further introduces a diffusion process to iteratively update the latent variables in the form of a stochastic differential equation. The basic idea of the diffusion process is to decompose the change of the latent variables into two parts: one part is controlled by the gradient of the potential energy function, whose main role is to guide the latent variables to converge along the semantic embedding direction; the other part is controlled by a stochastic noise term driven by a standard Wiener process, enabling the latent variables to explore moderately in the high-dimensional space to ensure that the generated content still has a certain degree of style variation between different instances. The introduction of this diffusion mechanism makes the optimization of the latent variables no longer a simple gradient descent process but a controlled evolution in a dynamic potential energy field, thus ensuring semantic consistency of the generated content while still allowing a certain degree of style deviation.

[0098] In the specific implementation process, a discretized iterative formula is adopted to simulate the diffusion process, enabling the latent variables to gradually evolve to a stable state within multiple time steps. The iterative update method ensures that each update can be adjusted based on the current potential energy gradient, and at the same time, by introducing an additional stochastic noise term, the optimization process will not fall into a local optimal solution but make an adaptive adjustment within a reasonable range. The gradient term is composed of Given, ensure that the optimization direction is always constrained by the text semantic embedding, while maintaining the contraction effect as a whole, making the optimization path more in line with the generation requirements of cultural and creative content. And the random noise term is responsible for providing additional exploration ability during the optimization process, so that the latent variables will not be limited to a fixed path during the evolution process, but can freely transform within a certain range, so that the generated cultural and creative content can show richer changes in style expression. The parameter D, as the diffusion coefficient, is used to control the influence degree of random noise on the optimization process. A smaller D value means that the evolution process of the latent variables is mainly constrained by the potential energy gradient, making the generated content more stable, while a larger D value will enhance the influence of the noise term, thus increasing the variation range of the generated content. Therefore, the setting of the D value needs to be adjusted according to specific application requirements to achieve a balance between stability and diversity. Compared with traditional text-to-image generation methods, the optimization method of the present invention not only relies on deterministic gradient optimization, but also combines a random diffusion mechanism, making the generated cultural and creative content show richer changes between different instances. Traditional optimization methods usually rely on a fixed gradient direction. Although this method can ensure that the generated content is consistent with the text description, it often leads to insufficient variation between different instances of the generated results. However, this method introduces controlled random diffusion during the optimization process, so that the generated content under the same text description can show different visual features to a certain extent, thus enhancing the expression ability of cultural and creative content. For example, in the art stylization task, this method can ensure that under the same text input, the generated images can still maintain basic style consistency, but still have different artistic expressions at the detail level, thus avoiding overly mechanical repetition of the generated content. In addition, in the digital cultural heritage restoration task, this method can ensure that between different instances, the generated cultural elements can maintain basic historical consistency, while showing different artistic changes in local details to better match different visual performance requirements.

[0099] Example 8: In step 3, the frequency modulation function has the formula:

[0100]

[0101] where ω is the frequency variable, belonging to the set of real numbers a 0 is the constant base value of amplitude modulation; a 1 is the proportional coefficient related to in amplitude modulation, which is a real constant; the center frequency ω baseis the reference center frequency, which is a preset constant; <·> represents the inner product operation; b is the proportionality coefficient; the frequency standard deviation σ 0 is the reference standard deviation, which is a positive preset constant; the local texture map is obtained through the inverse Fourier transform

[0102] where x is the texture map position coordinate; represents the operation of taking the real part; s is the imaginary unit; Ω x is the texture map spatial domain; x ∈ Ω x ; the phase function where c 0 and c 1 are constants.

[0103] Specifically, the frequency modulation function is the core part of this step, which determines the frequency distribution of the final texture and the detail changes in different regions. This function is jointly composed of an amplitude modulation term and a center frequency control term, enabling the frequency characteristics to be adaptively adjusted according to the state of the latent variable. In the amplitude modulation part, controls the overall intensity of the frequency amplitude, where a 0 is the reference amplitude, and by introducing the norm of the latent variable, the amplitude can be dynamically adjusted according to the state of the latent variable. This mechanism ensures that when the latent variable is closer to the text semantic embedding, the overall amplitude of the texture is more stable, and when the latent variable is far from the text embedding, the intensity of the texture can change, thus showing richer texture features in local details. Compared with the traditional fixed amplitude modulation method, this method enables the generated images to maintain semantic consistency among different instances while showing more diverse styles through the adaptive control of the latent variable. The setting of the center frequency is equally crucial, which determines the main frequency components of the texture and affects the detail features of the final generated image. In this method, the center frequency is given by where ω base is the reference center frequency, and by calculating the inner product of the latent variable and the text semantic embedding, the center frequency can be adaptively adjusted according to the state of the latent variable. The mathematical significance of this mechanism is that The inner product with e reflects the similarity between the latent variable and the text semantics. Therefore, when the latent variable is closer to the text semantic embedding, the central frequency is closer to the preset reference value. When the latent variable changes, the central frequency can be adjusted accordingly, ensuring that the texture features of the generated content can be adaptively adjusted with the change of text semantics. Compared with the traditional method of fixed central frequency, this method enables the generated image to better match the style requirements of different text descriptions and shows more delicate texture changes in local details through the control of the central frequency based on the latent variable. In addition, to further control the diffusion range of the frequency, this method introduces the frequency standard deviation where σ 0 is the reference standard deviation, and in the denominator By measuring the distance between the latent variable and the text semantic embedding, the standard deviation can be dynamically adjusted according to the state of the latent variable. When the latent variable is closer to the text semantic embedding, is larger, meaning that the frequency diffusion is wider, making the texture softer; when the latent variable is far from the text semantic embedding, is smaller, making the frequency distribution more concentrated, thus enhancing the local detail features. This mechanism ensures that the generated cultural and creative content can maintain overall consistency among different instances while showing rich changes in local details, thereby enhancing the artistic expressiveness of the generated results.

[0104] After constructing the frequency modulation function, we convert it into the local texture map T(x) through the inverse Fourier transform. The role of the inverse Fourier transform is to convert the modulation information in the frequency domain back to the spatial domain, enabling the finally generated image to reflect the characteristics of frequency modulation at the pixel level. Among them, the phase function The phase changes at different positions are controlled, enabling the texture to exhibit different local features in different regions. This phase control mechanism ensures that the generated cultural and creative content can display different texture features in different regions, thereby enhancing the layering and dynamic change ability of the generated image. Compared with the traditional direct texture mapping method, this method performs modulation in the frequency domain through inverse Fourier transform, making the generated image richer in local details and more controllable. Compared with the traditional generation method based on pixel space, the core innovation of this method lies in using frequency domain modulation technology, enabling the changes in latent variables to be directly mapped to frequency characteristics, so that the generated cultural and creative content can exhibit more abundant artistic expressiveness in local details. Traditional methods usually rely on the end-to-end training of neural networks, while this method directly controls in the frequency domain space through mathematical analysis methods, making the generated content not only more natural in style but also have higher fineness in local details. For example, in the digital art stylization task, this method can ensure that under different text inputs, the generated content maintains consistency in the overall style while showing different texture changes in local details. In addition, in the digital restoration task of cultural heritage, this method can ensure that different artistic features are displayed in local areas on the premise of consistent overall visual style, thus better matching the visual expression needs of historical culture.

[0105] Example 9: In step 4, use the obtained local texture map to generate the structural layout of the two-dimensional image; let the image be I(x, y, v), and its evolution is driven by the partial differential equation of the diffusion and reaction mechanism:

[0106]

[0107] where the local diffusion coefficient where κ 0 >0, is a preset value; τ>0, is a preset value, ensuring that the diffusion weakens at the texture mutation; v is the number of iterations; the attenuation rate where μ 0 >0, is a preset value; the texture driving weight λ T >0, is a preset value; the initial condition is selected as I(x, y, 0) = I 0 (x, y), which is a zero-mean noise image. After T 3 iterations, the image content I init (x, y) = I(x, y, T 3 ).

[0108] Specifically, the key to this step lies in defining a partial differential equation based on the diffusion and reaction mechanism to describe the change behavior of the image during the evolution process. The evolution of the image is controlled by the time variable t. Define the state of the image at any moment as I(x, y, v), where (x, y) are the spatial coordinates of the image, and v represents the number of iterations. The core driving force of the entire evolution process comes from the local diffusion, decay mechanism, and the generation term driven by the texture, enabling the image to gradually form the final form that conforms to the text semantic description during the adaptive adjustment process. The diffusion term is responsible for smoothing the image content, allowing the local pixel values to spread within the neighborhood, thereby ensuring the continuity of the overall structure. The reaction term -μ(x, y, T)I(x, y, v) is responsible for selectively suppressing certain features in the image, thus ensuring that the final image content can exhibit more distinct local features. In addition, the external driving term λ T T(x) is directly controlled by the local texture map T(x). Its main role is to ensure that the generated image content can reflect specific visual styles and structural layouts in the local area, making the final generated image not just a smooth evolution result, but able to exhibit visual features with specific styles. The definition of the diffusion coefficient κ(x, y, T) ensures that the diffusion process has different intensities in different regions, allowing the image content to expand more strongly in some regions while maintaining higher local details in other regions. Specifically, the definition form of this diffusion coefficient is where κ 0 is a preset value that controls the global intensity of diffusion, and the exponential term ensures that in regions with large texture gradients (i.e., regions with local texture mutations), the diffusion intensity is low, thereby retaining the detail information in these regions. The core role of this mechanism is to prevent the image from being over-smoothed during the evolution process, ensuring that the generated cultural and creative content can exhibit clear local features while maintaining good coherence in the smooth regions. Compared with traditional global diffusion methods, the innovation of this method lies in adaptively adjusting the diffusion coefficient through the texture gradient, enabling different regions of the image to be adjusted according to local features, and thus showing more flexible adaptability in different style generation tasks.

[0109] In addition, the definition of the decay rate μ(x, y, T) is also a key factor affecting the quality of the final image generation. This parameter controls the decay intensity of the image during the evolution process, ensuring that key features in the image are retained to a certain extent. Specifically, its definition is μ 0 (1 - exp(-||T(x, y)|| 2 ))), where μ 0is a global adjustment parameter, and the exponential term ensures that when the value of the local texture map is small, the decay rate approaches zero, allowing sufficient diffusion in this area. In areas with stronger local texture, the decay rate increases, suppressing the structural changes in the image to a certain extent and thus preserving key visual details. This decay control mechanism ensures that during the image generation process, the visual features of specific areas are not lost due to excessive diffusion, enabling the finally generated image to exhibit clear structural features in local details. The texture-driven weight λ T controls the influence degree of the local texture map on the evolution of the entire image. When λ T takes a larger value, the local texture map has a stronger influence on the finally generated image, enabling the image content to closely follow the structure of the local texture map. When λ T takes a smaller value, the image generation process is more affected by the diffusion term and the decay term, making the overall style softer. Therefore, by adjusting the value of λ T , the overall style of the image content can be flexibly controlled, enabling the finally generated cultural and creative image to not only exhibit stylized content with clear detail features but also maintain high stability in the overall structure. During the implementation process, the initial image I(x, y, 0) is set as a zero-mean noise image. After T 3 rounds of iteration, the final image content I init (x, y) = I(x, y, T 3 ) is obtained. This iterative process ensures that the image can be optimized based on the current state at each time step and finally converges to a stable style structure. Compared with the traditional neural network direct mapping method, the advantage of this method lies in its physical simulation-based partial differential equation framework, which makes the image generation process not only have strong mathematical interpretability but also provide higher controllability and self-adaptive adjustment ability. Traditional image generation methods usually rely on data-driven end-to-end training, while this method enables the image generation process to evolve under specific rules through explicit diffusion-reaction equations, thus ensuring a good balance between style consistency and detail expression ability in the generated results.

[0110] Although the specific implementation manners of the present invention are described above, those skilled in the art should understand that these specific implementation manners are only examples. Without departing from the principles and essence of the present invention, those skilled in the art can make various omissions, substitutions, and changes to the details of the above methods and systems. For example, combining the above method steps so as to perform substantially the same function in a substantially the same method to achieve substantially the same result belongs to the scope of the present invention. Therefore, the scope of the present invention is only defined by the appended claims.

Claims

1. A method for generating digital cultural creative content based on artificial intelligence, characterized in that: The method comprises: Step 1: Divide the input text description into several semantic segments according to predefined syntactic and semantic rules, calculate the local semantic score of each semantic segment using predefined word vector mapping and activation function, and calculate the global text embedding using the local semantic score of each semantic segment and the preset weight; Step 2: Add the global text embedding to the noise generated by the standard normal distribution to construct the initial latent variable; regard the initial latent variable as being on an implicit manifold, and construct a local loss function with the global text embedding as a reference, calculate the deviation between the initial latent variable and the text embedding through the preset Riemannian metric, and use the Riemannian gradient descent method to perform multi-step iterative updates on the initial latent variable to obtain the latent variable after manifold transformation; Step 3: Take the transformed latent variables as the initial state, construct a potential energy function including text semantic constraints and global regularization terms in the latent space, and realize the diffusion update of the latent variables through discretized stochastic differential equations; take the latent variables after diffusion update as input, determine the texture amplitude, center frequency and frequency diffusion parameters according to the preset frequency modulation function, and use the inverse Fourier transform method to generate a local texture map; Step 4: Use the generated local texture map as a local driving factor to construct a partial differential equation based on the diffusion and reaction mechanism. The partial differential equation evolves the image brightness distribution in time and space under preset initial conditions to achieve adaptive synthesis of the image structure layout and generate image content.

2. The method for generating digital cultural creative content based on artificial intelligence as claimed in claim 1, characterized in that: In step 1, let the input text description be T = {w1, w2, ..., w j , ..., w M }; where w j Represents the jth word, j is an integer subscript index ranging from 1 to M, and M is the total number of words; predefined word vector mapping The word w j Map to a real vector; Represents the vocabulary, which represents the set of all possible words; is the dimension of the word vector, indicating that word w j In the word vector space, feature values; map each word to the vector space to obtain the word vector φ(w j ); Using predefined syntactic and semantic rules, the text is decomposed into N semantic segments, the i-th semantic segment is is the jth word in the i-th semantic segment; word index n i,j Indicates the position of the word in the entire text; L i is the number of words in the i-th semantic segment; N is the total number of semantic segments after the text is divided.

3. The method for generating digital cultural creative content based on artificial intelligence as claimed in claim 2, characterized in that: In step 1, a semantic scoring function is defined for each semantic segment: Among them, ||·||1 represents the L1 norm operation; ψ(T i ) is the local semantic score; and the Sigmoid activation function is used to construct the fragment semantic vector: s i =σ(ψ(T i )-θ)v0; Among them, s i is the segment semantic vector of the i-th semantic segment; σ(·) is the Sigmoid activation function; θ is the activation threshold, which is a preset scalar used to translate ψ(T i ) output, thereby controlling the response interval of the Sigmoid activation function; v0 is the basic semantic direction vector, which is a The preset vector in d s is the dimension of the semantic vector, which determines the size of the feature space of global text embedding; global text embedding is obtained by weighted average of the semantic vectors of each segment: in, represents global text embedding; ||·||2 represents the L2 norm.

4. The method for generating digital cultural creative content based on artificial intelligence as claimed in claim 3, characterized in that: In step 2, the initial latent variables are constructed by the following process: for k = 1, ..., d s ,definition: U k ,V k ~Uniform(0,1); Among them, U k and V k is the kth noise component δ k The random variables used in the generation are independently sampled from the uniform distribution Uniform(0,1); all noise components form the noise vector Indicates d s dimensional unit matrix to ensure that there is no correlation between the noise dimensions; the noise vector is used to construct the initial latent variable ξ0: ξ0=e+λδ; Among them, λ adjusts the noise amplitude and its value range is 0.1 to 0.

5.

5. The method for generating digital cultural creative content based on artificial intelligence as claimed in claim 4, characterized in that: In step 2, the initial latent variable ξ0 is assumed to be distributed on the implicit manifold Its local geometry is described by a metric constructed based on the constraints of the global text embedding, defining the local loss function for:

6. The method for generating digital cultural creative content based on artificial intelligence as claimed in claim 5, characterized in that: In step 2, the deviation between the latent variable and the text embedding is calculated by the preset Riemannian metric, and the latent variable is updated in multiple steps by using the Riemannian gradient descent method. The process of obtaining the latent variable after manifold transformation includes: The default Riemann metric is g: Among them, ξ m is the mth component of ξ0; l is the lth component of ξ0; e m is the mth component of e; e l is the lth component of e; g ml is the element in the mth row and lth column of the Riemann metric g; The latent variables are updated using Riemannian gradient descent through the following formula: Among them, g -1 represents the inverse matrix of the Riemann metric; the latent variables after the manifold transformation are finally obtained as T1 is the total number of iterations; is the gradient operator; is the latent variable ξ obtained at the t-1th iteration t-1 The local loss function of ξ t is the latent variable obtained at the t-th iteration.

7. The method for generating digital cultural creative content based on artificial intelligence according to claim 6, characterized in that: In step 3, the obtained On this basis, a diffusion process is introduced to further balance text semantic constraints and generation exploration, and the following potential energy function with a global regularization term is constructed: Among them, δ1>0 is a preset value, providing a global contraction effect; the stochastic differential equation of the diffusion process is introduced: Where D>0, is the preset diffusion coefficient, W(t) is the standard Wiener process; let, define the discretization step length Δt, and the iteration formula is: in, After T1 iterations, the latent variables after diffusion update are finally obtained 8. The method for generating digital cultural creative content based on artificial intelligence as claimed in claim 7, characterized in that: In step 3, the frequency modulation function The formula is: Among them, ω is a frequency variable, belonging to the real number set a0 is the constant base value of amplitude modulation; a1 is the constant base value of amplitude modulation The relevant proportionality coefficient is a real constant; the center frequency ω base is the reference center frequency, which is a preset constant; <·> represents the inner product operation; b is the proportional coefficient; the frequency standard deviation σ0 is the benchmark standard deviation, which is a positive preset constant. The local texture map is obtained by inverse Fourier transform. Among them, x is the texture map position coordinate; Indicates real part operation; s is the imaginary unit; Ω x is the texture image space domain; x∈Ω x ; Phase function Where c0 and c1 are constants.

9. The method for generating digital cultural creative content based on artificial intelligence as claimed in claim 8, characterized in that: In step 4, the local texture map is used To generate the structural layout of a two-dimensional image; let the image be I(x, y, v), and its evolution is driven by the partial differential equation of the diffusion and reaction mechanism: The local diffusion coefficient is Among them, κ0>0 is the preset value; τ>0 is the preset value to ensure that the diffusion becomes weaker at the texture mutation point; v is the number of iterations; attenuation rate Among them, μ0>0 is the preset value; texture drive weight λ T >0, which is the preset value; the initial condition is selected as I(x, y, 0) = I0(x, y), which is a zero-mean noise image. After T3 iterations, the image content I is obtained. init (x, y) = I(x, y, T3).

Citation Information

Patent Citations

  • Image reconstruction and editing-based diffusion network fusing semantic enhancement clip

    CN117496289A

  • Single-exposure compression imaging method based on double-domain mean regression diffusion model

    CN118429448A

  • Text-to-image synthesis method based on efficient attention generative adversarial network

    CN119169153A

  • Multi-subject image generation method based on text description

    CN119274181A

  • Intelligent beauty feature optimization processing method and system for high-definition face changing

    CN119399080A

Cited By

  • Intelligent digital culture creative content generation method based on reinforcement learning

    CN121958579A