Image generation method and system based on style feature injection
By combining a multi-level style extraction model with a generative adversarial network, the problem of insufficient alignment between text semantics and image style is solved, and fine-grained style control and efficient image generation are achieved. The generated images perform excellently in visual effects and content accuracy.
Patent Information
- Application Number
- CN202510728023.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-03
AI Technical Summary
In existing technologies, the alignment of text semantics and image style in the feature space is insufficient, resulting in a mismatch between the generated image content and style, coarse style control granularity, and a lack of iterative optimization mechanism, making it difficult to meet the generation needs of complex scenes.
A pre-trained multi-level style extraction model is used to perform multi-dimensional feature analysis on the reference image, extracting color, texture, composition, and lighting feature vectors. The triplet information generated by the text description is weightedly fused and input into a generative adversarial network for multi-scale feature splicing. The generative adversarial network uses a deformable AdaIN module to inject style parameters and generate an image.
Fine-grained style control is achieved, and the generated images are highly matched with the text description, with content accuracy, style consistency and visual naturalness, making them suitable for high-precision art creation and scene modeling.
Smart Images

Figure CN120689466A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence image generation technology, and specifically relates to an image generation method and system based on style feature injection. Background Art
[0002] The following problems are common in existing technologies: 1. Insufficient cross-modal mapping: The alignment of text semantics and image style in the feature space is insufficient, resulting in a mismatch between the generated image content and style; 2. Coarse granularity of style control: Traditional methods only adjust the style through weight distribution or pixel fusion, and cannot achieve fine-grained control of multiple levels (such as color, texture, and composition); 3. Low generation robustness: The lack of an iterative optimization mechanism makes it difficult to meet the generation requirements of complex scenes.
[0003] Therefore, improvements are needed. Summary of the Invention
[0004] In order to solve the above technical problems, the present application provides an image generation method and system based on style feature injection.
[0005] The first object of the invention of this application is achieved through the following technical solutions:
[0006] An image generation method based on style feature injection, comprising:
[0007] Through the pre-trained multi-level style extraction model, multi-dimensional feature analysis is performed on the reference image specified by the user to extract the style feature vector, which includes the color feature vector, texture feature vector, composition feature vector, and illumination feature vector;
[0008] Obtain the text description input by the user and output triple information including scene entities, entity attributes and spatial relationships;
[0009] Performing weighted fusion on the triplet information and the style feature vector to obtain a style-enhanced semantic embedding vector;
[0010] The style-enhanced semantic embedding vector is input into a generative adversarial network for splicing to output an image.
[0011] In a preferred embodiment, the steps of performing multi-dimensional feature analysis on a reference image specified by a user terminal through a pre-trained multi-level style extraction model to extract a style feature vector, wherein the style feature vector includes a color feature vector, a texture feature phasor, a composition feature vector, and an illumination feature vector, include:
[0012] The pre-trained multi-level style extraction model includes a color feature extraction layer, a texture feature extraction layer, a composition feature extraction layer, and an illumination feature extraction layer;
[0013] The color feature extraction layer:
[0014] Performing color space conversion on the reference image, converting from RGB color space to HSV color space or Lab color space;
[0015] Construct a pixel distribution histogram to count the pixel distribution of each color channel in the image;
[0016] Calculate color feature parameters based on the pixel distribution histogram, wherein the color feature parameters include main color tone, color contrast and color gradient law;
[0017] Encoding the color feature parameters into a color feature vector, and outputting the color feature vector;
[0018] The texture feature extraction layer:
[0019] Designing a Gabor filter bank for the reference image, wherein the Gabor filter bank includes multiple Gabor filters of different directions and scales;
[0020] Performing a convolution operation on the reference image and a Gabor filter bank to obtain texture response maps in multiple directions;
[0021] Post-processing the texture response map, wherein the post-processing includes normalization and binarization operations;
[0022] Counting statistical characteristics of the texture response map and outputting texture feature parameters, wherein the texture feature parameters include mean, variance, and energy;
[0023] The texture feature parameters are encoded into a texture feature vector, and the texture feature vector is output.
[0024] In a preferred embodiment, the steps of performing multi-dimensional feature analysis on a reference image specified by a user terminal through a pre-trained multi-level style extraction model to extract a style feature vector, wherein the style feature vector includes a color feature vector, a texture feature phasor, a composition feature vector, and an illumination feature vector, further include:
[0025] The composition feature extraction layer:
[0026] Use a deep learning segmentation network to perform semantic segmentation on the reference image to obtain the various subject areas in the picture;
[0027] Performing morphological processing on the main area, wherein the morphological processing includes corrosion, expansion, opening operation, and closing operation;
[0028] Calculating the position information of the subject area, the position information including the center coordinates, the bounding box, and the relative position relationship between the areas;
[0029] Based on the position information of the main area, identify visual guide lines and balance features in the picture, the visual guide lines include diagonal lines and thirds lines, and the balance features include symmetry and center of gravity position;
[0030] Encoding the position information, visual guide lines, and balance features into a composition feature vector, and outputting the composition feature vector;
[0031] The illumination feature extraction layer:
[0032] Use the ambient light estimation model to analyze the reference image and estimate the light source direction, light source intensity and light source type;
[0033] Detecting a shadow area of the reference image, wherein the shadow area includes a position, a shape, and a projection direction of the shadow;
[0034] Identifying the highlight area of the reference image and calculating the brightness value and distribution range of the highlight area;
[0035] The light source direction, light source intensity, shadow area and highlight area are encoded into a lighting feature vector, and the lighting feature vector is output.
[0036] In a preferred embodiment, the step of obtaining a text description input by a user and outputting triplet information including scene entities, entity attributes, and spatial relationships includes:
[0037] Preprocessing the text description, wherein the preprocessing includes removing stop words and punctuation marks, and performing word segmentation and part-of-speech tagging;
[0038] Based on natural language processing technology, the preprocessed text description is parsed, and the parsing includes:
[0039] Identify and extract scene entities, including people, objects, and background elements;
[0040] Extracting entity attributes of each scene entity, wherein the entity attributes include color, shape, size, and material;
[0041] Identify and extract spatial relationships between scene entities, including position relationships, direction relationships, and distance relationships;
[0042] The extracted scene entities, entity attributes and spatial relationships are structured and represented according to a preset format to generate triple information.
[0043] In a preferred embodiment, the step of performing weighted fusion of the triple information and the style feature vector to obtain a style-enhanced semantic embedding vector includes:
[0044] The triple information is input into the graph neural network, and a semantic vector E with a dimension of d is generated through node embedding and relationship aggregation. semantic ;
[0045] The style feature vector is mapped to the d-dimensional space through the fully connected layer, and the style vector E is obtained after splicing. style ;
[0046] Calculate the cross attention weight of the semantic vector and the style vector:
[0047] Among them, W Q ,W K is the learnable parameter matrix;
[0048] Based on the cross attention weight, the style vector is weighted: E style\_attn =A·E style ;
[0049] Introducing a learnable dynamic weight parameter α, the contribution ratio of the style vector to the semantic vector is adjusted through the gating mechanism: α = σ(W g [E semantic ;E style ]), where σ is the Sigmoid function, W g is the gating weight matrix;
[0050] Weighted fusion generates style-enhanced semantic embedding vectors: E fused =α·E style\_attn +(1-α)·E semantic .
[0051] In a preferred embodiment, the step of inputting the style-enhanced semantic embedding vector into a generative adversarial network, performing splicing, and outputting an image includes:
[0052] The generative adversarial network includes a generator;
[0053] The generator includes an initial convolutional layer, a residual block sequence, an upsampling layer and an output layer;
[0054] Inputting the style-enhanced semantic embedding vector into the initial convolutional layer to generate a low-resolution feature map, wherein the low-resolution feature map contains high-frequency details, and the high-frequency details include edge features, local textures, and noise patterns;
[0055] AdaIN is inserted after the initial convolutional layer to perform style modulation on the low-resolution feature map and output a shallow feature map.
[0056] In a preferred embodiment, the step of inputting the style-enhanced semantic embedding vector into a generative adversarial network, performing splicing, and outputting an image further includes:
[0057] Inputting the shallow feature map into the residual block sequence, improving the resolution through an upsampling layer, and outputting a high-resolution feature map, wherein the high-resolution feature map contains global semantics, including object contours, spatial layout, and illumination distribution;
[0058] Insert an AdaIN layer after each upsampling layer to re-inject the style of the high-resolution feature map and output a deep feature map;
[0059] Upsampling the shallow feature map to the target resolution through transposed convolution, and concatenating it with the deep feature map in the channel dimension to generate a fused feature map;
[0060] The fused feature map is channel compressed and an RGB image is output.
[0061] The second object of the invention of this application is achieved through the following technical solutions:
[0062] An image generation system based on style feature injection, comprising:
[0063] Extraction module: Uses a pre-trained multi-level style extraction model to perform multi-dimensional feature analysis on the reference image specified by the user and extract style feature vectors. The style feature vectors include color feature vectors, texture feature vectors, composition feature vectors, and illumination feature vectors.
[0064] Acquisition module: obtains the text description input by the user and outputs triple information including scene entities, entity attributes and spatial relationships;
[0065] Fusion module: performs weighted fusion of the triple information and the style feature vector to obtain a style-enhanced semantic embedding vector;
[0066] Stitching module: Input the style-enhanced semantic embedding vector into the generative adversarial network, perform stitching, and output the image.
[0067] The third object of the invention of this application is achieved through the following technical solutions:
[0068] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned image generation method based on style feature injection are implemented.
[0069] The fourth object of the invention of this application is achieved through the following technical solutions:
[0070] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above-mentioned image generation method based on style feature injection.
[0071] In summary, this application includes at least one of the following beneficial technical effects:
[0072] A decoupled feature analysis of the reference image is performed through a multi-level style extraction model, and style vectors are extracted from four dimensions: color, texture, composition, and lighting to achieve fine-grained style control. Combined with the structured triples generated by the text parsing module, the triplet information and the style feature vector are weightedly fused to obtain a style-enhanced semantic embedding vector. The style-enhanced semantic embedding vector is input into a generative adversarial network. The generative adversarial network adopts a multi-scale residual architecture and injects style parameters between multiple resolution levels through a deformable AdaIN module to generate shallow feature maps and deep feature maps. The shallow feature maps and deep feature maps are then spliced to output an RGB image. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 This is a flowchart of an implementation of an embodiment of an image generation method based on style feature injection in the present application;
[0074] Figure 2 This is a flowchart for implementing step S20 in an embodiment of an image generation method based on style feature injection in the present application;
[0075] Figure 3 This is a flowchart for implementing step S40 in an embodiment of an image generation method based on style feature injection in the present application;
[0076] Figure 4 This is another implementation flowchart of step S40 in an embodiment of an image generation method based on style feature injection of the present application;
[0077] Figure 5 This is a principle block diagram of a computer device of the present application. DETAILED DESCRIPTION
[0078] The following is combined with Figure 1-5 This application is described in further detail.
[0079] In one embodiment, if Figure 1 As shown, the present application discloses an image generation method based on style feature injection, which specifically includes the following steps:
[0080] S10: Using a pre-trained multi-level style extraction model, perform multi-dimensional feature analysis on the reference image specified by the user to extract a style feature vector, wherein the style feature vector includes a color feature vector, a texture feature vector, a composition feature vector, and an illumination feature vector;
[0081] S20: Obtain the text description input by the user terminal and output triple information including scene entities, entity attributes and spatial relationships;
[0082] S30: performing weighted fusion on the triplet information and the style feature vector to obtain a style-enhanced semantic embedding vector;
[0083] S40: Input the style-enhanced semantic embedding vector into a generative adversarial network for splicing and outputting an image.
[0084] In this embodiment, a decoupled feature analysis is performed on the reference image through a multi-level style extraction model, and style vectors are extracted from four dimensions: color, texture, composition, and lighting to achieve fine-grained style control. Combined with the structured triples generated by the text parsing module, the triplet information and the style feature vector are weightedly fused to obtain a style-enhanced semantic embedding vector. The style-enhanced semantic embedding vector is input into a generative adversarial network. The generative adversarial network adopts a multi-scale residual architecture and injects style parameters between multiple resolution levels through a deformable AdaIN module to generate shallow feature maps and deep feature maps. The shallow feature maps and deep feature maps are then spliced to output an RGB image.
[0085] Step S10 includes:
[0086] SB1: wherein the pre-trained multi-level style extraction model includes a color feature extraction layer, a texture feature extraction layer, a composition feature extraction layer, and an illumination feature extraction layer;
[0087] SB2: the color feature extraction layer:
[0088] Performing color space conversion on the reference image, converting from RGB color space to HSV color space or Lab color space;
[0089] SB3: Construct a pixel distribution histogram to count the pixel distribution of each color channel in the image;
[0090] SB4: Calculate color feature parameters based on the pixel distribution histogram, the color feature parameters including main color tone, color contrast and color gradient rule;
[0091] SB5: Encode the color feature parameters into a color feature vector, and output the color feature vector;
[0092] SB6: Texture feature extraction layer:
[0093] Designing a Gabor filter bank for the reference image, wherein the Gabor filter bank includes multiple Gabor filters of different directions and scales;
[0094] SB7: performing a convolution operation on the reference image and a Gabor filter bank to obtain texture response maps in multiple directions;
[0095] SB8: performing post-processing on the texture response map, wherein the post-processing includes normalization and binarization operations;
[0096] SB9: Calculate the statistical characteristics of the texture response map and output texture feature parameters, which include mean, variance, and energy;
[0097] SB10: Encode the texture feature parameters into a texture feature vector and output the texture feature vector.
[0098] In this embodiment, a pre-trained multi-level style extraction model is used to perform fine-grained feature analysis on the reference image. The color feature extraction layer (SB1-SB5) converts the image from RGB space to Lab / HSV space, constructs a pixel distribution histogram to quantify the main color tone, color contrast and gradient pattern, and encodes it into a color feature vector to accurately capture the color style. The texture feature extraction layer (SB6-SB10) generates a texture response map through convolution of a multi-directional and multi-scale Gabor filter group, which is encoded into a texture feature vector after normalization and statistical characteristics (mean, variance, energy) analysis, effectively characterizing the complexity and directionality of the local texture.
[0099] Step S10 further includes:
[0100] SH1: the composition feature extraction layer:
[0101] Use a deep learning segmentation network to perform semantic segmentation on the reference image to obtain the various subject areas in the picture;
[0102] SH2: performing morphological processing on the main area, wherein the morphological processing includes corrosion, expansion, opening operation, and closing operation;
[0103] SH3: Calculate the position information of the subject area, the position information including the center coordinates, the bounding box, and the relative position relationship between the areas;
[0104] SH4: Based on the position information of the main body area, identify visual guide lines and balance features in the picture, the visual guide lines include diagonal lines and thirds lines, and the balance features include symmetry and center of gravity position;
[0105] SH5: Encode the position information, visual guide lines, and balance features into a composition feature vector, and output the composition feature vector;
[0106] SH6: the illumination feature extraction layer:
[0107] Use the ambient light estimation model to perform illumination analysis on the reference image to estimate the light source direction, light source intensity and light source type;
[0108] SH7: Detecting a shadow area of the reference image, wherein the shadow area includes a position, a shape, and a projection direction of the shadow;
[0109] SH8: Identify the highlight area of the reference image and calculate the brightness value and distribution range of the highlight area;
[0110] SH9: Encode the light source direction, light source intensity, shadow area and highlight area into a lighting feature vector, and output the lighting feature vector.
[0111] In this embodiment, a pre-trained multi-level style extraction model is used to perform structured analysis on the reference image. In the composition feature extraction layer, a deep learning segmentation network is used to perform semantic segmentation on the main area of the picture. Morphological processing (such as corrosion and dilation) is combined to optimize the region boundary. By calculating the center coordinates, bounding box and relative position relationship of the main area, visual guide lines (such as thirds and diagonals) and balance features (such as symmetry index and center of gravity offset) are identified and encoded into composition feature vectors to accurately represent the spatial layout rules of the picture. In the lighting feature extraction layer, the direction and intensity of the light source are analyzed based on the ambient light estimation model. By detecting the geometric shape of the shadow area and calculating the brightness distribution of the highlight area, parameters such as the shadow projection direction and highlight coverage range are extracted and encoded into lighting feature vectors to dynamically restore the real lighting scene.
[0112] Figure 2 , step S20 includes:
[0113] S201: Preprocessing the text description, wherein the preprocessing includes removing stop words and punctuation marks, and performing word segmentation and part-of-speech tagging;
[0114] S202: Parsing the pre-processed text description based on natural language processing technology, wherein the parsing includes:
[0115] S203: Identify and extract scene entities, where the scene entities include people, objects, and background elements;
[0116] S204: extracting entity attributes of each scene entity, wherein the entity attributes include color, shape, size, and material;
[0117] S205: Identify and extract spatial relationships between scene entities, where the spatial relationships include position relationships, direction relationships, and distance relationships;
[0118] S206: The extracted scene entities, entity attributes and spatial relationships are structured according to a preset format to generate triple information.
[0119] In this embodiment, the input text is first preprocessed (S201), with stop words and punctuation removed, followed by word segmentation and part-of-speech tagging, providing standardized input for semantic parsing. Based on dependency parsing and named entity recognition techniques, scene entities (such as people, objects, and background elements) and their attributes (including physical properties such as color, shape, and material) are accurately extracted from the text. A spatial relationship reasoning model (such as a rule-based or deep learning location parser) is then used to identify the relative position, direction, and distance relationships between entities (S202-S205), ultimately generating structured triple information (S206). This process converts free text into a machine-parseable semantic framework. By explicitly defining entity attributes and spatial constraints, it effectively eliminates ambiguity in natural language descriptions, providing a precise semantic control foundation for image generation. After combining with the style feature vector, the generative model can accurately restore the visual attributes and spatial layout of scene entities based on structured triples, ensuring that the generated image is highly matched with the text description in terms of object details (such as material texture), spatial logic (such as relative position consistency) and global composition (such as entity proportion relationship), significantly improving the content accuracy and scene rationality of the generated results.
[0120] Step S30 includes:
[0121] S301: Input the triple information into the graph neural network, and generate a semantic vector E with a dimension of d through node embedding and relationship aggregation semantic ;
[0122] S302: Map the style feature vector to the d-dimensional space through the fully connected layer, and obtain the style vector E after splicing style ;
[0123] S303: Calculate the cross attention weight of the semantic vector and the style vector:
[0124] Among them, W Q ,W K is the learnable parameter matrix;
[0125] S304: Based on the cross attention weight, weight the style vector: E style\_attn =A·E style ;
[0126] S305: Introduce a learnable dynamic weight parameter α and adjust the contribution ratio of the style vector and the semantic vector through the gating mechanism: α = σ(W g [E semantic ;E style ]), where σ is the Sigmoid function, W g is the gating weight matrix;
[0127] S306: Weighted fusion generates style-enhanced semantic embedding vector: E fused =α·E style\_attn +(1-α)·E semantic .
[0128] In this embodiment, efficient alignment of semantic and style features is achieved through a cross-modal fusion mechanism (S30). First, structured triples are input into a graph neural network, entity semantics are captured through node embedding, and spatial constraints are integrated based on the relationship aggregation layer to generate a dimensionally aligned semantic vector. At the same time, multi-dimensional style features (color, texture, composition, and lighting) are mapped to a unified latent space through a fully connected layer to form a style vector. The association weights between the semantic vector and the style vector are calculated through a cross-attention mechanism, key style elements are dynamically identified (for example, the "material" entity is preferentially associated with texture features), and the style vector is weighted. Further, a learnable gated weight parameter is introduced, and a Sigmoid function is used to adaptively adjust the fusion ratio of style and semantics (for example, increasing the style weight in complex scenes to enhance artistic expression). The resulting style-enhanced semantic embedding vector not only retains the precise semantic constraints of the text description (for example, reducing the object position error to pixel-level accuracy), but also deeply integrates the style characteristics of the reference image, enabling the generative adversarial network to output images with both content accuracy, style consistency, and visual naturalness.
[0129] Figure 3 , step S40 includes:
[0130] S401: The generative adversarial network includes a generator;
[0131] S402: The generator includes an initial convolutional layer, a residual block sequence, an upsampling layer and an output layer;
[0132] S403: Inputting the style-enhanced semantic embedding vector into the initial convolutional layer to generate a low-resolution feature map, where the low-resolution feature map contains high-frequency details, including edge features, local textures, and noise patterns;
[0133] S404: Insert AdaIN after the initial convolutional layer, perform style modulation on the low-resolution feature map, and output a shallow feature map.
[0134] In this embodiment, multi-level image synthesis is achieved through the generator of a generative adversarial network (S401-S402). The core of this process is the combination of style modulation and cross-level feature fusion. The style-enhanced semantic embedding vector is first input into the initial convolutional layer to generate a low-resolution feature map (S403), capturing high-frequency details (such as edge sharpness, local texture, and noise perturbations). The AdaIN layer (S404) dynamically adjusts the feature distribution to align shallow-level features with the reference style.
[0135] Figure 4 , step S40 further includes:
[0136] S405: Inputting the shallow feature map into the residual block sequence, improving the resolution through the upsampling layer, and outputting a high-resolution feature map, wherein the high-resolution feature map contains global semantics, and the global semantics include object contours, spatial layout, and illumination distribution;
[0137] S406: inserting an AdaIN layer after each upsampling layer, re-injecting style into the high-resolution feature map, and outputting a deep feature map;
[0138] S407: Upsampling the shallow feature map to the target resolution through transposed convolution, and performing channel-wise splicing with the deep feature map to generate a fused feature map;
[0139] S408: Perform channel compression on the fused feature map and output an RGB image.
[0140] In this embodiment, the shallow features are upsampled step by step through a sequence of residual blocks (S405), and an AdaIN layer is inserted after each level (S406) to re-inject the style of the high-resolution feature map, gradually constructing global semantics (such as the geometric accuracy of the object outline, the topological relationship of the spatial layout, and the natural transition of the illumination distribution). Finally, the shallow features are upsampled to the target resolution through transposed convolution and spliced with the deep features in the channel dimension (S407), fusing high-frequency details (such as the fine depiction of fabric texture) and global structure (such as the reasonable arrangement of scene objects), and then outputting the RGB image through channel compression (S408). This method ensures global style consistency while retaining local authenticity through multi-level style modulation and cross-scale feature fusion, significantly improving the problems of blurred details and style fragmentation in traditional methods, and generating images that are visually both delicate and harmonious, suitable for high-precision artistic creation and scene modeling needs.
[0141] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0142] In one embodiment, a system for generating an image based on style feature injection is provided. The system corresponds to the method for generating an image based on style feature injection in the above embodiment. The system includes:
[0143] For the specific definition of an image generation system based on style feature injection, please refer to the definition of an image generation method based on style feature injection above, and will not be repeated here. The various modules in the above-mentioned image generation system based on style feature injection can be implemented in whole or in part through software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of the above modules.
[0144] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store shallow feature maps and deep feature maps. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for generating an image based on style feature injection is implemented.
[0145] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a method for generating an image based on style feature injection is implemented.
[0146] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, a method for generating an image based on style feature injection is provided.
[0147] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0148] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
Claims
1. An image generation method based on style feature injection, characterized in that: include: Through the pre-trained multi-level style extraction model, multi-dimensional feature analysis is performed on the reference image specified by the user to extract the style feature vector, which includes the color feature vector, texture feature vector, composition feature vector, and illumination feature vector; Obtain the text description input by the user and output triple information including scene entities, entity attributes and spatial relationships; Performing weighted fusion on the triplet information and the style feature vector to obtain a style-enhanced semantic embedding vector; The style-enhanced semantic embedding vector is input into a generative adversarial network for splicing to output an image.
2. The image generation method based on style feature injection according to claim 1, characterized in that: The pre-trained multi-level style extraction model is used to perform multi-dimensional feature analysis on the reference image specified by the user end to extract the style feature vector, which includes a color feature vector, a texture feature vector, a composition feature vector, and an illumination feature vector. include: The pre-trained multi-level style extraction model includes a color feature extraction layer, a texture feature extraction layer, a composition feature extraction layer, and an illumination feature extraction layer; The color feature extraction layer: Performing color space conversion on the reference image, converting from RGB color space to HSV color space or Lab color space; Construct a pixel distribution histogram to count the pixel distribution of each color channel in the image; Calculate color characteristic parameters based on the pixel distribution histogram, wherein the color characteristic parameters include main color tone, color contrast and color gradient law; Encoding the color feature parameters into a color feature vector, and outputting the color feature vector; The texture feature extraction layer: Designing a Gabor filter bank for the reference image, wherein the Gabor filter bank includes multiple Gabor filters of different directions and scales; Performing a convolution operation on the reference image and a Gabor filter bank to obtain texture response maps in multiple directions; performing post-processing on the texture response map, wherein the post-processing includes normalization and binarization operations; Counting the statistical characteristics of the texture response map and outputting texture feature parameters, wherein the texture feature parameters include mean, variance, and energy; The texture feature parameters are encoded into a texture feature vector, and the texture feature vector is output.
3. The image generation method based on style feature injection according to claim 2, characterized in that: The method further includes: performing multi-dimensional feature analysis on a reference image specified by a user terminal through a pre-trained multi-level style extraction model to extract a style feature vector, wherein the style feature vector includes a color feature vector, a texture feature phasor, a composition feature vector, and an illumination feature vector; The composition feature extraction layer: Use a deep learning segmentation network to perform semantic segmentation on the reference image to obtain the various subject areas in the picture; Performing morphological processing on the main area, wherein the morphological processing includes corrosion, expansion, opening operation, and closing operation; Calculating the position information of the subject area, the position information including the center coordinates, the bounding box, and the relative position relationship between the areas; Based on the position information of the main area, identify visual guide lines and balance features in the picture, the visual guide lines include diagonal lines and thirds lines, and the balance features include symmetry and center of gravity position; Encoding the position information, visual guide lines, and balance features into a composition feature vector, and outputting the composition feature vector; The illumination feature extraction layer: Use the ambient light estimation model to analyze the reference image and estimate the light source direction, light source intensity and light source type; Detecting a shadow area of the reference image, wherein the shadow area includes a position, a shape, and a projection direction of the shadow; Identifying the highlight area of the reference image and calculating the brightness value and distribution range of the highlight area; The light source direction, light source intensity, shadow area and highlight area are encoded into a lighting feature vector, and the lighting feature vector is output.
4. The image generation method based on style feature injection according to claim 1, characterized in that: The step of obtaining a text description input by a user terminal and outputting triple information including scene entities, entity attributes, and spatial relationships includes: Preprocessing the text description, wherein the preprocessing includes removing stop words and punctuation marks, and performing word segmentation and part-of-speech tagging; Based on natural language processing technology, the preprocessed text description is parsed, and the parsing includes: Identify and extract scene entities, including people, objects, and background elements; Extracting entity attributes of each scene entity, wherein the entity attributes include color, shape, size, and material; Identify and extract spatial relationships between scene entities, including position relationships, direction relationships, and distance relationships; The extracted scene entities, entity attributes and spatial relationships are structured and represented according to a preset format to generate triple information.
5. The image generation method based on style feature injection according to claim 1, characterized in that: The step of performing weighted fusion of the triple information and the style feature vector to obtain a style-enhanced semantic embedding vector includes: The triple information is input into the graph neural network, and a semantic vector E with a dimension of d is generated through node embedding and relationship aggregation. semantic ; The style feature vector is mapped to the d-dimensional space through the fully connected layer, and the style vector E is obtained after splicing. style ; Calculate the cross attention weight of the semantic vector and the style vector: Among them, W Q ,W K is the learnable parameter matrix; Based on the cross attention weight, the style vector is weighted: E style\_attn =A·E style ; Introducing a learnable dynamic weight parameter α, the contribution ratio of the style vector to the semantic vector is adjusted through the gating mechanism: α = σ(W g [E semantic ;E style ]), where σ is the Sigmoid function, W g is the gating weight matrix; Weighted fusion generates style-enhanced semantic embedding vectors: E fused =α·E style\_attn +(1-α)·E semantic .
6. The image generation method based on style feature injection according to claim 1, characterized in that: The step of inputting the style-enhanced semantic embedding vector into a generative adversarial network, performing splicing, and outputting an image includes: The generative adversarial network includes a generator; The generator includes an initial convolutional layer, a residual block sequence, an upsampling layer and an output layer; Inputting the style-enhanced semantic embedding vector into the initial convolutional layer to generate a low-resolution feature map, wherein the low-resolution feature map contains high-frequency details, and the high-frequency details include edge features, local textures, and noise patterns; AdaIN is inserted after the initial convolutional layer to perform style modulation on the low-resolution feature map and output a shallow feature map.
7. The image generation method based on style feature injection according to claim 6, characterized in that: The step of inputting the style-enhanced semantic embedding vector into a generative adversarial network, performing splicing, and outputting an image further includes: Inputting the shallow feature map into the residual block sequence, improving the resolution through an upsampling layer, and outputting a high-resolution feature map, wherein the high-resolution feature map contains global semantics, including object contours, spatial layout, and illumination distribution; Insert an AdaIN layer after each upsampling layer to re-inject the style of the high-resolution feature map and output a deep feature map; Upsampling the shallow feature map to the target resolution through transposed convolution, and concatenating it with the deep feature map in the channel dimension to generate a fused feature map; The fused feature map is channel compressed and an RGB image is output.
8. An image generation system based on style feature injection, characterized in that: include: Extraction module: Uses a pre-trained multi-level style extraction model to perform multi-dimensional feature analysis on the reference image specified by the user and extract style feature vectors. The style feature vectors include color feature vectors, texture feature vectors, composition feature vectors, and illumination feature vectors. Acquisition module: obtains the text description input by the user and outputs triple information including scene entities, entity attributes and spatial relationships; Fusion module: performs weighted fusion of the triple information and the style feature vector to obtain a style-enhanced semantic embedding vector; Stitching module: Input the style-enhanced semantic embedding vector into the generative adversarial network, perform stitching, and output the image.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the image generation method based on style feature injection as described in claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the image generation method based on style feature injection according to claims 1 to 7.
Citation Information
Patent Citations
Stylized image generation method and device, computer equipment and storage medium
CN116012488A
Image generation method and device, computer equipment and storage medium
CN116363242A
Digital art creation auxiliary system
CN117788670A
Image generation method and device, equipment, storage medium and program product
CN118015111A
Image generation method and device, equipment, storage medium and program product
CN118035493A
Cited By
Game image generation system and method based on artificial intelligence enabling
CN121366215A
An artificial intelligence enabled game image generation system and method
CN121366215B
Digitized full-automatic wall painting repairing method fused with feature extractor
CN121544496A
A full-automatic mural digital restoration method of fusing feature extractors
CN121544496B
Artificial intelligence-fused creative visual style generation system and method
CN121600120A