An AI image generation method and system based on image analysis and language description

By combining language description and image analysis, correlation relationship data is generated and layout optimization is performed, the problem that it is difficult to understand semantic association and spatial layout logic in image generation in the prior art is solved, and high-precision image generation is achieved.

CN120014096BActive Publication Date: 2025-06-27HUAYI TIMES TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510496710.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-06-27
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing image generation methods based on image analysis and language description have limitations in understanding image semantic associations and spatial layout logic, and it is difficult to generate object placement relationships that conform to physical laws.

Method used

By obtaining the subject object, spatial orientation predicate and object object in the language description, the first description vector is generated, and a core element set is generated through the correlation analysis model clustering. Combining the visual characteristics of the reference image, the association data is generated through cross-modal matching, the object collection is reorganized, and layout optimization is used to ensure that the generated image conforms to physical laws.

Benefits of technology

The accuracy of target image generation is improved, so that the generated image meets both semantic accuracy and visual rationality, and avoids the problems of missing details in pure text generation and spatial logic confusion in pure image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014096B_ABST
    Figure CN120014096B_ABST
Patent Text Reader

Abstract

The present invention discloses an AI image generation method and system based on image analysis and language description, belonging to the technical field of image generation, which can improve the accuracy of target image generation. The method includes obtaining the language description of the target scene to be generated, identifying and extracting the subject object, spatial orientation predicate, and object object in the language description, and generating a corresponding first description vector; clustering the first description vector to generate a core element set; generating a second description vector corresponding to each object according to the obtained reference image; matching the second description vector with the core element set to generate associated relationship data; reorganizing each object based on the associated relationship data to generate a set of reorganized objects; optimizing the layout of the set of reorganized objects based on the spatial relationship between each object and the association degree between core elements to obtain the layout information of the set of reorganized objects; and generating a target image according to the layout information of the set of reorganized objects and the first description vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image generation, and specifically to an AI image generation method and system based on image analysis and language description. Background Art

[0002] With the rapid development of generative artificial intelligence technology, image generation methods based on deep learning have been widely applied in fields such as creative design and virtual scene construction. Traditional image generation technologies are mainly divided into two major technical routes: generation methods based on image content analysis and generation methods based on language description.

[0003] In the generation method based on image analysis, researchers usually use generative adversarial networks (GANs) or diffusion models (Diffusion Models) to extract and reconstruct features of the input image. Such methods can effectively capture low-level features such as the texture and color of visual elements, but have obvious limitations in understanding the semantic associations and spatial layout logic of images. For example, in the scene reconstruction task, existing methods are difficult to infer the physical-law-compliant object placement relationships from a single reference image.

[0004] The generation method based on language description maps text semantics to the latent space through cross-modal models such as CLIP, and uses architectures such as VQ-VAE to achieve text-to-image conversion. Although such methods perform outstandingly in responding to abstract semantic requirements, they have problems such as uncontrollable visual details and ambiguous spatial relationships. When faced with descriptions containing clear spatial indications such as "a red sofa is placed in the front left and a square decorative painting is hung on the right wall", the generated results often show object position offsets or attribute mismatches.

[0005] The disclosure of the above background art content is only used to assist in understanding the concept and technical solution of the present invention, and it does not necessarily belong to the prior art of this patent application. Without clear evidence indicating that the above content was publicly available on the filing date of this patent application, the above background art should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention

[0006] This application provides an AI image generation method and system based on image analysis and language description, which can improve the accuracy of target image generation.

[0007] To achieve the above object, the embodiments of this application disclose the following technical solutions:

[0008] In the first aspect, the embodiments of this application provide an AI image generation method based on image analysis and language description, including the following steps:

[0009] Obtain the language description of the target scene to be generated, and input the language description into a pre-trained language model to identify and extract the subject object, spatial orientation predicate, and object object in the language description, and generate a corresponding first description vector;

[0010] Cluster the first description vector through a pre-trained association analysis model to generate a set of core elements;

[0011] Obtain a reference image and input it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object;

[0012] Match the second description vector with the set of core elements, calculate the similarity between each second description vector and each core element in the set of core elements, and confirm the association relationship between each object and multiple core elements based on the similarity to generate association relationship data;

[0013] Recombine each object based on the association relationship data to generate a set of recombined objects;

[0014] Optimize the layout of the set of recombined objects based on the spatial relationship between each object and the association degree between core elements to obtain the layout information of the set of recombined objects;

[0015] Generate a target image according to the layout information of the set of recombined objects and the first description vector.

[0016] In the embodiments of the present application, high-level semantic constraints (such as subject objects, object objects, and spatial orientation relationships) are provided through language descriptions, and low-level visual features (such as textures and shapes) are provided through reference images. The two are aligned through cross-modal matching, so that the generated image simultaneously satisfies semantic accuracy and visual rationality, effectively avoiding the lack of details in pure text generation and the problem of chaotic spatial logic in pure image generation. In this way, the accuracy of target image generation can be improved.

[0017] In addition, gradient optimization of the object positions based on the association relationship data can ensure that the generated image conforms to physical laws (such as avoiding object collisions).

[0018] In a second aspect, the embodiments of the present application provide an AI image generation system based on image analysis and language description, including:

[0019] A first generation module, configured to obtain the language description of the target scene to be generated, and input the language description into a pre-trained language model to identify and extract the subject object, spatial orientation predicate, and object object in the language description, and generate a corresponding first description vector;

[0020] A second generation module, configured to cluster the first description vector through a pre-trained association analysis model to generate a set of core elements;

[0021] A third generation module, configured to obtain a reference image and input it into a pre-trained image feature extraction model to generate second description vectors corresponding to respective objects;

[0022] A fourth generation module, configured to match the second description vectors with the core element set, calculate the similarities between the second description vectors and the respective core elements in the core element set, confirm the association relationships between the respective objects and the multiple core elements based on the similarities, and generate association relationship data;

[0023] A first recombination module, configured to recombine the respective objects based on the association relationship data to generate a set of recombined objects;

[0024] A first optimization module, configured to perform layout optimization on the set of recombined objects based on the spatial relationships between the respective objects and the association degrees between the core elements to obtain layout information of the set of recombined objects;

[0025] A fifth generation module, configured to generate a target image according to the layout information of the set of recombined objects and the first description vectors.

[0026] In a third aspect, an embodiment of the present application provides an electronic device, including one or more processors; a storage device storing one or more programs thereon; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any one of the technical solutions of the first aspect.

[0027] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, having a computer program stored thereon, and when the computer program is executed by a processor, the method described in any one of the technical solutions of the first aspect is implemented.

[0028] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in any one of the technical solutions of the first aspect is implemented.

[0029] Wherein, for the technical effects brought by any one of the design manners in the second aspect to the fifth aspect, reference may be made to the technical effects brought by different design manners in the first aspect, which will not be elaborated herein. Description of the Drawings

[0030] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained according to the provided drawings.

[0031] Figure 1 A schematic flowchart of an AI image generation method based on image analysis and language description provided for some embodiments of this application;

[0032] Figure 2 A schematic structural diagram of an AI image generation method based on image analysis and language description provided for some embodiments of this application;

[0033] Figure 3 A schematic structural diagram of an electronic device suitable for implementing some embodiments of this application. Detailed implementation manners

[0034] Now, specific embodiments of the present invention will be described in detail. Although the present invention is described in conjunction with these specific embodiments, it should be understood that it is not intended to limit the present invention to these specific embodiments. On the contrary, these embodiments are intended to cover alternatives, modifications, or equivalent embodiments that may be included within the spirit and scope of the invention defined by the claims. In the following description, numerous specific details are set forth in order to provide a comprehensive understanding of the present invention. The present invention may be practiced without some or all of these specific details.

[0035] When used in conjunction with the terms "comprising", "the method comprises", or similar language in this specification and the appended claims, the singular forms "a", "an", "the" include plural references unless the context clearly dictates otherwise. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs.

[0036] Application overview: With the rapid development of generative artificial intelligence technology, image generation methods based on deep learning have been widely applied in fields such as creative design and virtual scene construction. Traditional image generation technologies are mainly divided into two major technical routes: generation methods based on image content analysis and generation methods based on language description.

[0037] In the generation method based on image analysis, researchers usually use generative adversarial networks (GANs) or diffusion models (Diffusion Models) to extract and reconstruct features of the input image. Such methods can effectively capture low-level features such as the texture and color of visual elements, but have obvious limitations in understanding the semantic associations and spatial layout logic of images. For example, in the scene reconstruction task, existing methods are difficult to infer the physical-law-compliant object placement relationships from a single reference image.

[0038] The generation method based on language description maps text semantics to the latent space through cross-modal models such as CLIP, and uses architectures such as VQ-VAE to achieve text-to-image conversion. Although such methods perform outstandingly in response to abstract semantic requirements, there are problems such as uncontrollable visual details and ambiguous spatial relationships. When faced with descriptions containing clear spatial indications such as "a red sofa is placed in the front left and a square decorative painting is hung on the right wall", the generated results often show object position offsets or attribute mismatches.

[0039] In view of the above technical problems, the general idea of the technical solution provided in this application is as follows: Provide an AI image generation method based on image analysis and language description, including the following steps: Obtain the language description of the target scene to be generated, and input the language description into a pre-trained language model to identify and extract the subject object, spatial orientation predicate, and object object in the language description, and generate a corresponding first description vector; Cluster the first description vector through a pre-trained association analysis model to generate a core element set; Obtain a reference image and input it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object; Match the second description vector with the core element set, calculate the similarity between each second description vector and each core element in the core element set, and based on the similarity, confirm the association relationship between each object and multiple core elements to generate association relationship data; Based on the association relationship data, reorganize each object to generate a reorganized object set; Optimize the layout of the reorganized object set based on the spatial relationship between each object and the association degree between the core elements to obtain the layout information of the reorganized object set; Generate a target image according to the layout information of the reorganized object set and the first description vector.

[0040] This method provides high-level semantic constraints through language description (such as subject object, object object, spatial orientation relationship), and provides low-level visual features through reference images (such as texture and shape). Align the two through cross-modal matching, so that the generated image meets both semantic accuracy and visual rationality, effectively avoiding the detail loss of pure text generation and the spatial logic confusion of pure image generation. In this way, the accuracy of target image generation can be improved. In addition, gradient optimization of the object position based on the association relationship data can ensure that the generated image conforms to physical laws (such as avoiding object collisions).

[0041] After introducing the basic principle of this application, the various non-limiting implementation manners of this application will be specifically introduced below with reference to the accompanying drawings of the specification. Please refer to Figure 1 , The embodiment of this application provides an AI image generation method based on image analysis and language description, including the following steps:

[0042] S101: Obtain the language description of the target scenario to be generated, input the language description into a pre-trained language model, identify and extract the subject object, spatial orientation predicate, and object object in the language description, and generate a corresponding first description vector;

[0043] Specifically, in some embodiments, the execution subject of the AI image generation method based on image analysis and language description (e.g., a computer device) can input the language description into the pre-trained language model through the following steps, identify and extract the subject object, spatial orientation predicate, and object object in the language description, and generate a corresponding first description vector:

[0044] The first step is to perform dependency syntactic analysis on the language description to extract a set of semantic triples:

[0045] ;

[0046] Among them, represents the subject object (Subject), represents the object object (Object), represents the spatial orientation predicate (Predicate, such as "on the left", "above");

[0047] The second step is to generate a semantic vector through a pre-trained multimodal encoder:

[0048] ;

[0049] ; is the preset dimension of the semantic vector;

[0050] The third step is to map the spatial orientation predicate to geometric coordinates:

[0051] ;

[0052] Among them, represents the two-dimensional spatial coordinates corresponding to the predicate ;

[0053] The fourth step is to fuse semantic and geometric information to generate an enhanced vector:

[0054] ;

[0055] Among them, is the first weight matrix, is the second weight matrix, , LayerNorm represents the layer normalization operation;

[0056] The fifth step is to generate the first description vector through attention pooling:

[0057] ;

[0058] Among them, is the attention weight vector, and Softmax is the normalization function. In this way, the relationship between objects can be explicitly modeled through the semantic triple (s, p, o) (such as "the sofa is on the left side of the table"), avoiding the fuzzy representation of spatial relationships by traditional text encoders. Mapping the predicate p to coordinates (such as "left side" → x < 0.5) can enable the orientation words in the language description to directly participate in the layout calculation, reduce position ambiguity, provide an initial position for subsequent layout optimization, and accelerate convergence. Redundant descriptions (such as irrelevant modifiers) are automatically filtered through the Softmax weight, improving the representation purity of the core elements.

[0059] S102: Cluster the first description vector through a pre-trained association analysis model to generate a set of core elements;

[0060] Specifically, in some embodiments, the above-mentioned execution subject can cluster the first description vector through a pre-trained association analysis model to generate a set of core elements through the following steps:

[0061] The first step is to construct a semantic graph , the vertex set , and the edge weight is calculated as:

[0062] ;

[0063] Among them, is a learnable weight matrix, is the balance coefficient of semantic and spatial similarity, is the Sigmoid activation function, is the coordinate intersection over union calculation;

[0064] The second step is to iteratively update the vertex features through a graph attention network:

[0065] ;

[0066] Among them, is the neighborhood similarity threshold, and GAT represents the graph attention network;

[0067] The third step is to perform spectral clustering on the updated vertex vectors:

[0068] ;

[0069] Among them, is the preset number of clustering centers, represents the th clustering center;

[0070] Step 4: Filter low-density clustering centers to generate a core element set:

[0071] ;

[0072] Among them, is the clustering radius threshold (such as the Euclidean distance threshold), is the minimum number of samples, and Card represents the set cardinality; the core element set . In this way, a semantic graph can be constructed (vertices = semantic triple vectors, edge weights = semantic similarity + spatial intersection-over-union), and the neighborhood information can be aggregated through a graph attention network (GAT). Then, spectral clustering is performed on the updated vertices to filter out low-density clusters and obtain the core element set. Specifically, when calculating the edge weights, both semantic vector similarity and spatial compatibility are considered to avoid clustering bias caused by relying on only a single modality. When the GAT is iteratively updated, only the edges with weights > τ (high-confidence associations) are retained to exclude noise interference and improve the robustness of clustering. Compared with local methods such as K-means, spectral clustering can discover the potential community structure in the semantic graph (such as separating "living room furniture" from "kitchen utensils").

[0073] S103: Obtain a reference image and input it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object;

[0074] Specifically, in some embodiments, the reference image can be manually uploaded by an operator through the interaction module (such as an input box, an interaction interface) of the above-mentioned execution entity. Of course, the present application is not limited to this. In other embodiments, the reference image can also be obtained by the above-mentioned execution entity based on a language description by calling a preset AI search model.

[0075] Specifically, in some embodiments, the above-mentioned execution entity can input the reference image into a pre-trained image feature extraction model through the following steps to generate a second description vector corresponding to each object:

[0076] Step 1: Segment the reference image through an instance segmentation model to generate an object mask and a bounding box:

[0077] ;

[0078] Among them, is the input reference image tensor; is the binary mask of the th object, is the bounding box coordinate, normalized to the [0, 1] interval;

[0079] Step 2: Extract the mask region features:

[0080] ;

[0081] Among them, is the image feature of the th object, is the dimension of the image feature, represents element-wise multiplication;

[0082] In the third step, fuse the feature and the coordinate to generate the second description vector:

[0083] ;

[0084] Among them, MLP is a multi-layer perceptron, and [;] represents feature concatenation. Among them, only extracting the feature of the mask area can avoid the interference of background noise and improve the feature purity. Concatenating the bounding box coordinate and the image feature can make the second description vector contain the object position information, which is convenient for matching with the geometric coordinates of the core elements.

[0085] S104: Match the second description vector with the core element set, calculate the similarity between each second description vector and each core element in the core element set, and confirm the association relationship between each object and multiple core elements based on the similarity, so as to generate association relationship data;

[0086] Specifically, in some embodiments, the above execution subject can match the second description vector with the core element through the following steps, calculate the similarity between each second description vector and each core element, confirm the association relationship between each object and multiple core elements based on the similarity, and generate association relationship data:

[0087] In the first step, calculate the cross-modal similarity matrix:

[0088] ;

[0089] Among them, is the spatial constraint intensity coefficient (controlling the influence of coordinate difference on the similarity), is the th object center coordinate, represents the th clustering center of the core element set;

[0090] In the second step, perform bidirectional matching to generate association pairs:

[0091] ;

[0092] Among them, is the similarity matching threshold;

[0093] In the third step, generate association relationship data:

[0094] ;

[0095] Among them, represents the tensor product (outer product); .

[0096] Among them, the similarity requires both semantic matching (high cosine similarity) and spatial proximity (small coordinate distance) to avoid mismatches caused by relying solely on a single modality (such as semantic matching but position conflict). Performing bidirectional matching to generate associated pairs can ensure the uniqueness of the matching and prevent the confusion of one-to-many mappings.

[0097] S105: Based on the associated relationship data, reorganize each object to generate a set of reorganized objects;

[0098] Specifically, in some embodiments, the above-mentioned execution entity can reorganize each object based on the associated relationship data through the following steps to generate a set of reorganized objects:

[0099] The first step is to group the objects based on the associated relationship matrix:

[0100] ;

[0101] Among them, is the set of objects belonging to the th core element;

[0102] The second step is to aggregate the intra-group features to generate reorganized objects:

[0103] ;

[0104] Among them, is the number of objects within the group;

[0105] The third step is to generate a set of reorganized objects:

[0106] ;

[0107] In this way, the reorganized vector is still dimensional, aligned with the dimension of the core element , which is convenient for loss calculation in layout optimization. The set of reorganized objects not only retains the details of the reference image but also conforms to the core elements described in the language, solving the problem of inconsistent object attributes in traditional methods (such as "wooden table" being wrongly replaced by "glass table").

[0108] S106: Based on the spatial relationships between objects and the degree of association between core elements, optimize the layout of the set of reorganized objects to obtain the layout information of the set of reorganized objects;

[0109] Specifically, in some embodiments, the above-mentioned execution entity can optimize the layout of the recombined object set based on the spatial relationship between objects and the degree of association between core elements through the following steps to obtain the layout information of the recombined object set:

[0110] First step, evaluate the quality of the current coordinates through a preset layout optimization objective function, and the layout optimization objective function is:

[0111] ;

[0112] Among them, is the loss weight coefficient, is the collision distance threshold (controlling the object spacing), represents the th object's bounding box coordinates, is the matrix Frobenius norm;

[0113] Second step, update the layout coordinates through gradient descent:

[0114] ;

[0115] Among them, is the learning rate, is the gradient of the loss function with respect to the coordinates;

[0116] Third step, generate the layout information of the recombined object set ;

[0117] Among them, represents the set of final spatial position coordinates of the objects; ; Each represents the coordinates of a bounding box of an object; , are the upper left coordinates of the bounding box; , are the lower right coordinates of the bounding box. In this way, multi-objective joint optimization can be performed to simultaneously minimize the object collision term and maximize the semantic matching term, achieving a balance between layout rationality and semantic consistency, and making the object spacing in the generated image conform to physical laws (such as keeping a reasonable distance between a chair and a table).

[0118] S107: Generate a target image according to the layout information of the recombined object set and the first description vector.

[0119] Specifically, in some embodiments, the above-mentioned execution entity can generate a target image through the following steps according to the layout information and the first description vector:

[0120] In the first step, a preliminary image is generated through a preset conditional diffusion model, where the preset conditional diffusion model is:

[0121] ;

[0122] where is the noise image at the step, is the number of denoising steps, PE is the position encoding function, and CrossAttn is the cross-modal attention mechanism;

[0123] In the second step, iterative denoising is performed on the noise image:

[0124] ;

[0125] where is the preset noise scheduling parameter, is the standard Gaussian noise;

[0126] In the third step, the target image is obtained after multiple iterations of denoising:

[0127] ;

[0128] where , where are the height and width of the image respectively. In this way, the global content (such as the scene category) can be controlled through the semantic vector , the layout coordinates constrain the local object positions, and the diffusion model iteratively denoises to gradually refine the details. In this way, the spatial accuracy (matching degree with the language description) of the generated image can be effectively improved.

[0129] Please refer to Figure 2 , based on the same inventive concept as the AI image generation method based on image analysis and language description in the foregoing embodiment, the embodiment of the present application provides an AI image generation system based on image analysis and language description, including: a first generation module 201, configured to obtain the language description of the target scene to be generated, input the language description into a pre-trained language model, identify and extract the subject object, spatial orientation predicate, and object object in the language description, and generate a corresponding first description vector;

[0130] A second generation module 202, configured to cluster the first description vector through a pre-trained association analysis model to generate a core element set;

[0131] A third generation module 203, configured to obtain a reference image and input it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object;

[0132] The fourth generation module 204 matches the second description vectors with the core element set, calculates the similarity between each second description vector and each core element in the core element set, determines the association relationship between each object and multiple core elements based on the similarity, and generates association relationship data;

[0133] The first recombination module 205 recombines each object based on the association relationship data to generate a set of recombined objects;

[0134] The first optimization module 206 optimizes the layout of the set of recombined objects based on the spatial relationship between each object and the association degree between core elements, and obtains the layout information of the set of recombined objects;

[0135] The fifth generation module 207 generates a target image according to the layout information of the set of recombined objects and the first description vector.

[0136] In some embodiments, the first generation module 201 is specifically configured to: perform dependency syntactic analysis on the language description to extract a set of semantic triples:

[0137] ;

[0138] Among them, represents the subject object, represents the object object, represents the spatial orientation predicate (such as "on the left", "above");

[0139] Generate semantic vectors through a pre-trained multi-modal encoder:

[0140] ;

[0141] ; is a preset semantic vector dimension (such as 512 dimensions);

[0142] Map the spatial orientation predicate to geometric coordinates:

[0143] ;

[0144] Among them, represents the two-dimensional spatial coordinates corresponding to the predicate ;

[0145] Fuse semantic and geometric information to generate enhanced vectors:

[0146] ;

[0147] Among them, is the first weight matrix, is the second weight matrix, , LayerNorm represents the layer normalization operation;

[0148] Generate the first description vector through attention pooling:

[0149] ;

[0150] Among them, is the attention weight vector, and Softmax is the normalization function.

[0151] In some embodiments, the second generation module 202 is specifically configured to:

[0152] Construct a semantic graph , the vertex set , and the edge weight is calculated as:

[0153] ;

[0154] Among them, is the learnable weight matrix, is the balance coefficient of semantic and spatial similarity, is the Sigmoid activation function, is the coordinate intersection over union calculation;

[0155] Iteratively update the vertex features through the graph attention network:

[0156] ;

[0157] Among them, is the neighborhood similarity threshold, and GAT represents the graph attention network;

[0158] Perform spectral clustering on the updated vertex vectors:

[0159] ;

[0160] Among them, is the preset number of clustering centers, represents the th clustering center;

[0161] Filter out low-density clustering centers to generate a set of core elements:

[0162] ;

[0163] Among them, is the clustering radius threshold (such as the Euclidean distance threshold), is the minimum number of samples, and Card represents the set cardinality; the set of core elements .

[0164] In some embodiments, the third generation module 203 is specifically configured to:

[0165] Segment the reference image through an instance segmentation model to generate an object mask and a bounding box:

[0166] ;

[0167] Wherein, is the input reference image tensor; is the binary mask of the th object, is the bounding box coordinate, normalized to the [0, 1] interval;

[0168] Extract the mask region features:

[0169] ;

[0170] Wherein, is the image feature of the th object, is the image feature dimension, represents element-wise multiplication;

[0171] Fuse the features and coordinates to generate a second description vector:

[0172] ;

[0173] Wherein, MLP is a multi-layer perceptron, and [;] represents feature concatenation.

[0174] In some embodiments, the fourth generation module 204 is specifically configured to:

[0175] Calculate a cross-modal similarity matrix:

[0176] ;

[0177] Wherein, is the spatial constraint strength coefficient (controlling the influence of coordinate differences on similarity), is the center coordinate of the th object, represents the th clustering center of the core element set;

[0178] Perform bidirectional matching to generate association pairs:

[0179] ;

[0180] Wherein, is the similarity matching threshold;

[0181] Generate associated relationship data:

[0182] ;

[0183] Among them, represents the tensor product (outer product); .

[0184] In some embodiments, the first recombination module 205 is specifically configured to:

[0185] Group objects based on the associated relationship matrix:

[0186] ;

[0187] Among them, is the set of objects belonging to the th core element;

[0188] Aggregate the features within the group to generate a recombined object:

[0189] ;

[0190] Among them, is the number of objects within the group;

[0191] Generate a set of recombined objects:

[0192] ;

[0193] In some embodiments, the first optimization module 206 is specifically configured to:

[0194] Evaluate the quality of the current coordinates through a preset layout optimization objective function, and the layout optimization objective function is:

[0195] ;

[0196] Among them, is the loss weight coefficient, is the collision distance threshold (controlling the object spacing), represents the bounding box coordinates of the th object, is the matrix Frobenius norm;

[0197] Update the layout coordinates through gradient descent:

[0198] ;

[0199] Among them, is the learning rate, is the gradient of the loss function with respect to the coordinates;

[0200] Generate the layout information of the set of recombined objects ;

[0201] Among them, represents the set of final spatial position coordinates of objects; each represents the coordinates of a bounding box of an object; , are the upper left coordinates of the bounding box; , are the lower right coordinates of the bounding box;

[0202] In some embodiments, the fifth generation module 207 is specifically configured to:

[0203] Generate a preliminary image through a preset conditional diffusion model, where the preset conditional diffusion model is:

[0204] ;

[0205] Among them, is the noise image at the th step, is the number of denoising steps, PE is the position encoding function, and CrossAttn is the cross-modal attention mechanism;

[0206] Perform iterative denoising on the noise image:

[0207] ;

[0208] Among them, is the preset noise scheduling parameter, is the standard Gaussian noise;

[0209] After multiple iterations of denoising, obtain the target image:

[0210] ;

[0211] Among them, , where are the height and width of the image respectively.

[0212] It can be understood that the various modules recorded in this AI image generation system based on image analysis and language description correspond to the respective steps in the AI image generation method described in the reference Figure 1 . Therefore, the operations, features, and beneficial effects described above for the method also apply to the AI image generation system based on image analysis and language description and the modules included therein, and will not be elaborated here.

[0213] Please refer to Figure 3, based on the inventive concept of an AI image generation method based on image analysis and language description in the foregoing embodiments, an embodiment of the present application provides an electronic device. The electronic device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device includes a processing device 301 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the ROM 302 (Read Only Memory) or a program loaded from the storage device 308 into the RAM 303 (Random Access Memory). In the RAM 303, various programs and data required for the operation of the electronic device are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output interface (i.e., the I / O interface 305) is also connected to the bus 304.

[0214] Generally, the following devices can be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data.

[0215] Specifically, according to some embodiments of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such some embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the method of some embodiments of the present application are executed.

[0216] It should be noted that the computer-readable medium described in some embodiments of the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. In some embodiments of the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0217] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.

[0218] The above computer-readable medium may be included in the above electronic device; or may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device is caused to: obtain a language description of a target scene to be generated, and input the language description into a pre-trained language model, identify and extract a subject object, a spatial orientation predicate, and an object object in the language description, and generate a corresponding first description vector; perform clustering on the first description vector through a pre-trained association analysis model to generate a set of core elements; obtain a reference image, and input it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object; match the second description vector with the set of core elements, calculate the similarity between each second description vector and each core element in the set of core elements, confirm the association relationship between each object and multiple core elements based on the similarity, and generate association relationship data; based on the association relationship data, reorganize each object to generate a set of reorganized objects; optimize the layout of the set of reorganized objects based on the spatial relationship between each object and the association degree between the core elements to obtain the layout information of the set of reorganized objects; generate a target image according to the layout information of the set of reorganized objects and the first description vector.

[0219] Computer program code for performing the operations of some embodiments of the present application may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0220] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0221] The modules described in some embodiments of the present application can be implemented in software or in hardware. The described modules can also be provided in a processor: for example, they can be described as: a first generation module, a second generation module, a third generation module, a fourth generation module, a first recombination module, a first optimization module, a fifth generation module. Among them, the names of these modules do not constitute a limitation on the module itself in some cases. For example, the first generation module can also be described as a "language description extraction module".

[0222] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0223] Some embodiments of the present application also provide a computer program product, including a computer program that, when executed by a processor, implements any of the above AI image generation methods based on image analysis and language description.

[0224] Although the present invention has been described in detail above with general descriptions and specific embodiments, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.

Claims

1. An AI image generation method based on image analysis and language description, characterized in that: The following steps are involved: Obtaining a language description of a target scene to be generated, and inputting the language description into a pre-trained language model, identifying and extracting a subject object, a spatial orientation predicate, and an object object in the language description, and generating a corresponding first description vector; Clustering the first description vectors through a pre-trained association analysis model to generate a core element set; Obtain a reference image and input it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object; Matching the second description vector with the core element set, calculating the similarity between each second description vector and each core element in the core element set, confirming the association relationship between each object and multiple core elements based on the similarity, and generating association relationship data; Based on the association relationship data, reorganize the objects to generate a reorganized object set; Optimizing the layout of the reorganized object set based on the spatial relationship between the objects and the correlation between the core elements to obtain layout information of the reorganized object set; A target image is generated according to the layout information of the reorganized object set and the first description vector.

2. The AI ​​image generation method based on image analysis and language description according to claim 1, characterized in that: The steps of inputting the language description into a pre-trained language model, identifying and extracting the subject object, the spatial orientation predicate and the object object in the language description, and generating the corresponding first description vector include: Perform dependency syntactic analysis on the language description to extract a set of semantic triples: ; in, Represents the subject object, Represents an object. Predicates indicating spatial location; Generate semantic vectors through pre-trained multimodal encoder: ; ; is the preset semantic vector dimension; Map spatial orientation predicates to geometric coordinates: ; in, Representation predicate The corresponding two-dimensional space coordinates; Fusion of semantic and geometric information to generate enhanced vectors: ; in, is the first weight matrix, is the second weight matrix, , LayerNorm represents the layer normalization operation; Generate the first description vector through attention pooling: ; in, is the attention weight vector, and Softmax is the normalization function.

3. The AI ​​image generation method based on image analysis and language description according to claim 2, characterized in that: The step of clustering the first description vectors by using a pre-trained association analysis model to generate a core element set includes: Building a semantic graph , vertex set , the edge weight is calculated as: ; in, is the learnable weight matrix, is the balance coefficient between semantic and spatial similarity, is the Sigmoid activation function, Calculate the intersection and union ratio of coordinates; Iteratively update vertex features through the graph attention network: ; in, is the neighborhood similarity threshold, GAT represents graph attention network; Perform spectral clustering on the updated vertex vectors: ; in, is the preset number of cluster centers, Indicates Cluster centers; Filter low-density cluster centers to generate a core feature set: ; in, is the cluster radius threshold, is the minimum number of samples, Card represents the cardinality of the set; the core element set .

4. The AI ​​image generation method based on image analysis and language description according to claim 3 is characterized in that: The steps of obtaining a reference image and inputting it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object include: Segment the reference image using the instance segmentation model to generate object masks and bounding boxes: ; in, is the input reference image tensor; For the The binary mask of the object, is the bounding box coordinate, normalized to the [0,1] interval; Extract mask area features: ; in, For the The image features of an object, is the image feature dimension, Represents element-wise multiplication; Fusion features and coordinates generate the second description vector: ; Among them, MLP is a multi-layer perceptron, and [;] represents feature concatenation.

5. The AI ​​image generation method based on image analysis and language description according to claim 4 is characterized in that: The steps of matching the second description vector with the core elements, calculating the similarity between each second description vector and each core element, and confirming the association relationship between each object and a plurality of core elements based on the similarity, and generating association relationship data include: Calculate the cross-modal similarity matrix: ; in, is the spatial constraint strength coefficient, which is used to control the influence of coordinate differences on similarity. For the The center coordinates of the object, Represents the core element set Cluster centers; Perform bidirectional matching to generate associated pairs: ; in, is the similarity matching threshold; Generate association relationship data: ; in, represents tensor product; .

6. The AI ​​image generation method based on image analysis and language description according to claim 5, characterized in that: The steps of reorganizing the objects based on the association relationship data to generate a reorganized object set include: Group objects based on the association matrix: ; in, For the A collection of objects with core elements; Aggregate features within a group to generate reconstructed objects: ; in, is the number of objects in the group; Generate a reorganized set of objects: 。 7. An AI image generation system based on image analysis and language description, characterized in that: include: A first generation module is used to obtain a language description of a target scene to be generated, and input the language description into a pre-trained language model, identify and extract subject objects, spatial orientation predicates and object objects in the language description, and generate a corresponding first description vector; A second generating module, used for clustering the first description vectors through a pre-trained association analysis model to generate a core element set; A third generation module is used to obtain a reference image and input it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object; a fourth generating module, matching the second description vector with the core element set, calculating the similarity between each second description vector and each core element in the core element set, confirming the association relationship between each object and the plurality of core elements based on the similarity, and generating association relationship data; A first reorganization module, based on the association relationship data, reorganizes the objects to generate a reorganized object set; A first optimization module performs layout optimization on the reorganized object set based on the spatial relationship between the objects and the correlation between the core elements to obtain layout information of the reorganized object set; A fifth generating module generates a target image according to the layout information of the reorganized object set and the first description vector.

8. An electronic device, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processing device, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processing device, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • AI image generation method and system based on image analysis and language description

    CN118037888A

  • Natural scene image description generation method and system

    CN118298431A