AI image generation method and system based on image analysis and language description
By combining language description and cross-modal matching and layout optimization of reference images, the problem of difficult to understand semantic correlation and spatial layout logic in image generation in the prior art is solved, and high-precision target image generation is achieved, ensuring the semantic accuracy and visual rationality of the generated images, while also in line with physical laws.
Patent Information
- Application Number
- CN202510496710.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The existing image generation technology has limitations in understanding image semantic correlation and spatial layout logic, and it is difficult to infer object placement relationships that conform to physical laws through a single reference image. Moreover, the methods based on language description have problems such as uncontrollable visual details and blurred spatial relationships.
By obtaining the language description of the target scene to be generated, the subject object, spatial orientation predicate and object object are extracted, the first description vector is generated, and the core element set is generated through the pretrained correlation analysis model clustering. Combining the visual characteristics of the reference image, the association data is generated through cross-modal matching, the object collection is reorganized, and layout optimization is used to ensure that the generated image conforms to physical laws.
It improves the accuracy of target image generation, so that the generated image meets both semantic accuracy and visual rationality, avoids the problems of missing details of pure text generation and the spatial logic confusion of pure image generation, and ensures that the generated image conforms to physical laws.
Smart Images

Figure CN120014096A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image generation, and in particular to an AI image generation method and system based on image analysis and language description. Background Art
[0002] With the rapid development of generative artificial intelligence technology, image generation methods based on deep learning have been widely used in creative design, virtual scene construction and other fields. Traditional image generation technology is mainly divided into two technical routes: generation methods based on image content analysis and generation methods based on language description.
[0003] In the generation method based on image analysis, researchers usually use generative adversarial networks (GANs) or diffusion models to extract and reconstruct the input image. Such methods can effectively capture low-level features such as texture and color of visual elements, but have obvious limitations in understanding image semantic associations and spatial layout logic. For example, in the scene reconstruction task, it is difficult for existing methods to infer the physical object placement relationship from a single reference image.
[0004] The generation method based on language description maps text semantics to latent space through cross-modal models such as CLIP, and uses architectures such as VQ-VAE to achieve text-to-image conversion. Although this type of method performs well in responding to abstract semantic requirements, it has problems such as uncontrollable visual details and vague spatial relationships. When faced with descriptions containing clear spatial instructions such as "a red sofa is placed in the front left, and a square decorative painting is hung on the right wall", the generated results often show object position offset or attribute mismatch.
[0005] The disclosure of the above background technology content is only used to assist in understanding the concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above content has been disclosed on the filing date of this patent application, the above background technology should not be used to evaluate the novelty and creativity of the present application. Summary of the invention
[0006] The present application provides an AI image generation method and system based on image analysis and language description, which can improve the accuracy of target image generation.
[0007] To achieve the above objectives, the present application discloses the following technical solutions: In a first aspect, an embodiment of the present application provides an AI image generation method based on image analysis and language description, comprising the following steps: Obtaining a language description of a target scene to be generated, and inputting the language description into a pre-trained language model, identifying and extracting a subject object, a spatial orientation predicate, and an object object in the language description, and generating a corresponding first description vector; Clustering the first description vector through a pre-trained association analysis model to generate a core element set; Obtain a reference image and input it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object; Matching the second description vector with the core element set, calculating the similarity between each second description vector and each core element in the core element set, confirming the association relationship between each object and the multiple core elements based on the similarity, and generating association relationship data; Based on the association relationship data, each object is reorganized to generate a reorganized object set; Based on the spatial relationship between the objects and the correlation between the core elements, the layout of the reorganized object set is optimized to obtain the layout information of the reorganized object set; A target image is generated according to the layout information of the reorganized object set and the first description vector.
[0008] In the embodiment of the present application, high-level semantic constraints (such as subject objects, object objects, and spatial orientation relationships) are provided through language description, and low-level visual features (such as texture and shape) are provided by reference images. The two are aligned through cross-modal matching, so that the generated image satisfies both semantic accuracy and visual rationality, effectively avoiding the problem of missing details in pure text generation and the problem of spatial logic confusion in pure image generation. In this way, the accuracy of target image generation can be improved.
[0009] In addition, gradient optimization of object positions based on association relationship data can ensure that the generated image conforms to physical laws (such as avoiding object collisions).
[0010] In a second aspect, an embodiment of the present application provides an AI image generation system based on image analysis and language description, including: A first generation module is used to obtain a language description of a target scene to be generated, and input the language description into a pre-trained language model, identify and extract subject objects, spatial orientation predicates and object objects in the language description, and generate a corresponding first description vector; A second generation module is used to cluster the first description vector through a pre-trained association analysis model to generate a core element set; A third generation module is used to obtain a reference image and input it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object; A fourth generation module matches the second description vector with the core element set, calculates the similarity between each second description vector and each core element in the core element set, confirms the association relationship between each object and the multiple core elements based on the similarity, and generates association relationship data; A first reorganization module reorganizes the objects based on the association relationship data to generate a reorganized object set; The first optimization module optimizes the layout of the reorganized object set based on the spatial relationship between the objects and the correlation between the core elements to obtain the layout information of the reorganized object set; The fifth generation module generates a target image according to the layout information of the reorganized object set and the first description vector.
[0011] In a third aspect, an embodiment of the present application provides an electronic device, comprising one or more processors; a storage device on which one or more programs are stored; when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any technical solution of the first aspect.
[0012] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any technical solution of the first aspect is implemented.
[0013] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the method described in any technical solution of the first aspect.
[0014] Among them, the technical effects brought about by any design method in the second to fifth aspects can refer to the technical effects brought about by different design methods in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.
[0016] Figure 1 A flowchart of an AI image generation method based on image analysis and language description provided in some embodiments of the present application; Figure 2 A schematic diagram of the structure of an AI image generation method based on image analysis and language description provided in some embodiments of the present application; Figure 3 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present application. DETAILED DESCRIPTION
[0017] Specific embodiments of the present invention will now be mentioned in detail. Although the present invention is described in conjunction with these specific embodiments, it should be appreciated that it is not intended to limit the present invention to these specific embodiments. On the contrary, these embodiments are intended to cover substitutions, changes or equivalent embodiments that may be included in the spirit and scope of the invention defined by the claims. In the following description, a large number of specific details are set forth in order to provide a comprehensive understanding of the present invention. The present invention may be implemented without some or all of these specific details.
[0018] When used in conjunction with "including," "methods comprising," or similar language in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0019] Application Overview: With the rapid development of generative artificial intelligence technology, deep learning-based image generation methods have been widely used in creative design, virtual scene construction and other fields. Traditional image generation technology is mainly divided into two technical routes: generation methods based on image content analysis and generation methods based on language description.
[0020] In the generation method based on image analysis, researchers usually use generative adversarial networks (GANs) or diffusion models to extract and reconstruct the input image. Such methods can effectively capture low-level features such as texture and color of visual elements, but have obvious limitations in understanding image semantic associations and spatial layout logic. For example, in the scene reconstruction task, it is difficult for existing methods to infer the physical object placement relationship from a single reference image.
[0021] The generation method based on language description maps text semantics to latent space through cross-modal models such as CLIP, and uses architectures such as VQ-VAE to achieve text-to-image conversion. Although this type of method performs well in responding to abstract semantic requirements, it has problems such as uncontrollable visual details and vague spatial relationships. When faced with descriptions containing clear spatial instructions such as "a red sofa is placed in the front left, and a square decorative painting is hung on the right wall", the generated results often show object position offset or attribute mismatch.
[0022] In view of the above technical problems, the overall idea of the technical solution provided by the present application is as follows: an AI image generation method based on image analysis and language description is provided, comprising the following steps: obtaining a language description of a target scene to be generated, and inputting the language description into a pre-trained language model, identifying and extracting the subject object, spatial orientation predicate and object object in the language description, and generating a corresponding first description vector; clustering the first description vector through a pre-trained association analysis model to generate a core element set; obtaining a reference image and inputting it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object; matching the second description vector with the core element set, calculating the similarity between each second description vector and each core element in the core element set, confirming the association relationship between each object and multiple core elements based on the similarity, and generating association relationship data; reorganizing each object based on the association relationship data to generate a reorganized object set; optimizing the layout of the reorganized object set based on the spatial relationship between each object and the association between the core elements to obtain the layout information of the reorganized object set; generating a target image according to the layout information of the reorganized object set and the first description vector.
[0023] This method provides high-level semantic constraints (such as subject objects, object objects, and spatial orientation relationships) through language descriptions, and reference images provide low-level visual features (such as texture and shape). The two are aligned through cross-modal matching, so that the generated image satisfies both semantic accuracy and visual rationality, effectively avoiding the lack of details in pure text generation and the spatial logic confusion in pure image generation. In this way, the accuracy of target image generation can be improved. In addition, gradient optimization of object positions based on association relationship data can ensure that the generated image conforms to physical laws (such as avoiding object collisions).
[0024] After introducing the basic principles of the present application, various non-limiting implementation methods of the present application will be specifically introduced in conjunction with the accompanying drawings. Figure 1 , the embodiment of the present application provides an AI image generation method based on image analysis and language description, comprising the following steps: S101: Obtain a language description of a target scene to be generated, and input the language description into a pre-trained language model, identify and extract subject objects, spatial orientation predicates and object objects in the language description, and generate a corresponding first description vector; Specifically, in some embodiments, the execution subject (e.g., a computer device) of the AI image generation method based on image analysis and language description can input the language description into a pre-trained language model through the following steps, identify and extract the subject object, spatial orientation predicate, and object object in the language description, and generate a corresponding first description vector: The first step is to perform dependency syntactic analysis on the language description and extract a set of semantic triples: ; in, Represents the subject object (Subject), Represents an object (Object), Predicates that indicate spatial location (such as "on the left" and "above"); The second step is to generate semantic vectors through the pre-trained multimodal encoder: ; ; is the preset semantic vector dimension; The third step is to map the spatial orientation predicate into geometric coordinates: ; in, Representation predicate The corresponding two-dimensional space coordinates; The fourth step is to fuse semantic and geometric information to generate enhanced vectors: ; in, is the first weight matrix, is the second weight matrix, , LayerNorm represents the layer normalization operation; The fifth step is to generate the first description vector through attention pooling: ; in, is the attention weight vector, and Softmax is the normalization function. In this way, the relationship between objects can be explicitly modeled through the semantic triple (s, p, o) (such as "the sofa is on the left side of the table"), avoiding the fuzzy representation of spatial relationships by traditional text encoders. Mapping the predicate p to coordinates (such as "left side" → x<0.5) allows the directional words in the language description to directly participate in the layout calculation, reduce position ambiguity, provide an initialization position for subsequent layout optimization, and accelerate convergence. Redundant descriptions (such as irrelevant modifiers) are automatically filtered through the Softmax weight to improve the representation purity of core elements.
[0025] S102: clustering the first description vector through a pre-trained association analysis model to generate a core element set; Specifically, in some embodiments, the execution subject may cluster the first description vector using a pre-trained association analysis model to generate a core element set through the following steps: The first step is to build a semantic graph , vertex set , the edge weight is calculated as: ; in, is the learnable weight matrix, is the balance coefficient between semantic and spatial similarity, is the Sigmoid activation function, Calculate the intersection and union ratio of coordinates; The second step is to iteratively update vertex features through the graph attention network: ; in, is the neighborhood similarity threshold, GAT represents graph attention network; The third step is to perform spectral clustering on the updated vertex vectors: ; in, is the preset number of cluster centers, Indicates Cluster centers; The fourth step is to filter the low-density cluster centers to generate a set of core elements: ; in, is the cluster radius threshold (such as the Euclidean distance threshold), is the minimum number of samples, Card represents the cardinality of the set; the core element set . In this way, we can construct a semantic graph (vertex = semantic triple vector, edge weight = semantic similarity + spatial intersection-union ratio), aggregate neighborhood information through the graph attention network (GAT), and then perform spectral clustering on the updated vertices to filter low-density clusters to obtain a set of core elements. Specifically, the edge weight calculation takes into account both semantic vector similarity and spatial compatibility to avoid clustering bias caused by relying solely on a single modality. When GAT is iteratively updated, only edges with weights > τ (high confidence associations) are retained to eliminate noise interference and improve clustering robustness. Compared with local methods such as K-means, spectral clustering can discover potential community structures in semantic graphs (such as the separation of "living room furniture" and "kitchen utensils").
[0026] S103: Obtain a reference image and input it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object; Specifically, in some embodiments, the reference image can be manually uploaded through the interactive module (e.g., input box, interactive interface) of the above-mentioned execution subject. Of course, the present application is not limited to this. In other embodiments, the reference image can also be obtained by the above-mentioned execution subject by calling a preset AI search model based on a language description.
[0027] Specifically, in some embodiments, the execution subject may generate a second description vector corresponding to each object by performing the following steps and inputting the reference image into a pre-trained image feature extraction model: The first step is to segment the reference image through the instance segmentation model to generate object masks and bounding boxes: ; in, is the input reference image tensor; For the The binary mask of the object, is the bounding box coordinate, normalized to the [0,1] interval; The second step is to extract the mask area features: ; in, For the The image features of an object, is the image feature dimension, Represents element-wise multiplication; The third step is to fuse the features and coordinates to generate the second description vector: ; Among them, MLP is a multi-layer perceptron, and [;] represents feature concatenation. Among them, only extracting the features of the mask area can avoid background noise interference and improve feature purity. Concatenating the bounding box coordinates with the image features can make the second description vector contain object location information, which is convenient for matching with the geometric coordinates of the core elements.
[0028] S104: matching the second description vector with the core element set, calculating the similarity between each second description vector and each core element in the core element set, confirming the association relationship between each object and multiple core elements based on the similarity, and generating association relationship data; Specifically, in some embodiments, the execution subject may match the second description vector with the core elements through the following steps, calculate the similarity between each second description vector and each core element, confirm the association relationship between each object and multiple core elements based on the similarity, and generate association relationship data: The first step is to calculate the cross-modal similarity matrix: ; in, is the spatial constraint strength coefficient (controlling the effect of coordinate differences on similarity), For the The center coordinates of the object, Represents the core element set Cluster centers; The second step is to perform two-way matching to generate association pairs: ; in, is the similarity matching threshold; The third step is to generate association relationship data: ; in, represents tensor product (outer product); .
[0029] Among them, similarity requires both semantic matching (high cosine similarity) and spatial proximity (small coordinate distance) to avoid mismatches caused by relying on only a single modality (such as semantic matching but position conflict). Performing bidirectional matching to generate associated pairs can ensure matching uniqueness and prevent confusion in one-to-many mapping.
[0030] S105: reorganizing the objects based on the association relationship data to generate a reorganized object set; Specifically, in some embodiments, the execution subject may reorganize the objects based on the association relationship data through the following steps to generate a reorganized object set: The first step is to group objects based on the association matrix: ; in, For the A collection of objects with core elements; The second step is to aggregate the features within the group to generate reconstructed objects: ; in, is the number of objects in the group; The third step is to generate the reorganized object set: ; Thus, the reorganized vector Still Dimensions and core elements Dimension alignment facilitates loss calculation in layout optimization. The reorganized object set retains the details of the reference image and conforms to the core elements of the language description, solving the problem of inconsistent object attributes in traditional methods (such as "wooden table" is mistakenly replaced by "glass table").
[0031] S106: Optimizing the layout of the reorganized object set based on the spatial relationship between the objects and the correlation between the core elements to obtain layout information of the reorganized object set; Specifically, in some embodiments, the execution subject may optimize the layout of the reorganized object set based on the spatial relationship between the objects and the correlation between the core elements through the following steps to obtain the layout information of the reorganized object set: The first step is to evaluate the quality of the current coordinates through the preset layout optimization objective function. The layout optimization objective function is: ; in, is the loss weight coefficient, is the collision distance threshold (controls the distance between objects), Indicates The bounding box coordinates of the objects, is the matrix Frobenius norm; The second step is to update the layout coordinates by gradient descent: ; in, is the learning rate, is the gradient of the loss function with respect to the coordinates; The third step is to generate the layout information of the reorganized object set ; in, express The final spatial position coordinate set of an object; Each Represents the coordinates of an object's bounding box; , is the coordinate of the upper left corner of the bounding box; , is the coordinate of the lower right corner of the bounding box. In this way, multi-objective joint optimization can be achieved while minimizing object collisions. Item and maximize semantic matching Item 1 achieves a balance between layout rationality and semantic consistency, so that the distance between objects in the generated image conforms to the laws of physics (such as a reasonable distance between a chair and a table).
[0032] S107: Generate a target image according to the layout information of the reorganized object set and the first description vector.
[0033] Specifically, in some embodiments, the execution subject may generate a target image according to the layout information and the first description vector through the following steps: The first step is to generate a preliminary image through a preset conditional diffusion model, where the preset conditional diffusion model is: ; in, For the The noisy image of the step, is the number of denoising steps, PE is the position encoding function, and CrossAttn is the cross-modal attention mechanism; The second step is to perform iterative denoising on the noisy image: ; in, is the preset noise scheduling parameter, is standard Gaussian noise; The third step is to obtain the target image after multiple iterations of denoising: ; in, ,in are the image height and width respectively. In this way, the semantic vector Control global content (such as scene categories), layout coordinates By constraining the position of local objects, the diffusion model iteratively removes noise and gradually refines the details. In this way, the spatial accuracy of the generated image (matching the language description) can be effectively improved.
[0034] See also Figure 2 , based on the same inventive concept as the AI image generation method based on image analysis and language description in the aforementioned embodiment, the embodiment of the present application provides an AI image generation system based on image analysis and language description, including: a first generation module 201, used to obtain a language description of a target scene to be generated, and input the language description into a pre-trained language model, identify and extract the subject object, spatial orientation predicate and object object in the language description, and generate a corresponding first description vector; The second generating module 202 is used to cluster the first description vector through a pre-trained association analysis model to generate a core element set; The third generation module 203 is used to obtain a reference image and input it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object; The fourth generation module 204 matches the second description vector with the core element set, calculates the similarity between each second description vector and each core element in the core element set, confirms the association relationship between each object and the multiple core elements based on the similarity, and generates association relationship data; A first reorganization module 205 reorganizes the objects based on the association relationship data to generate a reorganized object set; A first optimization module 206 performs layout optimization on the reorganized object set based on the spatial relationship between the objects and the correlation between the core elements to obtain layout information of the reorganized object set; The fifth generating module 207 generates a target image according to the layout information of the reorganized object set and the first description vector.
[0035] In some embodiments, the first generating module 201 is specifically used to: perform dependency syntax analysis on the language description to extract a set of semantic triples: ; in, Represents the subject object (Subject), Represents an object (Object), Predicates that indicate spatial location (such as "on the left" and "above"); Generate semantic vectors through pre-trained multimodal encoder: ; ; is the preset semantic vector dimension (e.g., 512 dimensions); Map spatial orientation predicates to geometric coordinates: ; in, Representation predicate The corresponding two-dimensional space coordinates; Fusion of semantic and geometric information to generate enhanced vectors: ; in, is the first weight matrix, is the second weight matrix, , LayerNorm represents the layer normalization operation; Generate the first description vector through attention pooling: ; in, is the attention weight vector, and Softmax is the normalization function.
[0036] In some embodiments, the second generating module 202 is specifically used for: Building a semantic graph , vertex set , the edge weight is calculated as: ; in, is the learnable weight matrix, is the balance coefficient between semantic and spatial similarity, is the Sigmoid activation function, Calculate the intersection and union ratio of coordinates; Iteratively update vertex features through the graph attention network: ; in, is the neighborhood similarity threshold, GAT represents graph attention network; Perform spectral clustering on the updated vertex vectors: ; in, is the preset number of cluster centers, Indicates Cluster centers; Filter low-density cluster centers to generate a core feature set: ; in, is the cluster radius threshold (such as the Euclidean distance threshold), is the minimum number of samples, Card represents the cardinality of the set; the core element set .
[0037] In some embodiments, the third generation module 203 is specifically used to: Segment the reference image using the instance segmentation model to generate object masks and bounding boxes: ; in, is the input reference image tensor; For the The binary mask of the object, is the bounding box coordinate, normalized to the [0,1] interval; Extract mask area features: ; in, For the The image features of an object, is the image feature dimension, Represents element-wise multiplication; Fusion features and coordinates generate the second description vector: ; Among them, MLP is a multi-layer perceptron, and [;] represents feature concatenation.
[0038] In some embodiments, the fourth generating module 204 is specifically used for: Calculate the cross-modal similarity matrix: ; in, is the spatial constraint strength coefficient (controlling the effect of coordinate differences on similarity), For the The center coordinates of the object, Represents the core element set Cluster centers; Perform bidirectional matching to generate associated pairs: ; in, is the similarity matching threshold; Generate association relationship data: ; in, represents tensor product (outer product); .
[0039] In some embodiments, the first recombination module 205 is specifically used to: Group objects based on the association matrix: ; in, For the A collection of objects with core elements; Aggregate features within a group to generate reconstructed objects: ; in, is the number of objects in the group; Generate a reorganized set of objects: ; In some embodiments, the first optimization module 206 is specifically configured to: The quality of the current coordinates is evaluated by the preset layout optimization objective function, which is: ; in, is the loss weight coefficient, is the collision distance threshold (controls the distance between objects), Indicates The bounding box coordinates of the objects, is the matrix Frobenius norm; Update the layout coordinates via gradient descent: ; in, is the learning rate, is the gradient of the loss function with respect to the coordinates; Generate layout information of the reorganized object set ; in, express The final spatial position coordinate set of an object; Each Represents the coordinates of an object's bounding box; , is the coordinate of the upper left corner of the bounding box; , is the coordinate of the lower right corner of the bounding box; In some embodiments, the fifth generating module 207 is specifically used for: A preliminary image is generated through a preset conditional diffusion model, wherein the preset conditional diffusion model is: ; in, For the The noisy image of the step, is the number of denoising steps, PE is the position encoding function, and CrossAttn is the cross-modal attention mechanism; Perform iterative denoising on a noisy image: ; in, is the preset noise scheduling parameter, is standard Gaussian noise; After multiple iterations of denoising, the target image is obtained: ; in, ,in are the image height and width respectively.
[0040] It is understandable that the modules recorded in the AI image generation system based on image analysis and language description are similar to those in the reference Figure 1 The steps in the AI image generation method based on image analysis and language description described above correspond to each other. Therefore, the operations, features and beneficial effects described above for the method are also applicable to the AI image generation system based on image analysis and language description and the modules contained therein, and will not be repeated here.
[0041] See also Figure 3, based on the inventive concept of an AI image generation method based on image analysis and language description in the aforementioned embodiment, an embodiment of the present application provides an electronic device. The electronic device may include but is not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device includes a processing device 301 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a ROM 302 (read-only memory) or a program loaded from a storage device 308 to a RAM 303 (random access memory). In RAM 303, various programs and data required for the operation of the electronic device are also stored. The processing device 301, ROM 302, and RAM 303 are connected to each other via a bus 304. The input / output interface (i.e., an I / O interface 305) is also connected to the bus 304.
[0042] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data.
[0043] In particular, according to some embodiments of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present application are executed.
[0044] It should be noted that the computer-readable medium recorded in some embodiments of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present application, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In some embodiments of the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0045] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperTextTransferProtocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an adhoc peer-to-peer network), as well as any currently known or future developed network.
[0046] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains a language description of a target scene to be generated, and inputs the language description into a pre-trained language model, identifies and extracts a subject object, a spatial orientation predicate and an object object in the language description, and generates a corresponding first description vector; clusters the first description vector through a pre-trained association analysis model to generate a core element set; obtains a reference image and inputs it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object; matches the second description vector with the core element set, calculates the similarity between each second description vector and each core element in the core element set, confirms the association relationship between each object and multiple core elements based on the similarity, and generates association relationship data; reorganizes each object based on the association relationship data to generate a reorganized object set; optimizes the layout of the reorganized object set based on the spatial relationship between each object and the association between the core elements to obtain layout information of the reorganized object set; generates a target image based on the layout information of the reorganized object set and the first description vector.
[0047] Computer program code for performing the operations of some embodiments of the present application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect through the Internet).
[0048] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0049] The modules described in some embodiments of the present application may be implemented by software or hardware. The modules described may also be set in a processor: for example, they may be described as: a first generation module, a second generation module, a third generation module, a fourth generation module, a first reorganization module, a first optimization module, and a fifth generation module. The names of these modules do not, in some cases, constitute limitations on the modules themselves. For example, the first generation module may also be described as a "language description extraction module."
[0050] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0051] Some embodiments of the present application also provide a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned AI image generation methods based on image analysis and language description.
[0052] Although the present invention has been described in detail above by general description and specific embodiments, it is obvious to those skilled in the art that some modifications or improvements can be made to the present invention. Therefore, these modifications or improvements made without departing from the spirit of the present invention all belong to the scope of protection claimed by the present invention.
Claims
1. An AI image generation method based on image analysis and language description, characterized in that: The following steps are involved: Obtaining a language description of a target scene to be generated, and inputting the language description into a pre-trained language model, identifying and extracting a subject object, a spatial orientation predicate, and an object object in the language description, and generating a corresponding first description vector; Clustering the first description vectors through a pre-trained association analysis model to generate a core element set; Obtain a reference image and input it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object; Matching the second description vector with the core element set, calculating the similarity between each second description vector and each core element in the core element set, confirming the association relationship between each object and multiple core elements based on the similarity, and generating association relationship data; Based on the association relationship data, reorganize the objects to generate a reorganized object set; Optimizing the layout of the reorganized object set based on the spatial relationship between the objects and the correlation between the core elements to obtain layout information of the reorganized object set; A target image is generated according to the layout information of the reorganized object set and the first description vector.
2. The AI image generation method based on image analysis and language description according to claim 1, characterized in that: The steps of inputting the language description into a pre-trained language model, identifying and extracting the subject object, the spatial orientation predicate and the object object in the language description, and generating the corresponding first description vector include: Perform dependency syntactic analysis on the language description to extract a set of semantic triples: ; in, Represents the subject object, Represents an object. Predicates indicating spatial location; Generate semantic vectors through pre-trained multimodal encoder: ; ; is the preset semantic vector dimension; Map spatial orientation predicates to geometric coordinates: ; in, Representation predicate The corresponding two-dimensional space coordinates; Fusion of semantic and geometric information to generate enhanced vectors: ; in, is the first weight matrix, is the second weight matrix, , LayerNorm represents the layer normalization operation; Generate the first description vector through attention pooling: ; in, is the attention weight vector, and Softmax is the normalization function.
3. The AI image generation method based on image analysis and language description according to claim 2, characterized in that: The step of clustering the first description vectors by using a pre-trained association analysis model to generate a core element set includes: Building a semantic graph , vertex set , the edge weight is calculated as: ; in, is the learnable weight matrix, is the balance coefficient between semantic and spatial similarity, is the Sigmoid activation function, Calculate the intersection and union ratio of coordinates; Iteratively update vertex features through the graph attention network: ; in, is the neighborhood similarity threshold, GAT represents graph attention network; Perform spectral clustering on the updated vertex vectors: ; in, is the preset number of cluster centers, Indicates Cluster centers; Filter low-density cluster centers to generate a core feature set: ; in, is the cluster radius threshold, is the minimum number of samples, Card represents the cardinality of the set; the core element set .
4. The AI image generation method based on image analysis and language description according to claim 3 is characterized in that: The steps of obtaining a reference image and inputting it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object include: Segment the reference image using the instance segmentation model to generate object masks and bounding boxes: ; in, is the input reference image tensor; For the The binary mask of the object, is the bounding box coordinate, normalized to the [0,1] interval; Extract mask area features: ; in, For the The image features of an object, is the image feature dimension, Represents element-wise multiplication; Fusion features and coordinates generate the second description vector: ; Among them, MLP is a multi-layer perceptron, and [;] represents feature concatenation.
5. The AI image generation method based on image analysis and language description according to claim 4 is characterized in that: The steps of matching the second description vector with the core elements, calculating the similarity between each second description vector and each core element, and confirming the association relationship between each object and a plurality of core elements based on the similarity, and generating association relationship data include: Calculate the cross-modal similarity matrix: ; in, is the spatial constraint strength coefficient, which is used to control the influence of coordinate differences on similarity. For the The center coordinates of the object, Represents the core element set Cluster centers; Perform bidirectional matching to generate associated pairs: ; in, is the similarity matching threshold; Generate association relationship data: ; in, represents tensor product; .
6. The AI image generation method based on image analysis and language description according to claim 5, characterized in that: The steps of reorganizing the objects based on the association relationship data to generate a reorganized object set include: Group objects based on the association matrix: ; in, For the A collection of objects with core elements; Aggregate features within a group to generate reconstructed objects: ; in, is the number of objects in the group; Generate a reorganized set of objects: 。 7. An AI image generation system based on image analysis and language description, characterized in that: include: A first generation module is used to obtain a language description of a target scene to be generated, and input the language description into a pre-trained language model, identify and extract subject objects, spatial orientation predicates and object objects in the language description, and generate a corresponding first description vector; A second generating module, used for clustering the first description vectors through a pre-trained association analysis model to generate a core element set; A third generation module is used to obtain a reference image and input it into a pre-trained image feature extraction model to generate a second description vector corresponding to each object; a fourth generating module, matching the second description vector with the core element set, calculating the similarity between each second description vector and each core element in the core element set, confirming the association relationship between each object and the plurality of core elements based on the similarity, and generating association relationship data; A first reorganization module, based on the association relationship data, reorganizes the objects to generate a reorganized object set; A first optimization module performs layout optimization on the reorganized object set based on the spatial relationship between the objects and the correlation between the core elements to obtain layout information of the reorganized object set; A fifth generating module generates a target image according to the layout information of the reorganized object set and the first description vector.
8. An electronic device, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processing device, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processing device, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
AI image generation method and system based on image analysis and language description
CN118037888A
Natural scene image description generation method and system
CN118298431A
Artistic font generation method based on semantic input
CN119444925A
Target scene composition using generative ai
US20240127511A1
Cited By
Robot simulation task generation method, medium and equipment
CN121009708A