An image generation method, system and medium based on object semantics

Through an image generation method based on object semantics, combined with an image generation engine and a semantic analysis model, the problem of difficulty in generating diversified designs and high resource consumption in the prior art is solved, and the flexibility and controllability are improved.

CN119379842BActive Publication Date: 2025-05-30CHINA JILIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411987699.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-30
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

In scenarios such as interior design rendering, it is difficult to generate diverse designs in scenarios such as interior design rendering, and at the same time, it requires a large number of data sets with precise semantic labels, which increases the time and resource consumption of data preparation.

Method used

A method of image generation based on object semantics is proposed. Through the original image and design prompt words input by the user, the target object list and design parameters are extracted, combined with the image generation engine and semantic analysis model, the target effect image is generated, and image enhancement is performed through the weight matrix data.

Benefits of technology

It realizes the generation of diversified designs while keeping the layout unchanged, reducing the time and resource consumption of data preparation, and improving the flexibility and controllability of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119379842B_ABST
    Figure CN119379842B_ABST
Patent Text Reader

Abstract

The present invention proposes an object-semantics-based image generation method, system, and medium, aiming to generate new images that meet the expected effects by intelligently processing the original images and design prompts provided by users. The method first uses a target object extraction module to identify objects with independent semantics from the original images and generate corresponding contour masks; then the design prompts are sent into a semantic analysis model to obtain design parameters including subject object information and template effect information. Subsequently, according to these parameters, an image generation engine creates a series of local images and performs weight sorting based on the correlation degree of each local image with the corresponding target object. Finally, the local images are fused into the original image by combining a logical filtering function and a pyramid mapping function, and the quality of the synthesized image is optimized through an image enhancement model. The present invention can efficiently combine image layout with user design intentions and provide flexible and high-quality image creation solutions for fields such as interior design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer image intelligent generation, and particularly relates to an image generation method, system and medium based on object semantics. Background Art

[0002] In recent years, with the rapid development of artificial intelligence technology, image generation technology based on deep learning has become a research hotspot. However, traditional generation models have certain limitations in practical applications. Especially in scenarios such as interior design rendering that require precise layout control, current methods are often limited to subtle texture or color changes and are difficult to generate diverse designs while maintaining the layout unchanged. In addition, many existing methods require a large amount of datasets with precise semantic labels, increasing the time and resource consumption for data preparation. Therefore, in image generation tasks that require precise layout control and preserve semantic consistency, there is an urgent need for a technical solution that can efficiently extract object semantic information and deeply integrate it with the layout pattern to provide more flexible generation capabilities and lower application thresholds. Summary of the Invention

[0003] To solve the problems of the prior art, the present invention proposes an image generation method, system and medium based on object semantics.

[0004] Specifically, the present application relates to an image generation method based on object semantics, which is characterized by including the following steps:

[0005] S1: A user inputs an original image to be processed and a design prompt. According to the original image to be processed, a list of target objects in the image is extracted. Information of each target object in the list of target objects includes an object with independent semantics in the original image to be processed and an image mask identifying the layout area of the object.

[0006] The design prompt is input into a semantic analysis model to obtain target design parameters, where the target design parameters include main object information and matching template effect information.

[0007] S2: The list of target objects and the target design parameters are input into an image generation engine. The image generation engine initially generates a corresponding target effect image. The target effect image is composed of m non-connected local effect images in layout, where m <= n and n is the length of the list of target objects. Each local effect image is mapped to a unique target object in the list of target objects, and then the target effect image is fused into the original image according to the corresponding layout area.

[0008] S3: The image generation engine also determines the weights of the local effect images according to the target design parameters and performs weight sorting.

[0009] The sorted weight matrix data is used as the model parameters of the image enhancement model. The fused image and the target design parameters are input into the image enhancement model. The image enhancement model enhances each of the local effect images in the fused image, and the enhanced image is used as the target generated image.

[0010] Preferably, the semantic analysis model is trained based on multiple design sample data sets, where each sample in the design sample data set includes an original image, a design keyword, and a template effect image;

[0011] According to the cycle rule or the reviewed user feedback, trigger the update of the semantic analysis model.

[0012] Preferably, the image generation engine draws local effect images corresponding to the top n target objects with the highest degree of association with the subject object information in the target design parameters according to the target design parameters, where the value of n is the empirical effect value in the debugging stage of the image generation engine.

[0013] Preferably, the image mask is the contour mask of the object. The fusion of each local effect image is specifically image stitching according to the centroid of the contour mask.

[0014] Preferably, when n = 1, multiply the single data element value in the weight matrix data by a preset adjustment factor.

[0015] Preferably, among the parameters of the image enhancement model, the action surfaces of the weight matrix data include: model performance, degree of combination of image features, and degree of association of regional positions.

[0016] Preferably, the fusion generation method of each local effect image in step S1 is:

[0017]

[0018] Among them, is the local effect image, is the image corresponding to the target object layout area mapped to , is the image after local fusion, is the logical filtering function, and L is the pyramid mapping function that makes adapt to the dimension.

[0019] Preferably, the weight matrix data is sorted in the following manner in step S3:

[0020]

[0021] Among them, Sort is a sorting function, is the normalized pixel quantity of the nth target object in the original image, is the normalized value of the correlation degree between the nth target object and the main object information in the target design parameters.

[0022] This application also proposes an image generation system based on object semantics, which is characterized by including:

[0023] Target object extraction module: Extract a list of target objects in the image according to the original image to be processed input by the user. The information of each target object in the target object list includes an object with independent semantics in the original image to be processed and an image mask identifying the layout area of the object;

[0024] Semantic analysis model: Input the design prompt words input by the user into the semantic analysis model to obtain target design parameters, and the target design parameters include main object information and matching template effect information;

[0025] Image generation engine: Input the target object list and target design parameters into the image generation engine, and the image generation engine initially generates a corresponding target effect image. The target effect image consists of m non-connected local effect images, m <= n, where n is the length of the target object list, and each local effect image is mapped to a unique target object in the target object list;

[0026] Image fusion module: Fuse the target effect image into the original image according to the corresponding layout area;

[0027] Weight calculation module: The image generation engine also determines the weights of the local effect images according to the target design parameters and performs weight sorting;

[0028] Image enhancement model: Use the sorted weight matrix data as the model parameters of the image enhancement model. Input the fused image and the target design parameters into the image enhancement model, and the image enhancement model performs image enhancement on each local effect image in the fused image and uses the enhanced image as the target generated image

[0029] This application also relates to a computer-readable storage medium, on which program code is stored. When the program code is run by a processor, it executes the steps of the above-mentioned image generation method based on object semantics.

[0030] The beneficial technical effects of this invention patent include: in the requirement of generating an image based on an input image and a prompt, on the one hand, the object layout information carried by the image is extracted, and on the other hand, the semantic analysis of the design prompt input by the user is performed to obtain the target design parameters including the main object information and the matching template effect information. Inputting these two aspects of information into the image generation engine can consider the correlation between the semantics and the object layout in the original image while following the semantic instructions, so that the subsequent target effect image can better meet the layout requirements. In addition, the target effect image includes multiple local effect images. The target effect image is a subset of the target object in terms of the layout set. Since the actual number of local effect images depends on the empirical effect value of the image generation engine in the training and debugging stage, the image generation is more flexible and controllable. In addition, the application of the weight matrix data makes the processing in the image enhancement stage with high computing power requirements more purposeful, reduces unnecessary area processing, and makes the generated image more in line with the expected effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 : A flowchart framework diagram of the method according to an embodiment of the present invention.

[0032] Figure 2 : A diagram generated by the semantic model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] Figure 1 A flowchart framework diagram of the method according to an embodiment of the present invention is shown. As shown in the figure, in an image generation method based on object semantics, a user inputs an original image to be processed and a design prompt. The target object list in the image is extracted according to the original image to be processed. The information of each target object in the target object list includes an object with independent semantics in the original image to be processed and an image mask identifying the layout area of the object.

[0035] In some specific embodiments, first, the user uploads an original image to be processed, such as a photo of interior design. At the same time, the user also needs to provide a design prompt, such as "Install a modern minimalist chandelier". After receiving this information, the system will use a pre-trained target object extraction model to identify and extract the key object objects and their layout areas in the image. For example, in a photo of a living room, elements such as a sofa, a coffee table, and a TV wall may be identified, and a corresponding contour mask (i.e., an image mask) is generated for each element for subsequent operations.

[0036] Input the design prompt into the semantic analysis model to obtain target design parameters, where the target design parameters include main object information and matching template effect information.

[0037] In some specific embodiments, the design prompt "Install a modern minimalist chandelier" will be input into the semantic analysis model. This model is trained through a large number of sample data sets marked with different design styles. It can understand the design intent of natural language and convert it into specific visual design parameters. Optionally, the model will output a set of template effect information that conforms to the "modern minimalist" style, including but not limited to main color preferences, material texture suggestions, and layout optimization suggestions, etc. And "chandelier" is used as the main object information, which, together with the template effect information, is output as the target design parameters of the semantic analysis model.

[0038] Such as Figure 2 As shown, during the generation process of the semantic analysis model, first, a design sample data set containing labeled data needs to be obtained. Each sample in the data set includes an original image, an associated design keyword, and a corresponding template effect image. Optionally, the design keyword and the template effect image are provided by professional designers or collected from public resources and are manually reviewed to ensure their accuracy and consistency. The semantic analysis model adopts a deep learning architecture. In some embodiments, the Transformer architecture is used to capture the complex relationships between design prompts. The model is divided into two main parts: an encoder and a decoder. The encoder is responsible for converting the input design prompt into a fixed-length vector representation; while the decoder generates target design parameters based on this vector, including main object information (such as "chandelier") and matching template effect information (such as color scheme, material texture, etc.).

[0039] Optionally, according to the cycle rule or the reviewed user feedback, trigger the update of the semantic analysis model. In some specific embodiments, to keep the model up-to-date, the semantic analysis model will be retrained regularly (e.g., every quarter) or when the user feedback shows that the current model performance has declined. This involves re-evaluating the effectiveness of the existing data set and introducing new samples to reflect the latest design trends and technological progress.

[0040] Input the target object list and target design parameters into an image generation engine, and let the image generation engine initially generate corresponding target effect images. The target effect images are composed of m non-connected local effect images, where m <= n and n is the length of the target object list. Each local effect image is mapped to a unique target object in the target object list.

[0041] In some specific embodiments, using the target object list and design parameters obtained in the previous steps, the system will start the image generation engine to create a new design scheme. Here, the image generation engine will generate local effect images. If there is no direct matching object for "chandelier" in the design parameters in the target object list, the image generation engine will obtain the closest target object according to a preset matching algorithm. For example, it may match to the existing "ceiling" in the target object list, and then perform intelligent drawing of a chandelier with a "modern minimalist" style in combination with the layout area of the target object "ceiling". The drawing result is the local effect image. There can be multiple local effect images. In actual design, the user's design intention may cover the creation of multiple target objects. Therefore, the generated local effect images are mapped to a unique target object in the target object list.

[0042] The image generation engine draws local effect images corresponding to the top n target objects with the highest degree of association with the main object information in the target design parameters, where the value of n is the empirical effect value in the debugging stage of the image generation engine.

[0043] In some specific embodiments, the image generation engine will evaluate the degree of association between these objects and the main object information of "chandelier". The calculation of the degree of association can be based on factors such as the spatial relationship and functional relevance between objects.

[0044] The image mask is the contour mask of the object. The fusion of each local effect image is specifically to perform image stitching according to the centroid of the contour mask.

[0045] In some specific embodiments, in order to enable the newly added effect image to seamlessly integrate into the original image, a method based on the centroid of the contour mask is adopted for fusion. Specifically, for each local effect image, first calculate the geometric center (i.e., centroid) of its corresponding contour mask. Then, use a logical filtering function and a pyramid mapping function to position and adjust the local effect image according to the centroid position to ensure a smooth and natural transition between the new and old images. This method not only ensures the visual coherence of the synthesized image but also maintains the integrity of the original image content.

[0046] Considering that when integrating the local effect image into the original image, in addition to determining the region position, all newly generated parts must ensure seamless connection with other parts of the original image. In some specific embodiments, a method combining a logical filtering function and a pyramid mapping function is used to achieve smooth transition. Specifically: The target effect image is integrated into the original image according to the corresponding layout area, and the generation method of each local effect image is as follows:

[0047]

[0048] Among them, is the local effect image, is the image corresponding to the target object layout area mapped with and is the image after local fusion, is the logical filtering function, L is the pyramid mapping function that adapts to the dimension. In practical applications, the specific implementations of the logical filtering function and the pyramid mapping function vary depending on the application scenario and performance requirements. In some embodiments, the former adopts a simple weighted average or a more complex adaptive weighting strategy, while the latter may adopt different pyramid representation methods such as Gaussian pyramid and Laplacian pyramid. Combining with the logical filtering function, the pyramid mapping function can adaptively adjust the fusion strategy according to the difference between the local effect image and the background image, which helps to maintain the consistency and coherence of the overall image, while allowing the local details to match the surrounding environment more precisely, thereby improving the realism of the synthesized image. In addition, from the debugging effect, the fusion processing in this way can preserve important visual features at different levels, so that even under small local changes, the key information in the original image can be well retained, ensuring the integrity and semantic consistency of the image content.

[0049] As those skilled in the art can understand, there can be multiple local effect images. Since the importance of each local effect image in actual requirements is not completely the same, differential processing with emphasis can be performed during subsequent image processing to improve the overall computing performance. Therefore, in some specific embodiments, the image generation engine also determines the weights of the local effect images according to the target design parameters and performs weight sorting:

[0050]

[0051] Among them, Sort is the sorting function, is the normalized pixel number of the nth target object in the original image, is the normalized correlation value between the nth target object and the main object information in the target design parameters.

[0052] In practical applications, the specific implementations of the logical filtering function and the pyramid mapping function vary according to the application scenarios and performance requirements. In some embodiments, the former adopts simple weighted averaging or more complex adaptive weighting strategies, while the latter may adopt different pyramid representation methods such as Gaussian pyramids and Laplacian pyramids. In combination with the logical filtering function, the pyramid mapping function can adaptively adjust the fusion strategy according to the differences between the local effect image and the background image, which helps to maintain the consistency and coherence of the overall image, while allowing the local details to match the surrounding environment more precisely, thereby improving the realism of the synthesized image. In addition, from the debugging effect, the fusion processing in this way can preserve important visual features at different levels, so that even under small local changes, the key information in the original image can be well retained, ensuring the integrity and semantic consistency of the image content.

[0053] As can be understood by those skilled in the art, there can be multiple local effect images. Since the importance of each local effect image in actual requirements is not completely the same, differential processing with emphasis can be performed during subsequent image processing to improve the overall computational performance. Therefore, in some specific embodiments, the image generation engine also determines the weights of the local effect images according to the target design parameters and performs weight sorting:

[0054]

[0055] Among them, is the local effect image, is the image corresponding to the layout area of the target object mapped with , is the image after local fusion, is the logical filtering function, and L is the pyramid mapping function that makes adapt to dimensions;

[0056] Weight calculation module: The image generation engine also determines the weights of the local effect images according to the target design parameters and performs weight sorting:

[0057]

[0058] Among them, Sort is the sorting function, is the normalized pixel number of the nth target object in the original image, is the normalized correlation value between the nth target object and the main object information in the target design parameters;

[0059] Image enhancement model: The sorted weight matrix data is used as the model parameters of the image enhancement model. The fused image and the target design parameters are input into the image enhancement model. The image enhancement model performs image enhancement on each of the local effect images in the fused image, and the enhanced image is used as the target generated image.

[0060] An embodiment of the present disclosure provides a non-volatile computer storage medium, on which program code is stored; the computer storage medium stores computer-executable instructions, and the computer-executable instructions can execute the method steps described in the above embodiments.

[0061] It should be noted that the above computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of computer-readable storage media can include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, apparatus, or device. And in the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above. The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device. Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer.

[0062] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. Each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions. The units described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases.

[0063] The preferred embodiments of the present invention are described above to make the spirit of the present invention clearer and easier to understand, and are not intended to limit the present invention. Any modifications, substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope defined by the appended claims of the present invention.

Claims

1. A method for generating an image based on object semantics, characterized in that: The steps include: S1: A user inputs an original image to be processed and a design prompt word, and a target object list in the image is extracted according to the original image to be processed, wherein the information of each target object in the target object list includes an object with independent semantics corresponding to the original image to be processed and an image mask identifying a layout area of ​​the object; Inputting the design prompt words into a semantic analysis model to obtain target design parameters, wherein the target design parameters include subject object information and matching template effect information; S2: Input the target object list and target design parameters into an image generation engine, and the image generation engine preliminarily generates a corresponding target effect image, wherein the target effect image is composed of m local effect images with disconnected layouts, m<=n, where n is the length of the target object list, and each local effect image is mapped to a unique target object in the target object list, and then the target effect image is merged into the original image according to the corresponding layout area; S3: The image generation engine also determines the weights of the local effect images according to the target design parameters, and performs weight sorting, and uses the weight matrix data obtained by sorting as the model parameters of the image enhancement model, and inputs the fused image and the target design parameters into the image enhancement model, and the image enhancement model performs image enhancement on each of the local effect images in the fused image, and uses the enhanced image as the target generation image; The fusion generation method of each local effect image is as follows: in, It is a local effect image. is with The image corresponding to the layout area of ​​the mapped target object, is the local fusion image, is the logical filter function, L is adaptation Dimensional pyramid mapping function.

2. The method for generating an image based on object semantics according to claim 1, characterized in that: The semantic analysis model is obtained by training based on multiple design sample data sets, wherein each sample in the design sample data set includes an original image, a design keyword, and a template effect image; Trigger updates to the semantic analysis model based on cycle rules or user feedback after review.

3. The method for generating an image based on object semantics according to claim 1, characterized in that: The image generation engine draws local effect images corresponding to the first n target objects with the highest correlation with the main object information in the target design parameters according to the target design parameters, where the value of n is the empirical effect value of the image generation engine in the debugging stage.

4. The method for generating an image based on object semantics according to claim 1, characterized in that: The image mask is a contour mask of the object, and the fusion of each local effect image is specifically to perform image stitching according to the centroid of the contour mask.

5. The method for generating an image based on object semantics according to claim 1, characterized in that: When n=1, the value of a single data element in the weight matrix data is multiplied by a preset adjustment factor.

6. The method for generating an image based on object semantics according to claim 1, characterized in that: In the parameters of the image enhancement model, the effect of the weight matrix data includes: model performance, image feature combination degree and correlation degree of regional position.

7. The method for generating an image based on object semantics according to claim 1, characterized in that: In step S3, the weight matrix data is obtained by sorting in the following manner: Among them, Sort is the sorting function, is the normalized number of pixels of the nth target object in the original image, It is the normalized value of the correlation between the nth target object and the main object information in the target design parameters.

8. An image generation system based on object semantics, characterized in that: include: Target object extraction module: extracts a target object list in the image according to the original image to be processed input by the user, wherein the information of each target object in the target object list includes an object with independent semantics corresponding to the original image to be processed and an image mask identifying the layout area of ​​the object; Semantic analysis model: input the design prompt words input by the user into the semantic analysis model to obtain target design parameters, wherein the target design parameters include subject object information and matching template effect information; Image generation engine: input the target object list and target design parameters into the image generation engine, and the image generation engine preliminarily generates a corresponding target effect image, wherein the target effect image is composed of m local effect images with disconnected layouts, m<=n, where n is the length of the target object list, and each local effect image is mapped to a unique target object in the target object list; Image fusion module: fuses the target effect image into the original image according to the corresponding layout area; Weight calculation module: the image generation engine also determines the weight of the local effect image according to the target design parameters and performs weight sorting; Image enhancement model: the weight matrix data obtained by sorting is used as the model parameters of the image enhancement model, the fused image and the target design parameters are input into the image enhancement model, the image enhancement model performs image enhancement on each of the local effect images in the fused image, and the enhanced image is used as the target generated image; The fusion generation method of each local effect image is as follows: in, It is a local effect image. is with The image corresponding to the layout area of ​​the mapped target object, is the local fusion image, is the logical filter function, L is adaptation Dimensional pyramid mapping function.

9. A computer-readable storage medium having program codes stored thereon, wherein the program codes are executed by a processor to execute the steps of the method for generating an image based on object semantics according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image generation method and device and storage medium

    CN117635760A