Front-end interface generation method, system, medium and terminal based on diffusion model

Through a series process based on the diffusion model and segmentation model, combined with the designer's modification, the front-end picture is generated and converted into wireframe diagrams, the problem of insufficient controllability and designability of front-end interface generation in the existing technology is solved, and efficient and controllable front-end interface generation is achieved.

CN116842295BActive Publication Date: 2025-07-08SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310810368.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-04
Publication Date
2025-07-08
Estimated Expiration
2043-07-04

AI Technical Summary

Technical Problem

The existing diffusion model performs poorly in front-end interface generation tasks. The controllability and designability of direct end-to-end generation are limited by language model understanding, which makes it difficult to meet the needs of front-end design.

Method used

By obtaining and labeling web page image training data, using the series process of diffusion model and segmentation model, the first draft of the front-end picture is generated and converted into a wireframe diagram. Combined with the designer's modification, the combination of front-end components and wireframe diagrams is realized to generate the final front-end interface.

Benefits of technology

It improves the front-end design efficiency, reduces labor costs, and has strong controllability and designability of the generated front-end interface, avoids the limitation of language model comprehension, and reaches the level of general front-end design drafts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116842295B_ABST
    Figure CN116842295B_ABST
Patent Text Reader

Abstract

The present invention provides a front-end interface generation method, system, medium and terminal based on a diffusion model, including: obtaining picture training data; preprocessing the picture training data to obtain corresponding text training data; training a diffusion model according to the picture training data and the text training data to obtain a first diffusion model, and generating a preliminary front-end picture and a preliminary front-end wireframe through the first diffusion model; fusing the preliminary front-end picture and the preliminary front-end wireframe through a segmentation model to obtain a combined front-end picture; further fusing and training the combined front-end picture according to the first diffusion model to obtain a final front-end interface. The present invention uses web pictures crawled and labeled from the Internet as training data, fine-tunes the original model, and performs more excellently in the task of front-end generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method, system, medium, and terminal for generating a front-end interface based on a diffusion model. Background Art

[0002] The front end, also known as the web front end, is the front part of a website that runs on browsers such as PCs and mobile devices and presents web pages to users. Through HTML, CSS, JavaScript, and various derivative technologies, frameworks, and solutions, the user interface interaction of Internet products is realized. In the prior art, there are the Stable Diffusion model and some existing commercial front-end interface generation tools.

[0003] The literature "High-Resolution Image Synthesis with Latent Diffusion Models. CVPR 2022 arXiv:2112.10752, 2022." discloses that by decomposing the image formation process into sequential applications of denoising autoencoders, diffusion models (DMs) achieve state-of-the-art synthesis results in image data and other aspects. However, due to the direct use of end-to-end generation in latent space diffusion models and the Stable Diffusion diffusion model developed based on this technology, and the overly broad data used in training, they perform poorly in the task of generating front-end interfaces. The front-end design technology developed by Midjourney also uses end-to-end generation. Although the generated front-end pictures are relatively beautiful, they are limited by the understanding ability of the language model and cannot well meet the design requirements and frameworks.

[0004] In addition, the Adobe Photoshop plugin developed based on Stable Diffusion can achieve fine-tuning of pictures with natural language prompts, as well as object detection and removal, etc. However, it is mainly applicable to daily life photos and is still not as good as the wireframe-based design process for the front end.

[0005] Therefore, there is an urgent need in the market for a method and system for generating a front-end interface based on a diffusion model that can solve the problems of the controllability and designability of directly generating front-end pictures being limited by the understanding ability of the language model. Summary of the Invention

[0006] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method, system, medium, and terminal for generating a front-end interface based on a diffusion model.

[0007] According to a method for generating a front-end interface based on a diffusion model provided by the present invention, it includes:

[0008] Step S1: Obtain image training data;

[0009] Step S2: Preprocess the image training data to obtain corresponding text training data;

[0010] Step S3: Train a diffusion model based on the image training data and text training data to obtain a first diffusion model, and generate a preliminary front-end image and a preliminary front-end wireframe through the first diffusion model;

[0011] Step S4: Use a segmentation model to fuse the preliminary front-end image and the preliminary front-end wireframe to obtain a combined front-end image;

[0012] Step S5: Further fuse and train the combined front-end image according to the first diffusion model to obtain a final front-end interface.

[0013] Preferably, the preprocessing includes the following sub-steps:

[0014] Step S2.1: Perform text annotation on the image training data;

[0015] Step S2.2: Perform data cleaning based on the results of the text annotation and the image training data;

[0016] The data cleaning includes screening images with a blank area greater than or equal to 90% and removing images that do not contain preset keyword entries in the text annotation.

[0017] Preferably, the step S4 includes:

[0018] Step S4.1: Use a segmentation model to segment the generated preliminary front-end image to obtain required front-end components;

[0019] Step S4.2: Automatically fill the front-end components into the preliminary front-end wireframe according to the matching settings to obtain a combined front-end image.

[0020] Preferably, train an automatic fusion model for matching and fusion according to the position and size of the front-end components and the wireframe in the preliminary front-end wireframe;

[0021] For the text to be added, extract the text, remove it from the generated preliminary front-end image, and add it to the finally obtained front-end interface.

[0022] A front-end interface generation system based on a diffusion model provided by the present invention includes:

[0023] Module M1: Obtain image training data;

[0024] Module M2: Preprocess the picture training data to obtain corresponding text training data;

[0025] Module M3: Train a diffusion model based on the picture training data and text training data to obtain a first diffusion model, and generate a preliminary front-end picture and a preliminary front-end wireframe through the first diffusion model;

[0026] Module M4: Fuse the preliminary front-end picture and the preliminary front-end wireframe through a segmentation model to obtain a combined front-end picture;

[0027] Module M5: Further fuse and train the combined front-end picture according to the first diffusion model to obtain a final front-end interface.

[0028] Preferably, the preprocessing includes the following sub-modules:

[0029] Module M2.1: Perform text annotation on the picture training data;

[0030] Module M2.2: Perform data cleaning based on the results of the text annotation and the picture training data;

[0031] The data cleaning includes screening pictures with a blank area greater than or equal to 90% and excluding pictures that do not contain preset keyword entries in the text annotation.

[0032] Preferably, the module M4 includes:

[0033] Module M4.1: Use a segmentation model to segment the generated preliminary front-end picture to obtain required front-end components;

[0034] Module M4.2: Automatically fill the front-end components into the preliminary front-end wireframe according to the matching settings to obtain a combined front-end picture.

[0035] Preferably, train an automatic fusion model for matching and fusion according to the position and size of the front-end components and the wireframe in the preliminary front-end wireframe;

[0036] For the text to be added, extract the text, remove it from the generated preliminary front-end picture, and add it to the finally obtained front-end interface.

[0037] According to a computer-readable storage medium storing a computer program provided by the present invention, when the computer program is executed by a processor, the steps of the front-end interface generation method based on a diffusion model are implemented.

[0038] According to an intelligent mobile terminal provided by the present invention, it includes the computer-readable storage medium storing the computer program, or includes the front-end interface generation system based on a diffusion model.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] 1. The present invention uses web images crawled and annotated by the Internet as training data, fine-tunes the original model, performs better in the task of front-end generation, realizes the automatic generation from natural language requirement description to the front-end interface of the web page, improves the front-end design efficiency, and reduces the labor cost.

[0041] 2. The present invention uses a series connection process of a diffusion model and a segmentation model, and the generated front-end draft can be transformed into a front-end wireframe, and designers can also modify and fine-tune the wireframe.

[0042] 3. The present invention adopts a combination form of front-end components and wireframes. After generating the front-end draft, the required materials can be extracted and secondary design can be carried out in the given wireframe, which can directly generate the controllability and designability of the front-end image from end to end and is not limited by the understanding ability of the language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:

[0044] Figure 1 It is a schematic diagram of the working process of the present invention.

[0045] Figure 2 It is a schematic diagram of the working process of the wireframe extractor in the present invention.

[0046] Figure 3 It is a schematic diagram of the working process of the integrated training of front-end components and wireframes in the present invention.

[0047] Figure 4 It is a schematic diagram of the text extraction process in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0049] According to a front-end interface generation method based on a diffusion model provided by the present invention, as Figures 1 to 4 shown, it includes:

[0050] Step S1: Obtain image training data. The way to obtain the image training data includes web crawling.

[0051] Step S2: Preprocess the image training data to obtain corresponding text training data. Specifically, the preprocessing includes the following sub-steps:

[0052] Step S2.1: Perform text annotation on the image training data. The annotation method includes using the self-guided language image pre-training method to perform text annotation on the image training data crawled from the network.

[0053] Step S2.2: Based on the results of text annotation and the image training data, perform data cleaning. Data cleaning includes screening images with a blank area greater than or equal to 90% and removing images that do not contain preset keyword entries in the text annotation. Specifically, since the quality of the images crawled from the network is uneven, it is easy to crawl images that are not the front-end interface or even blank. First, use computer vision methods to screen out image data with a blank area exceeding 90%. Then, retrieve and remove entries in the text annotation that do not have keywords such as website, interface, webpage, etc. Finally, a training data set with higher quality is obtained.

[0054] Step S3: According to the image training data and the text training data, train the diffusion model to obtain the first diffusion model, and generate the initial draft of the front-end image and the initial draft of the front-end wireframe through the first diffusion model. Specifically, use the higher-quality training data set obtained in Step S2 and adopt the Low-Rank Adaptation (LoRA) pre-training technique to fine-tune the pre-trained diffusion model to obtain the first diffusion model. Further, the initial draft of the front-end image can be generated on the first diffusion model based on the text prompt. As Figure 2 shown, the initial draft of the front-end wireframe starts from the initial draft of the front-end image. First, use morphological operations to extract the corresponding border information; then cluster the detected borders to obtain a front-end design wireframe composed of several rectangular frames. Combining these rectangular frames, components at the corresponding positions can be extracted from the original image for category detection, and the results are saved to a file separately. The present invention uses a series connection process of a diffusion model and a segmentation model, and the generated initial draft of the front-end can be converted into a front-end wireframe, and designers can also modify and fine-tune the wireframe.

[0055] Step S4: Use the segmentation model to fuse the initial draft of the front-end image and the initial draft of the front-end wireframe to obtain the combined front-end image. Adopting the combination form of front-end components and wireframes, after generating the initial draft of the front-end, the required materials can be extracted, and secondary design can be carried out in the given wireframe, which can directly generate the controllability and designability of the front-end image from end to end and is not limited by the understanding ability of the language model. Specifically, as Figure 3 shown, Step S4 includes:

[0056] Step S4.1: Use the segmentation model to segment the generated preliminary front-end image to obtain the required front-end components.

[0057] Step S4.2: According to the matching settings, automatically fill the front-end components into the preliminary front-end wireframe to obtain the combined front-end image. Specifically, based on the position and size of the front-end components and the wireframe in the preliminary front-end wireframe, train an automatic fusion model for matching and fusion. For the text that needs to be added, as Figure 4 shown, extract the text and remove it from the generated preliminary front-end image and then add it to the finally obtained front-end interface. That is to say, for the text that needs to be added, extract it through a text extractor and remove it from the generated preliminary front-end image, and finally add the preliminary front-end image after removal to the finished design draft to avoid losing details during the diffusion process.

[0058] Step S5: Further fuse and train the combined front-end image according to the first diffusion model to obtain the final front-end interface.

[0059] The present invention trains and integrates a series of neural network models, including using a fine-tuned general diffusion model to complete the generation from a text prompt to a preliminary front-end image, combining a segmentation model to fuse the components in the preliminary image with the wireframe designed by the designer, and finally using an image-to-image diffusion model to improve the overall style. The generated front-end interface reaches the level of a general front-end design draft, saving the time of design and beautification, and at the same time greatly improving the controllability compared with directly using an end-to-end diffusion model.

[0060] The present invention also provides a front-end interface generation system based on a diffusion model. Those skilled in the art can implement the front-end interface generation system based on the diffusion model by executing the step process of the front-end interface generation method based on the diffusion model. That is, the front-end interface generation method based on the diffusion model can be understood as a preferred implementation manner of the front-end interface generation system based on the diffusion model.

[0061] A front-end interface generation system based on a diffusion model provided by the present invention includes:

[0062] Module M1: Obtain picture training data.

[0063] Module M2: Preprocess the picture training data to obtain the corresponding text training data. The preprocessing includes the following sub-modules: Module M2.1: Perform text annotation on the picture training data. Module M2.2: Based on the results of the text annotation and the picture training data, perform data cleaning. Data cleaning includes screening pictures with a blank area greater than or equal to 90% and removing pictures that do not contain preset keyword entries in the text annotation.

[0064] Module M3: Train the diffusion model based on the image training data and text training data to obtain the first diffusion model, and generate the initial draft of the front-end image and the initial draft of the front-end wireframe through the first diffusion model.

[0065] Module M4: Merge the initial draft of the front-end image and the initial draft of the front-end wireframe through the segmentation model to obtain the combined front-end image. Module M4 includes: Module M4.1: Use the segmentation model to segment the generated initial draft of the front-end image to obtain the required front-end components. Module M4.2: Automatically fill the front-end components into the initial draft of the front-end wireframe according to the matching settings to obtain the combined front-end image. Train the automatic fusion model for matching and fusion according to the position and size of the front-end components and the wireframe in the initial draft of the front-end wireframe; for the text to be added, extract the text and remove it from the original image and then add it to the finished design drawing.

[0066] Module M5: Further fuse and train the combined front-end image according to the first diffusion model to obtain the final front-end interface.

[0067] A computer-readable storage medium storing a computer program, and the steps of the front-end interface generation method implemented when the computer program is executed by a processor, according to the present invention.

[0068] An intelligent mobile terminal according to the present invention, including a computer-readable storage medium storing a computer program, or including a front-end interface generation system based on a diffusion model.

[0069] Those skilled in the art know that in addition to implementing the systems, devices, and their respective modules provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to make the systems, devices, and their respective modules provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. to implement the same program. Therefore, the systems, devices, and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structure within the hardware component; the modules for implementing various functions can also be regarded as both software programs for implementing the method and the structure within the hardware component.

[0070] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily with each other.

Claims

1. A front-end interface generation method based on a diffusion model, characterized in that, Including: Step S1: Obtain picture training data; Step S2: Preprocess the picture training data to obtain corresponding text training data; Step S3: Train a diffusion model based on the picture training data and text training data to obtain a first diffusion model, and generate a preliminary front-end picture draft and a preliminary front-end wireframe draft through the first diffusion model; The generation of the preliminary front-end picture draft and the preliminary front-end wireframe draft through the first diffusion model includes: The preliminary front-end picture draft can be generated on the first diffusion model based on a text prompt. The preliminary front-end wireframe draft starts from the preliminary front-end picture draft. First, morphological operations are used to extract corresponding border information; Next, the detected borders are clustered to obtain a front-end design wireframe composed of several rectangular frames; Step S4: Fuse the preliminary front-end picture draft and the preliminary front-end wireframe draft through a segmentation model to obtain a combined front-end picture; Step S4 includes: Step S4.1: Use a segmentation model to segment the generated preliminary front-end picture draft to obtain the required front-end components; Step S4.2: Automatically fill the front-end components into the preliminary front-end wireframe draft according to the matching settings to obtain a combined front-end picture; Step S5: Further fuse and train the combined front-end picture according to the first diffusion model to obtain the final front-end interface.

2. The front-end interface generation method based on a diffusion model according to claim 1, wherein The preprocessing includes the following sub-steps: Step S2.1: Perform text annotation on the picture training data; Step S2.2: Perform data cleaning based on the results of the text annotation and the picture training data; The data cleaning includes screening pictures with a blank area greater than or equal to 90% and excluding pictures that do not contain preset keyword entries in the text annotation.

3. The front-end interface generation method based on a diffusion model according to claim 1, wherein Train an automatic fusion model for matching and fusion according to the position and size of the front-end components and the wireframe in the preliminary front-end wireframe draft; For the text that needs to be added, extract the text, remove it from the generated preliminary front-end draft, and add it to the finally obtained front-end interface.

4. A front-end interface generation system based on a diffusion model, characterized in that, Including: Module M1: Obtain picture training data; Module M2: Preprocess the picture training data to obtain corresponding text training data; Module M3: Train a diffusion model based on the picture training data and text training data to obtain a first diffusion model, and generate a preliminary front-end picture draft and a preliminary front-end wireframe draft through the first diffusion model; The generation of the preliminary front-end picture draft and the preliminary front-end wireframe draft through the first diffusion model includes: The preliminary front-end picture draft is generated on the first diffusion model based on a text prompt. The preliminary front-end wireframe draft starts from the preliminary front-end picture draft. First, morphological operations are used to extract corresponding border information; Next, the detected borders are clustered to obtain a front-end design wireframe composed of several rectangular frames; Module M4: Fuse the preliminary front-end picture draft and the preliminary front-end wireframe draft through a segmentation model to obtain a combined front-end picture; Module M4 includes: Module M4.1: Use a segmentation model to segment the generated preliminary front-end picture draft to obtain the required front-end components; Module M4.2: Automatically fill the front-end component into the initial draft of the front-end wireframe according to the matching settings to obtain the combined front-end image; Module M5: Further fuse and train the combined front-end image according to the first diffusion model to obtain the final front-end interface.

5. The front-end interface generation system based on a diffusion model according to claim 4, wherein The preprocessing includes the following sub-modules: Module M2.1: Perform text annotation on the picture training data; Module M2.2: Perform data cleaning based on the results of the text annotation and the picture training data; The data cleaning includes screening pictures with a blank area greater than or equal to 90% and excluding pictures that do not contain the preset keyword entries in the text annotation.

6. The front-end interface generation system based on the diffusion model according to claim 4, wherein Train the automatic fusion model for matching and fusion according to the position and size of the front-end component and the wireframe in the initial draft of the front-end wireframe; For the text to be added, extract the text, remove it from the generated initial front-end draft, and add it to the finally obtained front-end interface.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the front-end interface generation method based on the diffusion model according to any one of claims 1 to 3.

8. An intelligent mobile terminal, characterized in that, It includes the computer-readable storage medium storing the computer program according to claim 7, or includes the front-end interface generation system based on the diffusion model according to any one of claims 4 to 6.

Citation Information

Patent Citations

  • Interactive interface generation method and device, electronic equipment and storage medium

    CN115730361A

  • Data spreading on charts

    US20150169735A1