Front-end code generation method and device, electronic equipment and nonvolatile storage medium
By dividing the page design diagram into sub-regions and using multimodal large models to generate static code, the problem of low efficiency in front-end code development is solved, and efficient and excellent quality front-end code generation is achieved.
Patent Information
- Application Number
- CN202510157382.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-16
AI Technical Summary
Front-end code development is inefficient, and the existing technology has failed to effectively solve this problem.
By obtaining the page design diagram, dividing it into multiple sub-regions, and determining the position structure information between the sub-regions. Then, a multimodal large model is used to analyze the sub-region and generate static codes corresponding to the sub-region. Finally, the static code is spliced based on the location structure information to generate a front-end code file.
It improves the efficiency and code quality of front-end code development, and can automatically generate front-end code that meets user needs based on design drawings and language descriptions.
Smart Images

Figure CN120010847A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a front-end code generation method, device, electronic device and non-volatile storage medium. Background Art
[0002] Front-end development is a vital part of modern software development. It is used to provide packaged services in the user interface in an interactive and user-friendly way. It is widely used in the development of web pages and mobile applications. However, the front-end code development method in related technologies has the technical problem of low efficiency.
[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0004] The embodiments of the present application provide a front-end code generation method, device, electronic device and non-volatile storage medium to at least solve the technical problem of low front-end code development efficiency in related technologies.
[0005] According to one aspect of an embodiment of the present application, a front-end code generation method is provided, including: obtaining a page design drawing, dividing the page design drawing into a plurality of sub-areas, and determining position structure information between the sub-areas, wherein the number of page elements contained in the sub-areas is less than a preset number threshold, and the position structure information is used to characterize the relative position relationship between the sub-areas; using a first model to analyze the sub-areas and generate static codes corresponding to the sub-areas, wherein the first model is obtained by training a large model with multimodal data, and the static codes can restore the page layout of the sub-areas after rendering, and each sub-area corresponds to a static code; according to the position structure information, the static codes are spliced to obtain a front-end code file corresponding to the page design drawing.
[0006] Optionally, dividing the page design drawing into multiple sub-areas includes: using an image segmentation model to determine all page elements contained in the page design drawing, and converting the page elements into a first mask in the form of a mask; filtering the first mask to obtain a second mask, wherein the filtering process is used to screen out interference masks in the first mask that are not related to the page structure; based on the second mask, determining the boundary area in the page design drawing, and determining the center line of the boundary area as a candidate boundary line, wherein the boundary area is a rectangular area in the page design drawing that does not intersect with any second mask and runs through the page design drawing in the horizontal direction or the vertical direction; in response to a control instruction input from the front-end interactive interface, determining a target boundary line among the candidate boundary lines, and dividing the page design drawing according to the target boundary line to obtain sub-areas, wherein the control instruction is used to indicate the candidate boundary line selected by the user.
[0007] Optionally, filtering the first mask to obtain the second mask includes: determining the size data of the first mask, and determining the first mask whose size data exceeds a preset size threshold range as an interference mask and screening out the first mask; determining the first mask containing a hole area inside the mask as an interference mask and screening out the first mask; determining the number of colors contained in the area corresponding to the first mask in the page design diagram, and determining the first mask with a color number of one as an interference mask and screening out the first mask.
[0008] Optionally, each candidate dividing line corresponds to a segmentation scheme, which is used to divide the page design according to the candidate dividing line. Figure 1 Divided into two; before determining the target dividing line among the candidate dividing lines, the method also includes: determining the border length of the dividing area corresponding to the candidate dividing line in the non-penetrating direction, wherein the penetrating direction is the horizontal direction or the vertical direction of the page design drawing; determining the vertical distance between the candidate dividing line and the page reference point, wherein the page reference point includes: the vertex in the upper left corner of the page design drawing; determining the scheme score of the segmentation scheme corresponding to the selected dividing line based on the border length and the vertical distance, wherein the border length is positively correlated with the scheme score, and the vertical distance is negatively correlated with the scheme score; sending the candidate dividing lines corresponding to a preset number of segmentation schemes with the highest scheme scores to the front-end interactive interface for display.
[0009] Optionally, after obtaining the front-end code file corresponding to the page design drawing, the method also includes: rendering based on the front-end code file to obtain the rendering result, and sending the rendering result to the front-end interactive interface for display; using a second model to analyze the modification text, and adjusting the code in the front-end code file based on the analysis result, wherein the modification text is the modification requirements for page elements and layout in the rendering result input by the user in the front-end interactive interface.
[0010] Optionally, the training step of the first model includes: obtaining a basic training data set, wherein the basic training data set contains multiple static code data of a target framework type, and the target framework type includes at least one of the following: a Flutter framework; rendering the static code data in the basic training data set to obtain a page design drawing corresponding to the static code data, and constructing a first training data set based on the static code data and the page design drawing, wherein the first training data set contains multiple first training samples, and each first training sample contains: static code data, and a page design drawing corresponding to the static code data; using the first training data set to train the basic large model to obtain a first model.
[0011] Optionally, the training steps of the second model include: using the basic large model to modify the static code data in the basic training data set, and generating a modification description text corresponding to the modification, wherein the modification description text is a natural language description of the modification of the static code data; based on the static code data before the modification, the modified static code data, and the modification description text, constructing a second training data set, wherein the second training data set contains multiple second training samples, and each second training sample contains: static code data, modified static code data, and modification description text; using the second training data set to train the basic large model to obtain a second model.
[0012] Optionally, obtaining a basic training data set includes: obtaining project code in an open source code platform of a target framework type, determining the front-end code in the project code for rendering the front-end interface, and removing the interactive logical relationship in the front-end code to obtain static code data; and / or, using a basic large model to randomly generate page development requirements, and generate static code corresponding to the page development requirements to obtain static code data; and / or, obtaining static page code of a non-target framework type, and using the basic large model to rewrite the static page code into static code of the target framework type to obtain static code data.
[0013] According to another aspect of the embodiment of the present application, a front-end code generation device is also provided, including: a design drawing segmentation module, used to obtain a page design drawing, divide the page design drawing into multiple sub-areas, and determine the position structure information between the sub-areas, wherein the number of page elements contained in the sub-areas is less than a preset number threshold, and the position structure information is used to characterize the relative position relationship between the sub-areas; a code generation module, used to use a first model to analyze the sub-areas and generate static codes corresponding to the sub-areas, wherein the first model is obtained by training a large model with multimodal data, and the static code can restore the page layout of the sub-areas after rendering, and each sub-area corresponds to a static code; a page splicing module, used to splice the static code according to the position structure information to obtain a front-end code file corresponding to the page design drawing.
[0014] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory and a processor, the processor being used to run a program stored in the memory, wherein the front-end code generation method is executed when the program is run.
[0015] According to another aspect of the embodiments of the present application, a non-volatile storage medium is provided, the non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the front-end code generation method by running the computer program.
[0016] According to another aspect of the embodiments of the present application, a computer program product is provided, including a computer program, which implements the steps of the front-end code generation method when the computer program is executed by a processor.
[0017] In an embodiment of the present application, a page design drawing is obtained, the page design drawing is divided into multiple sub-areas, and the position structure information between the sub-areas is determined, wherein the number of page elements contained in the sub-areas is less than a preset number threshold, and the position structure information is used to characterize the relative position relationship between the sub-areas; a first model is used to analyze the sub-areas and generate static codes corresponding to the sub-areas, wherein the first model is obtained after training the large model with multimodal data, and the static code can restore the page layout of the sub-areas after rendering, and each sub-area corresponds to a static code; according to the position structure information, the static code is spliced to obtain a front-end code file corresponding to the page design drawing. By utilizing the powerful generation ability and intelligent characteristics of the large model, the front-end code that meets the user's needs can be automatically generated according to the design drawing and language description, thereby achieving the purpose of improving development efficiency and code quality, thereby solving the technical problem of low front-end code development efficiency in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0019] Figure 1 It is a hardware structure block diagram of a computer terminal (or electronic device) for implementing a method for front-end code generation provided in an embodiment of the present application;
[0020] Figure 2 It is a schematic diagram of a method flow for front-end code generation provided according to an embodiment of the present application;
[0021] Figure 3 It is a schematic diagram of the reasoning process of a front-end static code generation method based on a multimodal large model provided according to an embodiment of the present application;
[0022] Figure 4 It is an example diagram of a page design diagram of an input image segmentation model provided according to an embodiment of the present application;
[0023] Figure 5 is a schematic diagram of an original output mask of a SAM model provided according to an embodiment of the present application;
[0024] Figure 6 is a schematic diagram of an example of a post-processed mask provided according to an embodiment of the present application;
[0025] Figure 7 is a schematic diagram of a boundary area found according to a mask provided in an embodiment of the present application;
[0026] Figure 8 It is a schematic diagram of a segmentation scheme provided to a user after scoring and sorting according to an embodiment of the present application;
[0027] Fig. 9 It is a schematic diagram of a data collection and large model offline training process provided according to an embodiment of the present application;
[0028] Fig.10 It is a structural diagram of a front-end code generation device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0031] In order to facilitate those skilled in the art to better understand the embodiments of the present application, some technical terms or nouns involved in the embodiments of the present application are explained as follows:
[0032] Large Language Model: It is an artificial intelligence model based on deep learning, which is specially used to process and generate natural language text. It can understand and generate human-like language by training on a large amount of text data. Large language model can be applied to a variety of tasks, such as text generation, translation, question-answering system and conversational robot.
[0033] Multimodal Large Model: The multimodal field believes that different representations of the same thing (such as text, images, sounds, etc.) complement each other in the model's understanding of the thing, and these representations should be used as training signal inputs for the model at the same time. A large model trained with this as the guiding ideology is called a multimodal large model, which can process and generate multiple types of data at the same time, and can also understand and generate richer and more complex information. For example, it can generate an image based on a text description, or generate a corresponding text description based on an image.
[0034] SFT (Supervised Fine-Tuning): refers to further training the model based on the existing model through supervised learning method using task-specific data. SFT fine-tuning can make the model better adapt to specific tasks and improve its performance and accuracy.
[0035] Image Segmentation Task: It is a computer vision technique used to divide an image into multiple meaningful regions or objects. Its purpose is to identify and separate different parts of an image so that each part has its own semantic label. Image segmentation is widely used in medical image analysis, autonomous driving, image editing, and object recognition. Commonly used methods include traditional edge detection and deep learning-based segmentation algorithms.
[0036] SAM model (Segment Anything Model): is an open-source, deep learning-based image segmentation model. The SAM model is trained with a large amount of image data and has a strong general image segmentation capability. It can accurately segment any object in any image without targeted training for specific visual scenes. The SAM model is widely used in a variety of computer vision tasks, such as medical image analysis and autonomous driving.
[0037] Flutter: is a front-end UI development framework based on the Dart language. Flutter provides a wealth of components and tools to build high-performance, high-fidelity mobile applications across platforms (including iOS, Android, Web, etc.).
[0038] In related technologies, traditional front-end development mainly relies on manual coding. Although this method is highly flexible, it is inefficient and requires a high level of technical skills from developers. In order to improve development efficiency and lower the technical threshold for development, low-code platforms are gradually emerging. When users use low-code platforms, they can build applications by simply dragging and dropping preset components. This development method has the advantages of being efficient and easy to use. However, the applications produced by low-code platforms are obviously insufficient in portability and customization, and it is difficult to meet complex needs. Large language models can learn and imitate any text with specific organizational rules, and code is such a typical example. As long as enough static code files are provided for large model training, the large model can generate compilable and excellent style code according to user requirements. Code generation based on large language models requires users to provide detailed language descriptions of page elements and functions, and still requires users to have a certain knowledge base of the terminology and ideas of front-end development.
[0039] In order to solve the above problems, the embodiments of the present application provide relevant solutions. By introducing the multimodal concept, a code text and its rendered page screenshot are regarded as different descriptions of the same thing. In the embodiments of the present application, static code generation based on a multimodal large model can restore the original code only through the screenshot of the static page. The following is a detailed description.
[0040] According to an embodiment of the present application, a method embodiment of front-end code generation is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0041] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal (or electronic device) for implementing a front-end code generation method. Figure 1 As shown, the computer terminal 10 (or electronic device) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0042] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or electronic device). As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0043] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the front-end code generation method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, the front-end code generation method is realized. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0044] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0045] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or electronic device).
[0046] In the above operating environment, the embodiment of the present application provides a front-end code generation method. Figure 2is a schematic diagram of a method flow for front-end code generation provided according to an embodiment of the present application, such as Figure 2 As shown, the method comprises the following steps:
[0047] Step S202, obtaining a page design diagram, dividing the page design diagram into a plurality of sub-regions, and determining position structure information between the sub-regions, wherein the number of page elements contained in the sub-regions is less than a preset number threshold, and the position structure information is used to characterize the relative position relationship between the sub-regions;
[0048] Step S204: using the first model to analyze the sub-region and generate a static code corresponding to the sub-region, wherein the first model is obtained by training the large model with multimodal data, and the static code can restore the page layout of the sub-region after rendering, and each sub-region corresponds to a static code;
[0049] Step S206, splicing the static code according to the position structure information to obtain the front-end code file corresponding to the page design drawing.
[0050] Through the above steps, by utilizing the powerful generation capabilities and intelligent characteristics of the large model, it is possible to automatically generate front-end code that meets user needs based on the design drawings and language descriptions, thereby achieving the purpose of improving development efficiency and code quality, and thus solving the technical problem of low front-end code development efficiency in related technologies.
[0051] The following further introduces the front-end code generation method in steps S202 to S206 of the embodiment of the present application.
[0052] Figure 3 is a schematic diagram of the reasoning process of a front-end static code generation method based on a multimodal large model provided in an embodiment of the present application, such as Figure 3 As shown, in the embodiment of the present application, a high-efficiency and general front-end static code generation method is proposed by utilizing a large language model, an image segmentation model, and a large multimodal model, so as to significantly improve the front-end development efficiency. Through the stages of design drawing segmentation, block multimodal code generation, page splicing, and page style fine-tuning, static code with correct logic and elegant and unified style can be generated. The specific process is as follows.
[0053] First, obtain the page design diagram and segment the page design diagram. A large number of experiments have shown that when the number of elements in the page design diagram exceeds a certain range, some elements will be omitted in the generation result of the multimodal large model. Therefore, in this embodiment, the complex web page design diagram can be decomposed into multiple relatively simple parts by segmenting the page design diagram into multiple sub-areas, so as to reduce the element complexity of each sub-area after segmentation, and improve the success rate of static code generation of each sub-area after segmentation by the multimodal large model. At the same time, the relative position information of each sub-area after segmentation will also be saved in this step. The specific steps are as follows.
[0054] In some embodiments of the present application, dividing a page design drawing into multiple sub-areas includes the following steps: using an image segmentation model to determine all page elements contained in the page design drawing, and converting the page elements into a first mask in the form of a mask; filtering the first mask to obtain a second mask, wherein the filtering process is used to screen out interference masks in the first mask that are not related to the page structure; based on the second mask, determining a boundary area in the page design drawing, and determining the center line of the boundary area as a candidate boundary line, wherein the boundary area is a rectangular area in the page design drawing that does not intersect with any second mask and runs through the page design drawing in the horizontal or vertical direction; in response to a control instruction input from the front-end interactive interface, determining a target boundary line among the candidate boundary lines, and dividing the page design drawing according to the target boundary line to obtain sub-areas, wherein the control instruction is used to indicate the candidate boundary line selected by the user.
[0055] Specifically, the user interface of an application can usually be interpreted as a tree structure: a complete page consists of several sub-parts arranged in a row or a column, and each sub-part can also be further divided in this way. Therefore, in this embodiment, the SAM model fine-tuned by SFT can be used to extract the elements in the design drawing, and then the dividing lines between the elements are captured through post-processing to restore the tree structure of the page.
[0056] For example, for a Figure 4 The page design diagram shown in FIG. 1 first uses the image segmentation model SAM to output the mask of all page elements therein (i.e., the first mask), such as Figure 5 As shown, the first mask can be filtered using a post-processing algorithm to remove the interference mask in the first mask to obtain the second mask. The specific steps are as follows.
[0057] In some embodiments of the present application, filtering the first mask to obtain the second mask includes the following steps: determining the size data of the first mask, and determining the first mask whose size data exceeds a preset size threshold range as an interference mask and screening out the first mask; determining the first mask containing a hole area inside the mask as an interference mask and screening out the first mask; determining the number of colors contained in the area corresponding to the first mask in the page design diagram, and determining the first mask with a color number of one as an interference mask and screening out the first mask.
[0058] Specifically, in this embodiment, the first mask that is too large or too small, contains holes, and has a single color in the corresponding original image area can be determined as a noise mask. These masks are usually small components that are irrelevant to the parsed page structure, and background elements that affect the extraction of the boundary line. By screening out these masks, it can be ensured that the generated code only focuses on the key elements of the design and avoids generating invalid code. Figure 5 The second mask obtained by filtering the first mask in is as follows Figure 6 shown.
[0059] After obtaining the second mask, it is necessary to further determine the rectangular areas in the page design that do not intersect with any second mask and run through the entire image in the horizontal or vertical direction, and determine these areas as boundary areas, such as Figure 7 The green strip shown in the figure is the dividing area determined in the page design diagram. The center line of the dividing area can then be determined as a candidate dividing line. Each candidate dividing line corresponds to a segmentation scheme, and multiple segmentation schemes that divide the image into two can be formed.
[0060] In the embodiment of the present application, a scoring formula may be used to list multiple horizontal and / or vertical segmentation schemes with the highest scores for the user to choose from, such as Figure 8 As shown, according to the user's selection (control instruction), the target dividing line among the candidate dividing lines is determined, and the page design drawing is segmented according to the target dividing line. The above segmentation operation can then be repeated for the segmented sub-areas to continue to refine the segmented sub-areas.
[0061] The specific steps for scoring the segmentation scheme are as follows.
[0062] As an optional implementation, each candidate dividing line corresponds to a segmentation scheme, and the segmentation scheme is used to divide the page design according to the candidate dividing line. Figure 1Divided into two; before determining the target dividing line among the candidate dividing lines, the method also includes the following steps: determining the border length of the dividing area corresponding to the candidate dividing line in the non-penetrating direction, wherein the penetrating direction is the horizontal direction or the vertical direction of the page design drawing; determining the vertical distance between the candidate dividing line and the page reference point, wherein the page reference point includes: the vertex in the upper left corner of the page design drawing; determining the scheme score of the segmentation scheme corresponding to the selected dividing line based on the border length and the vertical distance, wherein the border length is positively correlated with the scheme score, and the vertical distance is negatively correlated with the scheme score; sending the candidate dividing lines corresponding to a preset number of segmentation schemes with the highest scheme scores to the front-end interactive interface for display.
[0063] The scheme scoring mechanism can automatically screen out the optimal segmentation scheme, reducing manual intervention. It is suitable for front-end development processes with a high degree of automation. At the same time, the optimal scheme is displayed for users to choose from, which not only ensures the efficiency of code generation but also retains the flexibility of design.
[0064] After completing all the segmentation operations, the original page design drawing will be divided into a series of non-overlapping rectangular areas (sub-areas), and each sub-area will be input into the multimodal large model as a separate design drawing. In addition, in the segmentation step, the structure tree (i.e., position structure information) of the original design drawing will be restored according to the user's operation log for use in the page splicing step.
[0065] For each sub-area after segmentation, static code generation is performed through the first model (static generation model) to obtain a static code of a target framework type corresponding to the sub-area. In this embodiment, the above-mentioned target framework type is illustrated by taking the static code of the Widget class in the Flutter framework as an example.
[0066] The first model mentioned above is obtained by training the large model with multimodal data. The training process of the model is described in detail below. Fig. 9 As shown, the details are as follows.
[0067] In some embodiments of the present application, the training step of the first model includes: obtaining a basic training data set, wherein the basic training data set contains multiple copies of static code data of the target framework type, and the target framework type includes at least one of the following: Flutter framework; rendering the static code data in the basic training data set to obtain a page design drawing corresponding to the static code data, and constructing a first training data set based on the static code data and the page design drawing, wherein the first training data set contains multiple first training samples, and each first training sample contains: static code data, and a page design drawing corresponding to the static code data; using the first training data set to train the basic large model to obtain the first model.
[0068] Specifically, you first need to collect data (static page code for training) and build a basic training data set to fine-tune the multimodal large model. The specific steps are as follows.
[0069] In some embodiments of the present application, obtaining a basic training data set includes the following steps: obtaining project code in an open source code platform of a target framework type, determining the front-end code in the project code for rendering the front-end interface, and removing the interactive logical relationship in the front-end code to obtain static code data; and / or, using a basic large model to randomly generate page development requirements, and generate static code corresponding to the page development requirements to obtain static code data; and / or, obtaining static page code of a non-target framework type, and using the basic large model to rewrite the static page code into static code of the target framework type to obtain static code data.
[0070] In this embodiment, the channels for static code data of mobile phones include but are not limited to: 1) Pulling Flutter front-end code repositories from multiple open source code hosting platforms, extracting single files (project codes) that can render pages by tracking routing files in their projects, and further screening out the front-end codes of pages implemented using the StatelessWidget class, and cleaning the code data to remove the residual interaction-related logic therein, and obtaining completely static static code data; 2) Using a large language basic model (i.e., a basic large model) to randomly generate development requirement descriptions of thousands of different pages, and then based on these requirement descriptions, using the large language basic model to write corresponding StatelessWidget class code data; 3) Collecting static page codes and screenshot data implemented in non-Flutter languages from the open source data set WebSight, and calling the large language basic model to rewrite the codes therein into Flutter codes implemented using StatelessWidget.
[0071] The embodiment of the present application collects and organizes the existing Flutter code on a large scale, cleans the code to obtain a pure static version, and renders the corresponding screenshot data, forming a dedicated data set for generating Flutter static code from design drawings. In addition, the size of this data set is expanded by autonomous generation of large models and transcription from other languages. This data set plays an important role in understanding the correspondence between Flutter coding and its visual rendering results in a multimodal large model.
[0072] After obtaining the basic training data set, the static code data in the basic training data set can be rendered to obtain the page design corresponding to the static code data, forming a sample pair of (page design, StatelessWidget class code) (i.e., the first training sample mentioned above), thereby obtaining the first training data set. Through the first training data set, the multimodal large model is fine-tuned and trained, and the first model capable of generating Flutter page code based on the design can be obtained.
[0073] The first model obtained through training can generate corresponding static code according to the page design and sub-area, and render the code into a page and feed it back to the user. After obtaining the static code corresponding to the sub-area, the generated static code segments can be reassembled into a complete code file according to the position structure information retained in the segmentation link to obtain the front-end code file corresponding to the page design.
[0074] For example, you can use the built-in RowView and ColumnView components of Flutter to organize the generated codes into complete page front-end code files by rows or columns. This file can be directly ported to the user-side Flutter project for use.
[0075] Furthermore, after obtaining the front-end code file corresponding to the page design drawing, the generated code and rendering results can be visualized to the user, and a multi-round dialogue large language model (second model) is provided. The user can use natural language to propose modification requirements for page elements and layout, and the static code is continuously updated until it meets user requirements. The specific steps are as follows.
[0076] In some embodiments of the present application, after obtaining the front-end code file corresponding to the page design drawing, the method also includes the following steps: rendering based on the front-end code file to obtain the rendering result, and sending the rendering result to the front-end interactive interface for display; using a second model to analyze the modification text, and adjust the code in the front-end code file based on the analysis result, wherein the modification text is the modification requirements for page elements and layout in the rendering result input by the user in the front-end interactive interface.
[0077] Specifically, in this embodiment, the second model can be obtained by fine-tuning the large language model, wherein the training steps of the second model are as follows.
[0078] In some embodiments of the present application, the training step of the second model includes the following steps: using the basic large model to modify the static code data in the basic training data set, and generating a modification description text corresponding to the modification, wherein the modification description text is a natural language description of the modification of the static code data; based on the static code data before the modification, the modified static code data, and the modification description text, construct a second training data set, wherein the second training data set contains multiple second training samples, and each second training sample contains: static code data, modified static code data, and modification description text; using the second training data set to train the basic large model to obtain a second model.
[0079] In this embodiment, the second training data set can be constructed based on the basic training data set obtained when training the first model. Specifically, the static code data in the basic training data set is randomly fine-tuned and modified using the basic large model, for example, the style or type of a single component is modified, and a natural language modification requirement (i.e., the above-mentioned modification description text) is generated to form a sample group (i.e., the second training sample) of (code before modification, modification requirement, code after modification), thereby obtaining the second training data set. Afterwards, the large language model is fine-tuned and trained using the second training data set to obtain a second model capable of generating modified code based on the initial code and modification requirements.
[0080] During use, users can propose page modification requirements (modify text) for the generated front-end code file, and the trained second model can generate modified code based on the code and modification requirements. This process can be repeated until the user's needs are met.
[0081] The embodiment of the present application uses multiple large models to collaborate and automatically generate front-end static code under the Flutter framework, thereby significantly improving the efficiency and code quality of front-end development. Specifically, the present application can automatically generate code by automatically generating front-end projects according to product design drawings and requirement documents, realize the automated writing, modification, and deployment of front-end code, reduce development time and labor costs, reduce development time and labor costs, and can greatly improve the efficiency of front-end development; the generation process of the present application can handle various scenarios and designs, and can generate corresponding page codes for various application scenarios and design drawing forms (from simple hand-drawn drafts to complex real page design drawings), without the restrictions of scenarios and configurations, and utilizes the modal fusion capabilities of multimodal large models to generate highly adaptable code, while allowing users to fine-tune the code in a fine-grained manner through page style fine-tuning to meet design requirements, with strong adaptability and versatility.
[0082] In addition, the static page code produced by this application can easily add interactive functions by retaining the interface fields of each component, and supports organic integration with manually pre-written or other large language model-generated back-end code, so as to quickly complete the complete application construction.
[0083] According to an embodiment of the present application, an embodiment of a front-end code generation device is also provided. Fig.10 Schematic diagram of a front-end code generation device provided according to an embodiment of the present application. Fig.10 As shown, the device comprises:
[0084] The design drawing segmentation module 100 is used to obtain a page design drawing, segment the page design drawing into a plurality of sub-areas, and determine position structure information between the sub-areas, wherein the number of page elements contained in the sub-areas is less than a preset number threshold, and the position structure information is used to characterize the relative position relationship between the sub-areas;
[0085] The code generation module 102 is used to analyze the sub-region using the first model and generate static code corresponding to the sub-region, wherein the first model is obtained by training the large model using multimodal data, and the static code can restore the page layout of the sub-region after rendering, and each sub-region corresponds to a static code;
[0086] The page splicing module 104 is used to splice the static code according to the position structure information to obtain the front-end code file corresponding to the page design drawing.
[0087] In some embodiments of the present application, dividing a page design drawing into multiple sub-areas includes: using an image segmentation model to determine all page elements contained in the page design drawing, and converting the page elements into a first mask in the form of a mask; filtering the first mask to obtain a second mask, wherein the filtering process is used to screen out interference masks in the first mask that are not related to the page structure; based on the second mask, determining a boundary area in the page design drawing, and determining the center line of the boundary area as a candidate boundary line, wherein the boundary area is a rectangular area in the page design drawing that does not intersect with any second mask and runs through the page design drawing in the horizontal or vertical direction; in response to a control instruction input from the front-end interactive interface, determining a target boundary line among the candidate boundary lines, and dividing the page design drawing according to the target boundary line to obtain sub-areas, wherein the control instruction is used to indicate the candidate boundary line selected by the user.
[0088] In some embodiments of the present application, filtering the first mask to obtain the second mask includes: determining the size data of the first mask, and determining the first mask whose size data exceeds a preset size threshold range as an interference mask and screening out the first mask; determining the first mask containing a hole area inside the mask as an interference mask and screening out the first mask; determining the number of colors contained in the area corresponding to the first mask in the page design diagram, and determining the first mask with a color number of one as an interference mask and screening out the first mask.
[0089] In some embodiments of the present application, each candidate dividing line corresponds to a segmentation scheme, and the segmentation scheme is used to divide the page design according to the candidate dividing line. Figure 1 Divided into two parts; before determining the target dividing line among the candidate dividing lines, it also includes: determining the border length of the dividing area corresponding to the candidate dividing line in the non-penetrating direction, wherein the penetrating direction is the horizontal direction or the vertical direction of the page design drawing; determining the vertical distance between the candidate dividing line and the page reference point, wherein the page reference point includes: the vertex in the upper left corner of the page design drawing; determining the scheme score of the segmentation scheme corresponding to the selected dividing line based on the border length and the vertical distance, wherein the border length is positively correlated with the scheme score, and the vertical distance is negatively correlated with the scheme score; sending the candidate dividing lines corresponding to a preset number of segmentation schemes with the highest scheme scores to the front-end interactive interface for display.
[0090] In some embodiments of the present application, after obtaining the front-end code file corresponding to the page design drawing, it also includes: rendering based on the front-end code file to obtain the rendering result, and sending the rendering result to the front-end interactive interface for display; using a second model to analyze the modified text, and adjust the code in the front-end code file based on the analysis result, wherein the modified text is the modification requirements for page elements and layout in the rendering result input by the user in the front-end interactive interface.
[0091] In some embodiments of the present application, the training step of the first model includes: obtaining a basic training data set, wherein the basic training data set contains multiple copies of static code data of the target framework type, and the target framework type includes at least one of the following: Flutter framework; rendering the static code data in the basic training data set to obtain a page design drawing corresponding to the static code data, and constructing a first training data set based on the static code data and the page design drawing, wherein the first training data set contains multiple first training samples, and each first training sample contains: static code data, and a page design drawing corresponding to the static code data; using the first training data set to train the basic large model to obtain the first model.
[0092] In some embodiments of the present application, the training step of the second model includes: using the basic large model to modify the static code data in the basic training data set, and generating a modification description text corresponding to the modification, wherein the modification description text is a natural language description of the modification of the static code data; based on the static code data before the modification, the modified static code data, and the modification description text, constructing a second training data set, wherein the second training data set contains multiple second training samples, and each second training sample contains: static code data, modified static code data, and modification description text; using the second training data set to train the basic large model to obtain a second model.
[0093] In some embodiments of the present application, obtaining a basic training data set includes: obtaining project code in an open source code platform of a target framework type, determining the front-end code in the project code for rendering the front-end interface, and removing the interactive logical relationship in the front-end code to obtain static code data; and / or, using a basic large model to randomly generate page development requirements, and generate static code corresponding to the page development requirements to obtain static code data; and / or, obtaining static page code of a non-target framework type, and using the basic large model to rewrite the static page code into static code of the target framework type to obtain static code data.
[0094] It should be noted that the various modules in the above-mentioned front-end code generation device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.
[0095] It should be noted that the front-end code generation device provided in this embodiment can be used to execute Figure 2 The front-end code generation method shown, therefore, the relevant explanations and instructions on the above-mentioned front-end code generation method are also applicable to the embodiments of the present application and will not be repeated here.
[0096] An embodiment of the present application also provides a non-volatile storage medium, the non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the following front-end code generation method by running the computer program: obtaining a page design drawing, dividing the page design drawing into multiple sub-areas, and determining the position structure information between the sub-areas, wherein the number of page elements contained in the sub-areas is less than a preset number threshold, and the position structure information is used to characterize the relative position relationship between the sub-areas; using a first model, analyzing the sub-areas, and generating static codes corresponding to the sub-areas, wherein the first model is obtained by training a large model with multimodal data, and the static code can restore the page layout of the sub-areas after rendering, and each sub-area corresponds to a static code; according to the position structure information, the static code is spliced to obtain a front-end code file corresponding to the page design drawing.
[0097] The embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the front-end code generation method described in each embodiment of the present application: obtaining a page design drawing, dividing the page design drawing into multiple sub-areas, and determining the position structure information between the sub-areas, wherein the number of page elements contained in the sub-areas is less than a preset number threshold, and the position structure information is used to characterize the relative position relationship between the sub-areas; using a first model, analyzing the sub-areas, and generating static codes corresponding to the sub-areas, wherein the first model is obtained by training a large model with multimodal data, and the static code can restore the page layout of the sub-areas after rendering, and each sub-area corresponds to a static code; according to the position structure information, the static code is spliced to obtain a front-end code file corresponding to the page design drawing.
[0098] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0099] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0100] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0101] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0102] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0103] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.
[0104] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A front-end code generation method, characterized in that: include: Acquire a page design diagram, divide the page design diagram into a plurality of sub-regions, and determine position structure information between the sub-regions, wherein the number of page elements contained in the sub-regions is less than a preset number threshold, and the position structure information is used to characterize the relative position relationship between the sub-regions; Using a first model, the sub-region is analyzed to generate a static code corresponding to the sub-region, wherein the first model is obtained by training a large model using multimodal data, and the static code can restore the page layout of the sub-region after rendering, and each sub-region corresponds to a copy of the static code; The static codes are spliced according to the position structure information to obtain a front-end code file corresponding to the page design diagram.
2. The front-end code generation method according to claim 1, characterized in that: Dividing the page design diagram into multiple sub-areas includes: Using an image segmentation model, determining all page elements included in the page design diagram, and converting the page elements into a first mask in the form of a mask; Performing filtering processing on the first mask to obtain a second mask, wherein the filtering processing is used to filter out interference masks that are irrelevant to the page structure in the first mask; Determine a boundary region in the page design according to the second mask, and determine the midline of the boundary region as a candidate boundary line, wherein the boundary region is a rectangular region in the page design that does not intersect with any of the second masks and runs through the page design in a horizontal direction or a vertical direction; In response to a control instruction inputted from the front-end interactive interface, a target dividing line among the candidate dividing lines is determined, and the page design diagram is segmented according to the target dividing line to obtain the sub-areas, wherein the control instruction is used to indicate the candidate dividing line selected by the user.
3. The front-end code generation method according to claim 2, characterized in that: Filtering the first mask to obtain a second mask includes: Determine the size data of the first mask, and determine the first mask whose size data exceeds a preset size threshold range as the interference mask and screen out the interference mask; Determine the first mask containing a hole area inside the mask as the interference mask and filter it out; The number of colors contained in the area corresponding to the first mask in the page design is determined, and the first mask with the number of colors being one is determined as the interference mask and is screened out.
4. The front-end code generation method according to claim 2, characterized in that: Each candidate dividing line corresponds to a segmentation scheme, and the segmentation scheme is used to divide the page design diagram into two according to the candidate dividing line; Before determining the target dividing line among the candidate dividing lines, the method further includes: Determine the border length of the boundary area corresponding to the candidate boundary line in a non-penetrating direction, wherein the penetrating direction is a horizontal direction or a vertical direction of the page design; Determine the vertical distance between the candidate dividing line and a page reference point, wherein the page reference point includes: a vertex at the upper left corner of the page design; Determine, according to the frame length and the vertical distance, a scheme score of the segmentation scheme corresponding to the selected dividing line, wherein the frame length is positively correlated with the scheme score, and the vertical distance is negatively correlated with the scheme score; The candidate dividing lines corresponding to a preset number of the segmentation schemes with the highest scheme scores are sent to the front-end interactive interface for display.
5. The front-end code generation method according to claim 1, characterized in that: After obtaining the front-end code file corresponding to the page design diagram, the method further includes: Rendering is performed based on the front-end code file to obtain a rendering result, and the rendering result is sent to the front-end interactive interface for display; The second model is used to analyze the modification text, and the code in the front-end code file is adjusted based on the analysis results, wherein the modification text is the modification requirements for the page elements and layout in the rendering results input by the user in the front-end interactive interface.
6. The front-end code generation method according to claim 5, characterized in that: The training steps of the first model include: Obtain a basic training data set, wherein the basic training data set includes multiple copies of static code data of a target framework type, and the target framework type includes at least one of the following: a Flutter framework; Rendering the static code data in the basic training data set to obtain a page design corresponding to the static code data, and constructing a first training data set based on the static code data and the page design, wherein the first training data set includes a plurality of first training samples, each of which includes: the static code data and the page design corresponding to the static code data; The first training data set is used to train the basic large model to obtain the first model.
7. The front-end code generation method according to claim 6, characterized in that: The training step of the second model includes: Using the basic large model, modifying the static code data in the basic training data set, and generating a modification description text corresponding to the modification, wherein the modification description text is a natural language description of the modification of the static code data; Based on the static code data before modification, the static code data after modification, and the modification description text, a second training data set is constructed, wherein the second training data set includes a plurality of second training samples, and each of the second training samples includes: the static code data, the static code data after modification, and the modification description text; The second training data set is used to train the basic large model to obtain the second model.
8. The front-end code generation method according to claim 6, characterized in that: Obtaining the basic training data set includes: Obtaining the project code in the open source code platform of the target framework type, determining the front-end code for rendering the front-end interface in the project code, and removing the interactive logic relationship in the front-end code to obtain the static code data; and / or, using the basic large model to randomly generate page development requirements, and generating static codes corresponding to the page development requirements to obtain the static code data; And / or, obtaining a static page code of a type other than the target framework, and rewriting the static page code into a static code of the target framework type using the basic large model to obtain the static code data.
9. A front-end code generation device, characterized in that: include: A design drawing segmentation module, used to obtain a page design drawing, segment the page design drawing into a plurality of sub-areas, and determine position structure information between the sub-areas, wherein the number of page elements contained in the sub-areas is less than a preset number threshold, and the position structure information is used to characterize the relative position relationship between the sub-areas; A code generation module, configured to analyze the sub-region by using a first model, and generate a static code corresponding to the sub-region, wherein the first model is obtained by training a large model by using multimodal data, and the static code can restore the page layout of the sub-region after rendering, and each sub-region corresponds to a copy of the static code; The page splicing module is used to splice the static code according to the position structure information to obtain the front-end code file corresponding to the page design drawing.
10. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the front-end code generation method described in any one of claims 1 to 8 when running.
11. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the front-end code generation method described in any one of claims 1 to 8 by running the computer program.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the front-end code generation method described in any one of claims 1 to 8 are implemented.