Training data generation method and device, equipment, medium and product
By dividing the page into multiple sections and calculating the layout cost value, training data that is closer to the real user interface is generated, which solves the problem of insufficient training data coverage in the existing technology, improves the recognition accuracy of the model and reduces the workload of manual labeling.
Patent Information
- Application Number
- CN202510645608.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies for generating training data cannot cover a variety of scenarios, resulting in insufficient accuracy of deep learning models in user interface analysis, and large workload and low efficiency of manual labeling.
By dividing the page into multiple sections, filling the page components according to the layout properties, and calculating the page cost, the filled page with the optimal layout is selected for rendering to generate training data that is closer to the real user interface.
It improves the diversity and coverage of training data, enhances the recognition accuracy of deep learning models in user interface parsing, and reduces the workload of manual labeling.
Smart Images

Figure CN120611809A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of machine learning and artificial intelligence, and in particular to a method, apparatus, device, medium, and product for generating training data. Background Art
[0002] During front-end website (web) page development, developers need to analyze user interface design drafts and write corresponding target code. Frequent changes to the user interface often make this work tedious and inefficient. Deep learning models can quickly analyze the basic information of user interface components in the design draft. Therefore, a large amount of training data is required to train deep learning models and improve their accuracy.
[0003] Currently, some solutions for generating training data all transform existing user page images to obtain training data for training models. However, the training data obtained using this method cannot cover a variety of scenarios, resulting in inaccurate models trained based on this training data. Summary of the Invention
[0004] The present application provides a training data generation method, apparatus, device, medium and product to improve the comprehensiveness of generated training data, improve the accuracy of training a model based on the training data, and obtain the model.
[0005] In the first aspect, the present application provides a training data generation method, including: obtaining a page and page components, dividing the page into multiple page sections, and determining the page components corresponding to each page section; filling the page components corresponding to each page section into the corresponding page section according to different layout attributes, and determining the page cost value of the filled page obtained after each filling; the page cost value represents the layout matching degree of the page component under the filled page; the layout attribute represents the layout position and layout size of the page component; selecting the filled page corresponding to the smallest page cost value, rendering the filled page, and obtaining a rendered page image; and obtaining training data based on the rendered page image.
[0006] In one possible implementation, the page cost value of the filled page obtained after each filling is determined, including: determining the cost value of each page component in the filled page based on the position information and size information of each page component filled in the filled page; summing the cost values of each page component to obtain the page cost value of the filled page.
[0007] In one possible implementation, a page is divided into multiple page sections, and a page component corresponding to each page section is determined, including: dividing the page into a preset number of page sections based on a block segmentation algorithm; determining the area type of the page section according to the proportion of each page section in the page; and determining, based on the area type of the page section and based on the component style and component category of the page component, a page component that matches the area type of the page section as the page component corresponding to the page section.
[0008] In one possible implementation, the method further includes: determining the number range of page components in each page section; screening and obtaining the page components corresponding to the page section based on the number range of page components in the page section, and filling them into the corresponding page section.
[0009] In one possible implementation, the cost of each page component in the filled page is determined, including: calculating the boundary constraints corresponding to each page panel according to the position coordinates of the center points of each page panel in the filled page and the page range of the page; the boundary constraints represent the layout boundaries of each page panel; based on the boundary constraints and weight coefficients, a weighted summation calculation is performed according to the size of each page component in the filled page, the position coordinates of the center points of each page component, and the minimum distance between the center points of each page component and the layout boundary under the page panel in which it is located, to obtain the cost of each page component in the filled page; the weight coefficient includes the spacing weight of the page component, the size weight of the page component, the distance weight from the page component to the boundary of the page panel in which it is located, and the influence weight of the size of the page component on the boundary of the page panel in which it is located.
[0010] In one possible implementation, before rendering the filled page to obtain the rendered page image, it also includes: comparing the page cost value of the filled page with a preset threshold; rendering the filled page to obtain the rendered page image, specifically including: if the page cost value is not greater than the preset threshold, rendering the filled page to obtain the rendered page image.
[0011] In a possible implementation, the method further includes: if the page cost value is greater than a preset threshold, stopping page rendering for the filled page.
[0012] In one possible implementation, the fill page is rendered to obtain a rendered page image, including: defining a component style of each page component in the fill page; the component style represents the color, shape, effect, and font of the page component; saving the component style of each page component in the fill page and the component position of each page component to a configuration file; generating a type script based on the configuration file; and generating a renderable page according to the type script; rendering the renderable page to generate a rendered page image.
[0013] In one possible implementation, training data is obtained based on the rendered page image, including: obtaining pixel coordinates of each page component based on the rendered page image; performing data annotation on the rendered image based on the pixel coordinates of each page component to obtain annotation data corresponding to the rendered image; and performing data enhancement based on the rendered image and the annotation data corresponding to the rendered image to obtain training data.
[0014] In one possible implementation, the pixel coordinates of each page component are obtained based on the rendered page image, including: generating the pixel coordinates of the page component under the actual saved image based on the position coordinates and size of each page component under the rendered page image and the scaling ratio; the scaling ratio represents the scaling ratio between the rendered image and the actual saved image.
[0015] In one possible implementation, the training data is used to train a user interface parsing model; the user interface parsing model is used to parse the components and layout in the user's page design draft and convert them into front-end code and / or structured data, so that the user can develop a user interface based on the front-end code and / or structured data.
[0016] In the second aspect, the present application provides a training data generation device, including: an acquisition module, used to acquire pages and page components, divide the page into multiple page sections, and determine the page components corresponding to each page section; a filling module, used to fill the page components corresponding to each page area into the corresponding page area according to different layout attributes, and determine the page cost value of the filled page obtained after each filling; the page cost value represents the layout matching degree of the page component under the filled page; the layout attribute represents the layout position and layout size of the page component; a generation module, used to select the filling page corresponding to the minimum page cost value, render the filling page, and obtain the rendered page image; and obtain training data based on the rendered page image.
[0017] In a third aspect, the present application provides an electronic device comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; and the processor executes the computer-executable instructions stored in the memory to implement the above method.
[0018] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the above method.
[0019] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which is used to implement the above method when executed by a processor.
[0020] The training data generation method, device, equipment, medium and product provided by the present application are as follows: a blank page and a page component to be filled into the page are obtained according to the interface size of a conventional user interface; the obtained page is divided into multiple page sections, and the page components corresponding to each page section are determined; the page components are filled into the corresponding page sections according to different layout attributes, and the page cost value of the filled page obtained after each filling is determined; the page cost value represents the layout matching degree of the page component under the filled page; the filled page corresponding to the minimum page cost value is selected, and the layout of the filled page is considered to be optimal at this time; the filled page is rendered to obtain a rendered page image; and training data is obtained based on the rendered page image. In the scheme of the present application, the page image generated after rendering is closer to the real user interface, and the training data obtained based on the page image is more diversified and can cover more scenarios; the deep learning model for user interface analysis is trained based on the obtained training data, which can improve the recognition accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0022] Figure 1 A flowchart of a training data generation method is exemplified;
[0023] Figure 2 A flowchart of a training data generation method is exemplified;
[0024] Figure 3 This is a schematic diagram of page section segmentation in an example of this application;
[0025] Figure 4 A flowchart of a training data generation method is exemplified;
[0026] Figure 5 This is a flow chart of obtaining a fill page cost value according to an example of the present application;
[0027] Figure 6 This is a flow chart of generating a fill-in page according to an example of this application;
[0028] Figure 7 This is a flow chart of generating a rendered page image according to an example of the present application;
[0029] Figure 8 A schematic diagram of a process for generating training data is shown as an example;
[0030] Figure 9 A schematic diagram of the process of generating training data for an example of this application;
[0031] Figure 10 The following is a schematic diagram showing the structure of a training data generating device;
[0032] Figure 11 Schematic diagram of the structure of an electronic device is shown in FIG.
[0033] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0034] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0035] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings. The terms "first", "second", etc. in the specification and claims and the above-mentioned drawings in this application are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or precedence, unless otherwise indicated. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, for example, they can be implemented in an order other than those given in the diagrams or descriptions of the embodiments of this application. The terms "including" and "having" in the specification and claims and the above-mentioned drawings in this application and any of their variations are intended to cover but not exclusively include, for example, a product or device that includes a series of components is not necessarily limited to those components clearly listed, but may include other components that are not clearly listed or inherent to these products or devices. The term "module" used in this application refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or a combination of hardware or / and software code that can perform functions related to the component.
[0036] During front-end website (web) page development, developers need to analyze the user interface (UI) design draft and write the corresponding target code. In scenarios involving frequent UI changes, this work often becomes tedious and inefficient. Deep learning models can quickly analyze the basic information of UI components in the design draft (such as category, location, and size). However, the training process of deep learning models faces the following three problems: First, due to the personalization of page components and the privatization of design, there are very few public component datasets available for model training; second, UI components are diverse, and the model requires a large dataset for training; third, manually annotating datasets is a huge workload and inefficient.
[0037] At present, most of the solutions for generating training data (UI component datasets) use manual capture of web page images, and then manually annotate the information of page components through image annotation (LabelImg) tools. Other solutions use image processing methods to generate single component images, and generate different component data by enlarging or reducing the images. However, the above solutions have the following problems: (1) Manual sampling requires a lot of time and energy, and the periodicity is particularly long. The amount of data obtained is generally small, so the model is mostly overfitted when the data amount is small. (2) The model is trained on a randomly sampled page dataset, and the recognition accuracy is greatly reduced in personalized scenarios. (3) The component dataset obtained by a single image processing method does not have page layout information, so it has too much difference in the real page detection task, and the model recognition accuracy is extremely low.
[0038] The technical content provided by this application is intended to solve some technical problems of related technologies such as those mentioned above. In the training data generation method, device, equipment, medium and product provided by this application, a blank page and a page component to be filled into the page are obtained according to the interface size of a conventional user interface; the obtained page is divided into multiple page sections, and the page components corresponding to each page section are determined; according to different layout attributes, the page components are filled into the corresponding page sections, and the page cost value of the filled page obtained after each filling is determined; the page cost value represents the layout matching degree of the page component under the filled page; the filled page corresponding to the minimum page cost value is selected, and the layout of the filled page is considered to be optimal at this time; the filled page is rendered to obtain a rendered page image; and training data is obtained based on the rendered page image. In the scheme of this application, the page image generated after rendering is closer to the real user interface, and the training data obtained based on the page image can better cover a variety of scenarios; the deep learning model for user interface analysis is trained based on the obtained training data, which can improve the recognition accuracy of the model.
[0039] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0040] Figure 1 A flow chart of a training data generation method is shown as an example; Figure 1 As shown, the method includes:
[0041] Step 101: Obtain a page and page components, divide the page into multiple page sections, and determine the page components corresponding to each page section;
[0042] Step 102: Fill the corresponding page components of each page section into the corresponding page section according to different layout attributes, and determine the page cost value of the filled page obtained after each filling; the page cost value represents the layout matching degree of the page component on the filled page; the layout attribute represents the layout position and layout size of the page component;
[0043] Step 103: Select a filling page corresponding to the minimum page cost value, render the filling page to obtain a rendered page image; and obtain training data based on the rendered page image.
[0044] Specifically, first, based on the spatial distribution characteristics of the front-end page design, obtain the page and page components; divide the obtained page into multiple page sections, and determine the page components corresponding to each page section. Exemplarily, multiple page sections usually include: a top menu section, a middle main section, and a bottom tool section. Page components usually include basic interactive components: such as buttons, input boxes, selectors, etc.; navigation components: such as menus, tabs, pagers, sidebars, bottom navigation bars, etc.; information display components: such as cards, lists, tables, labels, progress bars, prompt boxes, notification pop-ups, etc.; feedback components: loading indicators, warning boxes, etc.; container components: folding panels, step bars, dividing lines, etc.; multimedia components: pictures, video players, carousels, etc.; advanced interactive components: drag components, tree controls, maps, charts, etc.; and other special components: color pickers, rich text editors, code editors, etc. In actual applications, different page sections need to be filled with corresponding page components; for example, the functions of the top menu section are main navigation, brand display, global operations, user login status, etc.; the corresponding page components include: Navigation category: main navigation menu, tabs (for content category switching); Operation category: search box, notification icon, etc. The functions of the middle main section are core content display, interactive operations, data presentation, etc.; the corresponding page components include: Content display category: cards, tables, lists, charts, etc.; Interaction category: forms, button groups, paginators, etc.; Dynamic content category: carousels, video / picture display areas, comment areas / bullet screens, etc. The functions of the bottom tool section are auxiliary navigation, copyright information, and quick operations (common on mobile devices); the corresponding page components include: Navigation category: footer links, back to top buttons, etc.; Information category: copyright statement, contact information, social media icons, etc.; Operation category (common on mobile devices): bottom fixed toolbar, floating action button, etc.
[0045] After determining the page components corresponding to each page section, fill the page components corresponding to the page section into the corresponding page section according to different layout attributes, and determine the page cost value of the filled page obtained after each filling; the page cost value represents the layout matching degree of the page component under the filled page; the smaller the page cost value, the higher the layout matching degree of the page component under the filled page, and vice versa. For example, different layout attributes represent different sizes and different positions of page components; a certain number of page components in each page section are filled into the corresponding page section according to different layout attributes. Determine the page cost value of the filled page obtained after each filling of the page components; select the filled page corresponding to the smallest page cost value, then it is considered that the layout matching degree of the page components in the filled page under the filled page is the highest, and the filled page is the optimal result. Afterwards, the filled page is rendered to obtain a rendered page image; and training data is obtained based on the rendered page image. In this example, based on the actual UI interface diagram, the blank page is divided into multiple page sections, and the page components corresponding to each page section are filled into the corresponding page section according to different layout attributes; the filling page with the smallest page cost value is selected from the filling pages obtained after each filling, and image rendering is performed to obtain a rendered page image; the rendered page image obtained is closer to the UI interface diagram in the real scene; then, training data is obtained based on the rendered page image, so that the training data is more in line with the real scene.
[0046] Optional, Figure 2 The flowchart of a training data generation method is shown as an example; a page is divided into multiple page sections, and the page components corresponding to each page section are determined, including:
[0047] Step 201: Divide the page into a preset number of page sections based on a block segmentation algorithm;
[0048] Step 202: Determine the area type of each page section based on the proportion of each page section in the page;
[0049] Step 203: According to the area type of the page section, based on the component style and component category of the page component, determine the page component that matches the area type of the page section as the page component corresponding to the page section.
[0050] Specifically, the acquired page is divided into a preset number of page sections based on a block segmentation algorithm; exemplarily, the block segmentation algorithm is a Voronoi diagram (Voronoi) segmentation algorithm; the Voronoi segmentation algorithm is a geometric algorithm based on space division, and its core idea is to divide the space into several regions based on a given set of seed points (or "sites"), and the distance from all points in each region to the seed point of the region is less than the distance to any other seed point. This division result is called a Voronoi diagram, and each region is called a Voronoi cell. The range of the number of page sections is determined based on the experience of web front-end page development; a value m is randomly selected from this range as the number of seed points, and based on the "scan line" idea, an "event queue" that changes continuously according to the scan line position is maintained, and finally the acquired page is divided into m page sections; the number of selected seed points determines the number of page sections to be divided. Figure 3 This is a schematic diagram of page section segmentation in an example of this application; Figure 3 The example in the figure shows a page divided into three sections when m = 3, with the black dots in the figure representing the center points of the sections. The Voronoi segmentation algorithm is primarily based on the spatial distribution characteristics of front-end page design. Unlike random arrangements, front-end pages typically exhibit a block-like distribution, and the geometric forms of page components often appear as irregular polygons. Based on this characteristic, the Voronoi segmentation algorithm can divide the page plane into several irregular polygonal regions based on a preset number of seed points. This division method not only conforms to the geometric characteristics of the actual page layout but, more importantly, effectively avoids the layout pattern of the generated dataset from becoming homogenous, thereby preventing the model from learning invalid features and improving the generalization ability of model training.
[0051] Furthermore, the regional type of the page section is determined based on the proportion of each page section in the entire page. Exemplarily, the regional types of the page section include the aforementioned three categories: the top menu section, the middle main section, and the bottom tool section. In actual applications, the top menu section usually accounts for about 10%, the middle main section usually accounts for about 75%, and the bottom tool section usually accounts for about 15%. The regional type of each page section divided according to the block segmentation algorithm is determined according to the proportion of the entire page occupied. According to the regional type of each page section, based on the component style and component category of the page component, the page component that matches the regional type of the page section is determined as the page component corresponding to the page section. In this example, the entire page is segmented based on the block segmentation algorithm, and then the page component corresponding to each page section is determined according to the regional type of the page section. This is more in line with the layout characteristics of the page section of the actual UI interface, avoids the homogeneity of the layout pattern of the generated training data, and improves the diversity of the subsequent training data generation; when using more diverse training data for model training, the generalization ability of the model can be improved.
[0052] Optionally, the method further includes:
[0053] Determine the range of page components in each page section;
[0054] According to the number range of page components in the page section, filter and obtain the page components corresponding to the page section, and fill them into the corresponding page section.
[0055] Specifically, based on the page section range of each page section (the height and width of the page and the boundaries of the page section), the range of the number of components that can be filled is determined. For example, the top menu section can be filled with 3-5 page components, the middle main section can be filled with 10-12 page components, and the bottom tool section can be filled with 6-8 page components. The target number of page components to be filled is selected within the range of the number of page components that can be filled in each page section. For example, 4 page components can be filled in the top menu section, 11 page components can be filled in the middle main section, and 7 page components can be filled in the bottom tool section; or, 5 page components can be filled in the top menu section, 10 page components can be filled in the middle main section, and 6 page components can be filled in the bottom tool section. Furthermore, based on the determined number of page components for each page section, the page components corresponding to the page section are filtered; for example, based on the range of the number of page components for each page section, the page components are filtered in the component library according to the type of page component, the page components under each page section are determined, and the filtered page components are filled into the corresponding page section. The number of page components under each page section cannot be too many, and the size cannot be too large, otherwise the distance between page components will be too close, and the page components will overlap, making the subsequently generated training data inconsistent with the real scene. At the same time, the number of page components under each page section cannot be too small to prevent the generated training data from being inconsistent with the actual UI interface scene. In this example, the number of page components for each page section is determined, and the page components are filtered according to the determined number of page components for each page section; then, the filtered page components are filled into the corresponding page sections according to different layout attributes, thereby improving the accuracy of generating the filled page.
[0056] Optional, Figure 4 A flow chart of a training data generation method is exemplified; determining a page cost value of a filled page obtained after each filling includes:
[0057] Step 301: Determine the cost of each page component in the filled page based on the position information and size information of each page component in the filled page;
[0058] Step 302: Sum the cost values of the various page components to obtain the page cost value of the filled page.
[0059] Specifically, to determine the page cost value of the filled page obtained after each filling, first determine the cost value of each page component in the filled page based on the position information and size information of each page component in the filled page obtained after each filling; obtain the cost value of each page component in the entire filled page based on the position information of each page component under each page section in the filled page, and the size (width and height) of each page component; sum up the cost value of each page component to obtain the page cost value of the entire filled page. For example, the page cost value of the filled page is C, and the cost value of each page component is C i ,but
[0060]
[0061] Here, i represents any page component, and n represents the number of page components. In this example, the cost value of the entire filled page is obtained by calculating the sum of the cost values of all page components in the filled page, which improves the accuracy of the page cost calculation.
[0062] Optionally, determine the cost of each page component in the populated page, including:
[0063] Calculate the boundary constraints corresponding to each page section based on the center point position coordinates of each page section in the filled page and the page range of the filled page; the boundary constraints represent the layout boundaries of each page section;
[0064] Based on the boundary constraints and weight coefficients, a weighted summation calculation is performed according to the size of each page component in the filled page, the position coordinates of the center point of each page component, and the minimum distance between the center point of each page component and the layout boundary under the page block in which it is located, to obtain the cost value of each page component in the filled page; the weight coefficient includes the spacing weight of the page component, the size weight of the page component, the distance weight from the page component to the boundary of the page block in which it is located, and the influence weight of the size of the page component on the boundary of the page block in which it is located.
[0065] Specifically, to determine the cost of each page component within the infill page, first, based on the center coordinates of each page section and the entire infill page's bounds (the infill page's boundaries), calculate the boundary constraints for each page section. The boundary constraints represent the layout boundaries of each page section. After determining the boundaries of each page section, calculate the cost of each page component based on the boundary constraints and weight coefficients.
[0066] For each page component, based on the boundary constraints and weight coefficients obtained above, according to the size (height and width) of each page component; according to the position coordinates of the center point of each page component, the distance between each page component and other page components in the page block where it is located is obtained; and according to the minimum distance from each page component to the boundary of the page block where it is located, a weighted summation calculation is performed to obtain the cost value of each page component. Exemplarily, the weight coefficient includes the spacing weight (α) of the page component, the size weight (β) of the page component, the distance weight (γ) from the page component to the boundary of the page block where the page component is located, and the influence weight (δ) of the size of the page component on the boundary of the page block where it is located. The calculation formula for the cost value of each page component is:
[0067]
[0068] Among them, C i represents the cost value of each page component in the filled page, i represents any current page component, j represents any page component in the page block where the current page component is located, N(i) represents the set of other page components in the page block where the current page component is located, α is the spacing weight of the aforementioned page components, β is the size weight of the aforementioned page components, γ is the distance weight from the page component to the boundary of the page block where the page component is located, δ is the influence weight of the size of the page component on the boundary of the page block where it is located, (x i ,y i ) represents the position coordinates of the center point of the current page component, (x j ,y j ) represents the position coordinates of the center point of any page component in the page section where the current page component is located, w i Indicates the width of the current page component, h i Indicates the height of the current page component, w j Indicates the width of any page component in the page section where the current page component is located. j Indicates the height of any page component in the page section where the current page component is located. border,i Indicates the minimum distance between the current page component and the border of the page section it is in. In this example, by calculating the borders of each page section that fills the page, and calculating the cost of each page component based on the boundary conditions and weight coefficients, the accuracy of the page component cost calculation is improved.
[0069] Combined with the above example, according to the number of page components corresponding to each page section, a component grid is constructed under each page section, and the component grid is used to fill the page components. Traverse each component grid and calculate the cost of the page components under the component grid. When the component grid, that is, the page components under the component grid, exceed the boundary of the page section, the cost of the page components under the component grid is increased; when the component grid, that is, the page components under the component grid, does not exceed the boundary of the page section, the cost of the page components under the component grid is reduced. Subsequently, the cost of the entire filled page is calculated based on the calculated cost of each component grid. Figure 5 FIG. 1 is a flow chart of obtaining a fill page cost value according to an example of the present application; Figure 5 As shown, the page components under each component grid are first traversed. Based on the calculated boundary constraints, it is determined whether the page components exceed the boundary constraints. If the boundary constraints are exceeded, the corresponding cost value increases; if the boundary constraints are not exceeded, the corresponding cost value decreases. The page cost of the filled page is then obtained by taking the weighted sum of the costs of each page component in the filled page.
[0070] Optionally, before rendering the filled page to obtain the rendered page image, the process further includes:
[0071] Compare the page cost of the filled page with a preset threshold;
[0072] Rendering the filled page to obtain a rendered page image specifically includes:
[0073] If the page cost is not greater than a preset threshold, the filled page is rendered to obtain a rendered page image.
[0074] Specifically, after determining the filled page based on the minimum cost value of the filled page obtained after each filling, the filled page is rendered. Before obtaining the rendered page image, the page cost value of the determined filled page needs to be compared with a preset threshold. If the page cost value of the filled page is not greater than the preset threshold, the filled page is rendered to obtain a rendered page image. Exemplarily, the cost value of the filled page is compared with the preset threshold. When the page cost value of the filled page is not greater than the preset threshold, the filled page is considered to be legally filled. At this time, the number, size, and spacing of the page components filled in the filled page are highly matched with the filled page. In this example, before performing image rendering on the obtained filled page to obtain the rendered page image, the legitimacy of the filled page is verified according to the cost value of the filled page, thereby further improving the accuracy of generating the rendered page image.
[0075] Optionally, the method further includes:
[0076] If the page cost is greater than a preset threshold, page rendering for the filled page is stopped.
[0077] Specifically, when the page cost value of the obtained filling page is compared with the preset threshold, and it is found that the page cost value of the filling page is greater than the preset threshold, it is considered that the filling page is illegally filled and the layout of the filling page is unreasonable. At this time, it is necessary to re-segment the page to obtain the corresponding multiple page sections and determine the page components under each page section, and re-fill the page sections with page components to obtain a new filling page. The page cost value of the new filling page is calculated again. When the page cost value of the new filling page is not greater than the preset threshold, the image is rendered again according to the new filling page to obtain a rendered image. If it is determined that the new filling page is still illegally filled, the step of re-segmenting the page is repeated again. Exemplarily, Figure 6 This is a flow chart of generating a fill page for an example of this application; Figure 6 As shown, first obtain the page; split the page into multiple page sections; determine the number of page components under each page section, filter the page components according to the number of page components under each page section, and determine the page components under each page section; fill the page components under each page section into the corresponding page section according to different layout attributes, and obtain the corresponding filled page after each filling. Calculate the page cost value of the filled page obtained after each filling, and select the minimum page cost value in the filled page for comparison with the preset threshold. If the page cost value is not greater than the preset threshold, the filled page corresponding to the minimum page cost value is used as the determined filled page, and the filled page is subsequently rendered to obtain the rendered page image; if the minimum page cost value in the filled page is greater than the preset threshold, the page is re-segmented. In this example, the page cost of the selected filling page with the smallest page cost is compared with the preset threshold to ensure that the obtained filling page is legally filled, and then the filling page is rendered. If the page cost of the selected filling page with the smallest page cost is greater than the preset threshold, the image rendering of the filling page is stopped, and the page is re-divided so that the subsequent generated training data is more in line with the real scenario.
[0078] Optionally, rendering the filled page to obtain a rendered page image includes:
[0079] Define the component style of each page component in the populated page; the component style represents the color, shape, effect, and font of the page component;
[0080] Save the component style and component position of each page component in the populated page to the configuration file;
[0081] Generate type scripts based on configuration files; and generate renderable pages based on type scripts;
[0082] Render the renderable page to generate a rendered page image.
[0083] Specifically, according to each page component in the filled page, based on the personalized style of the page component, the component style of each page component is customized; the component style represents the color, shape, effect, and font of the page component; and the relevant parameters are saved. For example, the relevant parameters are saved in a configuration file, and the configuration file can generate a type script (TypeScript); based on the generated TypeScript script, a renderable page is generated; the renderable page is rendered based on the image to generate a rendered page image. Figure 7 This is a flow chart of generating a rendered page image according to an example of this application; Figure 7 As shown, a determined filling page is obtained; based on each page component in the filling page, the component style of each page component in the filling page is defined; the component style information of each page component in the filling page and the location information of each page component are saved in a configuration file; based on the configuration file, a type script is generated, and based on the type script, a renderable page is generated; the renderable page is rendered to generate a rendered page image. In this example, based on each page component in the determined filling page, the component style of each page component is customized, and the component style of the defined page component and the location information of the page component are saved in a configuration file; the configuration file generates a TypeScript script; based on the TypeScript script, image rendering is performed to obtain a rendered image; the component style is customized to make the image closer to the real user interface; the subsequent training data obtained based on the rendered page image is closer to the real user interface.
[0084] Optional, Figure 8 The following is a flow chart showing an exemplary process for generating training data. The training data is obtained based on the rendered page image, including:
[0085] Step 401: Obtain pixel coordinates of each page component based on the rendered page image;
[0086] Step 402: annotate the rendered image according to the pixel coordinates of each page component to obtain the annotated data corresponding to the rendered image.
[0087] Step 403: Perform data enhancement based on the rendered image and the annotation data corresponding to the rendered image to obtain training data.
[0088] Specifically, based on the rendered page image, the pixel coordinates of each page component in the rendered page image are obtained. Based on the pixel coordinates of each page component, the rendered image is data-labeled to obtain the labeled data corresponding to the rendered image. Data enhancement is performed based on the rendered image and the labeled data corresponding to the rendered image to obtain training data. Exemplarily, after outputting a rendered image and obtaining the labeled data corresponding to the rendered image, in order to increase the diversity of the training data and improve the generalization ability of the model, it is necessary to perform data enhancement based on the rendered image and its corresponding labeled data to obtain a large amount of highly diverse training data. Among them, data enhancement includes rotation: rotating the image with data annotations, specifically rotating the training image with annotations within a certain angle range (considering that the offset angle of the page component is usually not large, set from -15 degrees to 15 degrees), and at the same time, performing coordinate conversion on the label position coordinates of the labeled data according to the same angle to enhance the directional invariance of the model so that the model can better handle page components at different angles. Translation: The image with data annotation is translated horizontally or vertically (to ensure the integrity of the page, the translation offset is set within 100 pixels) to improve the robustness to changes in the target position. Noise addition: Gaussian noise points are randomly added to the original image to enhance the accuracy of identifying image blur (to ensure that the main component features are not affected, an 11*11 square noise block is set) and prevent model overfitting. Flipping: The image with data annotation is flipped by mirroring to enhance the accuracy of the model in the process of symmetrical component recognition. Cropping: The image with data annotation is cropped (it has been verified that the cropping range is within 30% of the original image to ensure that the main information of the page is not lost). At the same time, the page component labels in the area are filtered out and their coordinates are repositioned to enhance the model's ability to recognize local components when dealing with complex page recognition.
[0089] Optionally, the pixel coordinates of each page component are obtained based on the rendered page image, including:
[0090] According to the position coordinates and size of each page component in the rendered image, based on the scaling ratio, the pixel coordinates of the page component in the actual saved image are generated; the scaling ratio represents the scaling ratio between the rendered image and the actual saved image.
[0091] Specifically, after obtaining the rendered page image, the actual saved image will be scaled compared to the page image obtained after rendering. Therefore, it is necessary to standardize the position and size information of the page components. Specifically, based on the characteristic that the relative position and size ratio of the page components in the layout of the filled page remain unchanged, the position coordinates and size of each page component under the rendered image are converted into pixel coordinates under the actual saved image based on the scaling ratio. Among them, the scaling ratio represents the scaling ratio between the rendered image and the actual saved image. Specifically, this conversion process is achieved by calculating the relative position ratio of the component in the page and mapping it to the pixel coordinate system of the screenshot image. The converted coordinate data will be stored in the label file, and the original image file will be saved at the same time, so as to construct a complete image-annotation data pair to provide standardized input for subsequent model training. This processing method ensures the consistency of component position information under screenshots of different resolutions and increases the generalization of the model to image size, while retaining the spatial relationship characteristics of the original layout.
[0092] Optionally, the training data is used to train a user interface parsing model; the user interface parsing model is used to parse the components and layout in the user's page design draft and convert them into front-end code and / or structured data, so that the user can develop a user interface based on the front-end code and / or structured data.
[0093] Specifically, the training data generated by the present application is used to train the user interface parsing model; the user interface parsing model is used to parse the components and layout in the user's page design draft; and is converted into front-end code or structured data, or front-end code and structured data; so that users can develop user interfaces based on the front-end code and structured data. The training data generated by the present application can be used for direct training of deep learning models and multimodal large models, and can quickly generate a UI component dataset containing annotated data. The style of the page component dataset can be controlled by defining the style of the page component, and can be directly provided to deep learning models and multimodal large models for training, thereby improving the accuracy of model training.
[0094] Combining the above examples, Figure 9 A flow chart showing the generation of training data for an example of this application; Figure 9As shown, first determine the page sections and page components under the page sections in the page; fill the page components into the corresponding page sections according to different layout attributes, calculate the page cost values of each filled page after each filling, and determine the filled page according to the page cost values of each filled page; define the component style of the page components in the filled page, save the relevant parameters, and build an HTML page; render the page to obtain the rendered page image; save the rendered image; obtain the pixel coordinates of each page component, and perform data annotation on the rendered image based on the pixel coordinates to obtain the annotation data corresponding to the rendered image; based on the rendered image and the corresponding annotation data; rotate, translate, crop, mirror, add noise and other data enhancements to generate a training data set.
[0095] The training data generation method provided in this embodiment is as follows: according to the interface size of a conventional user interface, a blank page and the page components to be filled into the page are obtained; the obtained page is divided into multiple page sections, and the page components corresponding to each page section are determined; according to different layout attributes, the page components are filled into the corresponding page sections, and the page cost value of the filled page obtained after each filling is determined; the page cost value represents the layout matching degree of the page components under the filled page; the filled page corresponding to the minimum page cost value is selected, and the layout of the filled page is considered to be optimal at this time; the filled page is rendered to obtain a rendered page image; and training data is obtained based on the rendered page image. In the scheme of the present application, the page image generated after rendering is closer to the real user interface, and the training data obtained based on the page image is more diversified and can cover more scenarios; the deep learning model for user interface analysis is trained based on the obtained training data, which can improve the recognition accuracy of the model.
[0096] Example 2
[0097] Figure 10 A schematic diagram of the structure of a training data generating device is shown as an example. Figure 10 As shown, the device includes:
[0098] The acquisition module 21 is used to acquire a page and page components, divide the page into multiple page sections, and determine the page components corresponding to each page section;
[0099] The filling module 22 is used to fill the page components corresponding to each page section into the corresponding page section according to different layout attributes, and determine the page cost value of the filled page obtained after each filling; the page cost value represents the layout matching degree of the page component under the filled page; the layout attribute represents the layout position and layout size of the page component;
[0100] The generating module 23 is configured to select a filling page corresponding to a minimum page cost, render the filling page to obtain a rendered page image, and obtain training data based on the rendered page image.
[0101] Specifically, first, based on the spatial distribution characteristics of the front-end page design, obtain the page and page components; divide the obtained page into multiple page sections, and determine the page components corresponding to each page section. Exemplarily, multiple page sections usually include: a top menu section, a middle main section, and a bottom tool section. Page components usually include basic interactive components: such as buttons, input boxes, selectors, etc.; navigation components: such as menus, tabs, pagers, sidebars, bottom navigation bars, etc.; information display components: such as cards, lists, tables, labels, progress bars, prompt boxes, notification pop-ups, etc.; feedback components: loading indicators, warning boxes, etc.; container components: folding panels, step bars, dividing lines, etc.; multimedia components: pictures, video players, carousels, etc.; advanced interactive components: drag components, tree controls, maps, charts, etc.; and other special components: color pickers, rich text editors, code editors, etc. In actual applications, different page sections need to be filled with corresponding page components; for example, the functions of the top menu section are main navigation, brand display, global operations, user login status, etc.; the corresponding page components include: Navigation category: main navigation menu, tabs (for content category switching); Operation category: search box, notification icon, etc. The functions of the middle main section are core content display, interactive operations, data presentation, etc.; the corresponding page components include: Content display category: cards, tables, lists, charts, etc.; Interaction category: forms, button groups, paginators, etc.; Dynamic content category: carousels, video / picture display areas, comment areas / bullet screens, etc. The functions of the bottom tool section are auxiliary navigation, copyright information, and quick operations (common on mobile devices); the corresponding page components include: Navigation category: footer links, back to top buttons, etc.; Information category: copyright statement, contact information, social media icons, etc.; Operation category (common on mobile devices): bottom fixed toolbar, floating action button, etc.
[0102] After determining the page components corresponding to each page section, fill the page components corresponding to the page section into the corresponding page section according to different layout attributes, and determine the page cost value of the filled page obtained after each filling; the page cost value represents the layout matching degree of the page component under the filled page; the smaller the page cost value, the higher the layout matching degree of the page component under the filled page, and vice versa. For example, different layout attributes represent different sizes and different positions of page components; a certain number of page components in each page section are filled into the corresponding page section according to different layout attributes. Determine the page cost value of the filled page obtained after each filling of the page components; select the filled page corresponding to the smallest page cost value, then it is considered that the layout matching degree of the page components in the filled page under the filled page is the highest, and the filled page is the optimal result. Afterwards, the filled page is rendered to obtain a rendered page image; and training data is obtained based on the rendered page image. In this example, based on the actual UI interface diagram, the blank page is divided into multiple page sections, and the page components corresponding to each page section are filled into the corresponding page section according to different layout attributes; the filling page with the smallest page cost value is selected from the filling pages obtained after each filling, and image rendering is performed to obtain a rendered page image; the rendered page image obtained is closer to the UI interface diagram in the real scene; then, training data is obtained based on the rendered page image, so that the training data is more in line with the real scene.
[0103] Optionally, the acquisition module 21 is configured to:
[0104] Split the page into a preset number of page sections based on the block segmentation algorithm;
[0105] Determine the area type of each page section based on the proportion of each page section in the page;
[0106] According to the area type of the page section, based on the component style and component category of the page component, a page component that matches the area type of the page section is determined as the page component corresponding to the page section.
[0107] Specifically, the acquired page is divided into a preset number of page sections based on a block segmentation algorithm; exemplarily, the block segmentation algorithm is a Voronoi diagram (Voronoi) segmentation algorithm; the Voronoi segmentation algorithm is a geometric algorithm based on space division, and its core idea is to divide the space into several regions based on a given set of seed points (or "sites"), and the distance from all points in each region to the seed point of the region is less than the distance to any other seed point. This division result is called a Voronoi diagram, and each region is called a Voronoi cell. The range of the number of page sections is determined based on the experience of web front-end page development; a value m is randomly selected from this range as the number of seed points, and based on the "scan line" idea, an "event queue" that changes continuously according to the scan line position is maintained, and finally the acquired page is divided into m page sections; the number of selected seed points determines the number of page sections to be divided.
[0108] The Voronoi segmentation algorithm is primarily based on the spatial distribution characteristics of front-end page designs. Unlike random arrangements, front-end pages typically exhibit a block-like distribution, and the geometric forms of page components often appear as irregular polygons. Based on this characteristic, the Voronoi segmentation algorithm can divide the page plane into several irregular polygonal regions based on a preset number of seed points. This division method not only conforms to the geometric characteristics of the actual page layout, but more importantly, it effectively prevents the layout pattern of the generated dataset from becoming homogenous, thereby preventing the model from learning invalid features and improving the generalization ability of model training.
[0109] Furthermore, the regional type of the page section is determined based on the proportion of each page section in the entire page. Exemplarily, the regional types of the page section include the aforementioned three categories: the top menu section, the middle main section, and the bottom tool section. In actual applications, the top menu section usually accounts for about 10%, the middle main section usually accounts for about 75%, and the bottom tool section usually accounts for about 15%. The regional type of each page section divided according to the block segmentation algorithm is determined according to the proportion of the entire page occupied. According to the regional type of each page section, based on the component style and component category of the page component, the page component that matches the regional type of the page section is determined as the page component corresponding to the page section. In this example, the entire page is segmented based on the block segmentation algorithm, and then the page component corresponding to each page section is determined according to the regional type of the page section. This is more in line with the layout characteristics of the page section of the actual UI interface, avoids the homogeneity of the layout pattern of the generated training data, and improves the diversity of the subsequent training data generation; when using more diverse training data for model training, the generalization ability of the model can be improved.
[0110] Optionally, the device further includes a screening module 24, which is configured to:
[0111] Determine the range of page components in each page section;
[0112] According to the number range of page components in the page section, filter and obtain the page components corresponding to the page section, and fill them into the corresponding page section.
[0113] Specifically, based on the page section range of each page section (the height and width of the page and the boundaries of the page section), the range of the number of components that can be filled is determined. For example, the top menu section can be filled with 3-5 page components, the middle main section can be filled with 10-12 page components, and the bottom tool section can be filled with 6-8 page components. The target number of page components to be filled is selected within the range of the number of page components that can be filled in each page section. For example, 4 page components can be filled in the top menu section, 11 page components can be filled in the middle main section, and 7 page components can be filled in the bottom tool section; or, 5 page components can be filled in the top menu section, 10 page components can be filled in the middle main section, and 6 page components can be filled in the bottom tool section. Furthermore, based on the determined number of page components for each page section, the page components corresponding to the page section are filtered; for example, based on the range of the number of page components for each page section, the page components are filtered in the component library according to the type of page component, the page components under each page section are determined, and the filtered page components are filled into the corresponding page section. The number of page components under each page section cannot be too many, and the size cannot be too large, otherwise the distance between page components will be too close, and the page components will overlap, making the subsequently generated training data inconsistent with the real scene. At the same time, the number of page components under each page section cannot be too small to prevent the generated training data from being inconsistent with the actual UI interface scene. In this example, the number of page components for each page section is determined, and the page components are filtered according to the determined number of page components for each page section; then, the filtered page components are filled into the corresponding page sections according to different layout attributes, thereby improving the accuracy of generating the filled page.
[0114] Optionally, a filling module 22 is used to:
[0115] Determine the cost of each page component in the filled page according to the position information and size information of each page component filled in the filled page;
[0116] Sum the cost values of each page component to get the page cost value of the filled page.
[0117] Specifically, to determine the page cost value of the filled page obtained after each filling, first determine the cost value of each page component in the filled page based on the position information and size information of each page component in the filled page obtained after each filling; obtain the cost value of each page component in the entire filled page based on the position information of each page component under each page section in the filled page, and the size (width and height) of each page component; sum up the cost value of each page component to obtain the page cost value of the entire filled page. For example, the page cost value of the filled page is C, and the cost value of each page component is C i ,but
[0118]
[0119] Here, i represents any page component, and n represents the number of page components. In this example, the cost value of the entire filled page is obtained by calculating the sum of the cost values of all page components in the filled page, which improves the accuracy of the page cost calculation.
[0120] Optionally, the filling module 22 is specifically used to:
[0121] Calculate the boundary constraints corresponding to each page section based on the center point position coordinates of each page section in the filled page and the page range of the filled page; the boundary constraints represent the layout boundaries of each page section;
[0122] Based on the boundary constraints and weight coefficients, a weighted summation calculation is performed according to the size of each page component in the filled page, the position coordinates of the center point of each page component, and the minimum distance between the center point of each page component and the layout boundary under the page block in which it is located, to obtain the cost value of each page component in the filled page; the weight coefficient includes the spacing weight of the page component, the size weight of the page component, the distance weight from the page component to the boundary of the page block in which it is located, and the influence weight of the size of the page component on the boundary of the page block in which it is located.
[0123] Specifically, to determine the cost of each page component within the infill page, first, based on the center coordinates of each page section and the entire infill page's bounds (the infill page's boundaries), calculate the boundary constraints for each page section. The boundary constraints represent the layout boundaries of each page section. After determining the boundaries of each page section, calculate the cost of each page component based on the boundary constraints and weight coefficients.
[0124] For each page component, based on the boundary constraints and weight coefficients obtained above, according to the size (height and width) of each page component; according to the position coordinates of the center point of each page component, the distance between each page component and other page components in the page block where it is located is obtained; and according to the minimum distance from each page component to the boundary of the page block where it is located, a weighted summation calculation is performed to obtain the cost value of each page component. Exemplarily, the weight coefficient includes the spacing weight (α) of the page component, the size weight (β) of the page component, the distance weight (γ) from the page component to the boundary of the page block where the page component is located, and the influence weight (δ) of the size of the page component on the boundary of the page block where it is located. The calculation formula for the cost value of each page component is:
[0125]
[0126] Among them, C i represents the cost value of each page component in the filled page, i represents any current page component, j represents any page component in the page block where the current page component is located, N(i) represents the set of other page components in the page block where the current page component is located, α is the spacing weight of the aforementioned page components, β is the size weight of the aforementioned page components, γ is the distance weight from the page component to the boundary of the page block where the page component is located, δ is the influence weight of the size of the page component on the boundary of the page block where it is located, (x i ,y i ) represents the position coordinates of the center point of the current page component, (x j ,y j ) represents the position coordinates of the center point of any page component in the page section where the current page component is located, w i Indicates the width of the current page component, h i Indicates the height of the current page component, w j Indicates the width of any page component in the page section where the current page component is located. j Indicates the height of any page component in the page section where the current page component is located. border,i Indicates the minimum distance between the current page component and the border of the page section it is in. In this example, by calculating the borders of each page section that fills the page, and calculating the cost of each page component based on the boundary conditions and weight coefficients, the accuracy of the page component cost calculation is improved.
[0127] Optionally, the device further includes a comparison module 25, which is used to:
[0128] Compare the page cost of the filled page with a preset threshold;
[0129] The generation module 23 is specifically used to:
[0130] If the page cost is not greater than a preset threshold, the filled page is rendered to obtain a rendered page image.
[0131] Specifically, after determining the filled page based on the minimum cost value of the filled page obtained after each filling, the filled page is rendered. Before obtaining the rendered page image, the page cost value of the determined filled page needs to be compared with a preset threshold. If the page cost value of the filled page is not greater than the preset threshold, the filled page is rendered to obtain a rendered page image. Exemplarily, the cost value of the filled page is compared with the preset threshold. When the page cost value of the filled page is not greater than the preset threshold, the filled page is considered to be legally filled. At this time, the number, size, and spacing of the page components filled in the filled page are highly matched with the filled page. In this example, before performing image rendering on the obtained filled page to obtain the rendered page image, the legitimacy of the filled page is verified according to the cost value of the filled page, thereby further improving the accuracy of generating the rendered page image.
[0132] Optionally, if the page cost is greater than a preset threshold, page rendering of the filled page is stopped.
[0133] Specifically, when the page cost of the obtained filled page is compared with a preset threshold and the page cost of the filled page is found to be greater than the preset threshold, the filled page is considered to be illegal and the layout of the filled page is unreasonable. At this point, the page needs to be re-segmented to obtain multiple corresponding page sections, determine the page components under each page section, and re-fill the page sections with page components to obtain a new filled page. The page cost of the new filled page is calculated again. If the page cost of the new filled page is not greater than the preset threshold, the image is rendered again based on the new filled page to obtain a rendered image. If the new filled page is still determined to be illegal, the page re-segmentation step is repeated again. In this example, the page cost of the selected filled page with the smallest page cost is compared with the preset threshold to ensure that the obtained filled page is legal, and then the filled page is rendered. If the page cost of the selected filled page with the smallest page cost is greater than the preset threshold, the image rendering of the filled page is stopped and the page is re-segmented, so that the subsequent training data generated is more consistent with real-world scenarios.
[0134] Optionally, the generating module 23 is used to:
[0135] Define the component style of each page component in the populated page; the component style represents the color, shape, effect, and font of the page component;
[0136] Save the component style and component position of each page component in the populated page to the configuration file;
[0137] Generate type scripts based on configuration files; and generate renderable pages based on type scripts;
[0138] Render the renderable page to generate a rendered page image.
[0139] Specifically, according to each page component in the filled page, based on the personalized style of the page component, the component style of each page component is customized; the component style represents the color, shape, effect, and font of the page component; and the relevant parameters are saved. For example, the relevant parameters are saved in a configuration file, and the configuration file can generate a type script (TypeScript); based on the generated TypeScript script, a renderable page is generated; the renderable page is rendered based on the image to generate a rendered page image. In this example, according to each page component in the determined filled page, the component style of each page component is customized, and the component style of the defined page component and the location information of the page component are saved in the configuration file; the configuration file generates a TypeScript script; based on the TypeScript script, image rendering is performed to obtain a rendered image; the style of the component is customized to make the image closer to the real user interface; the training data subsequently obtained based on the rendered page image is closer to the real user interface.
[0140] Optionally, the generating module 23 is used to:
[0141] According to the rendered page image, the pixel coordinates of each page component are obtained;
[0142] According to the pixel coordinates of each page component, the rendered image is annotated with data to obtain the annotated data corresponding to the rendered image;
[0143] Data enhancement is performed based on the rendered image and the labeled data corresponding to the rendered image to obtain training data.
[0144] Specifically, based on the rendered page image, the pixel coordinates of each page component in the rendered page image are obtained. Based on the pixel coordinates of each page component, the rendered image is data-labeled to obtain the labeled data corresponding to the rendered image. Data enhancement is performed based on the rendered image and the labeled data corresponding to the rendered image to obtain training data. Exemplarily, after outputting a rendered image and obtaining the labeled data corresponding to the rendered image, in order to increase the diversity of the training data and improve the generalization ability of the model, it is necessary to perform data enhancement based on the rendered image and its corresponding labeled data to obtain a large amount of highly diverse training data. Among them, data enhancement includes rotation: rotating the image with data annotations, specifically rotating the training image with annotations within a certain angle range (considering that the offset angle of the page component is usually not large, set from -15 degrees to 15 degrees), and at the same time, performing coordinate conversion on the label position coordinates of the labeled data according to the same angle to enhance the directional invariance of the model so that the model can better handle page components at different angles. Translation: The image with data annotation is translated horizontally or vertically (to ensure the integrity of the page, the translation offset is set within 100 pixels) to improve the robustness to changes in the target position. Noise addition: Gaussian noise points are randomly added to the original image to enhance the accuracy of identifying image blur (to ensure that the main component features are not affected, an 11*11 square noise block is set) and prevent model overfitting. Flipping: The image with data annotation is flipped by mirroring to enhance the accuracy of the model in the process of symmetrical component recognition. Cropping: The image with data annotation is cropped (it has been verified that the cropping range is within 30% of the original image to ensure that the main information of the page is not lost). At the same time, the page component labels in the area are filtered out and their coordinates are repositioned to enhance the model's ability to recognize local components when dealing with complex page recognition.
[0145] Optionally, the generating module 23 is used to:
[0146] According to the position coordinates and size of each page component in the rendered image, based on the scaling ratio, the pixel coordinates of the page component in the actual saved image are generated; the scaling ratio represents the scaling ratio between the rendered image and the actual saved image.
[0147] Specifically, after obtaining the rendered page image, the actual saved image will be scaled compared to the page image obtained after rendering. Therefore, it is necessary to standardize the position and size information of the page components. Specifically, based on the characteristic that the relative position and size ratio of the page components in the layout of the filled page remain unchanged, the position coordinates and size of each page component under the rendered image are converted into pixel coordinates under the actual saved image based on the scaling ratio. Among them, the scaling ratio represents the scaling ratio between the rendered image and the actual saved image. Specifically, this conversion process is achieved by calculating the relative position ratio of the component in the page and mapping it to the pixel coordinate system of the screenshot image. The converted coordinate data will be stored in the label file, and the original image file will be saved at the same time, so as to construct a complete image-annotation data pair to provide standardized input for subsequent model training. This processing method ensures the consistency of component position information under screenshots of different resolutions and increases the generalization of the model to image size, while retaining the spatial relationship characteristics of the original layout.
[0148] Optionally, the training data is used to train a user interface parsing model; the user interface parsing model is used to parse the components and layout in the user's page design draft and convert them into front-end code and / or structured data, so that the user can develop a user interface based on the front-end code and / or structured data.
[0149] Specifically, the training data generated by the present application is used to train the user interface parsing model; the user interface parsing model is used to parse the components and layout in the user page design draft; and is converted into front-end code or structured data, or front-end code and structured data; so that users can develop user interfaces based on the front-end code and structured data. The training data generated by the present application can be used for direct training of deep learning models and multimodal large models, and can quickly generate a UI component dataset containing annotated data. The style of the component dataset can be controlled by parameters and can be directly provided to deep learning models and multimodal large models for training, thereby improving the accuracy of model training.
[0150] The training data generation device provided in this embodiment obtains a blank page and the page components to be filled into the page according to the interface size of a conventional user interface; divides the obtained page into multiple page sections and determines the page components corresponding to each page section; fills the page components into the corresponding page sections according to different layout attributes, and determines the page cost value of the filled page obtained after each filling; the page cost value represents the layout matching degree of the page components under the filled page; selects the filled page corresponding to the minimum page cost value, and at this time, it is considered that the layout of the filled page is optimal; renders the filled page to obtain a rendered page image; and obtains training data based on the rendered page image. In the scheme of the present application, the page image generated after rendering is closer to the real user interface, and the training data obtained based on the page image is more diversified and can cover more scenarios; the deep learning model for user interface analysis is trained based on the obtained training data, which can improve the recognition accuracy of the model.
[0151] Example 3
[0152] Figure 11 exemplarily shows a structural diagram of an electronic device, the device comprising:
[0153] The device includes a processor 291 and a memory 292; a communication interface 293, and a bus 294. The processor 291, memory 292, and communication interface 293 can communicate with each other via bus 294. Communication interface 293 can be used for information transmission. Processor 291 can invoke logic instructions in memory 292 to execute the method described above.
[0154] In addition, the logic instructions in the memory 292 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0155] Memory 292, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of the present application. Processor 291 executes the software programs, instructions, and modules stored in memory 292 to execute functional applications and data processing, thereby implementing the methods in the above-mentioned method examples.
[0156] Memory 292 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal device. Memory 292 may also include high-speed random access memory and non-volatile memory.
[0157] An embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method in any embodiment.
[0158] An embodiment of the present application further provides a computer program product, including a computer program, which is used to implement the method in any embodiment when executed by a processor.
[0159] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.
[0160] It should be further noted that, although the various steps in the flowchart are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps may be executed in other orders. Moreover, at least a portion of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed in the same time period, but may be executed in different time periods. The execution order of these sub-steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0161] It should be understood that the above-described device embodiments are merely illustrative, and the device of the present application may also be implemented in other ways. For example, the division of units / modules in the above-described embodiments is merely a logical functional division, and actual implementations may employ other division methods. For example, multiple units, modules, or components may be combined or integrated into another system, or some features may be omitted or not implemented.
[0162] In addition, unless otherwise specified, the functional units / modules in the various embodiments of the present application may be integrated into a single unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The aforementioned integrated units / modules may be implemented in the form of hardware or software program modules.
[0163] If the integrated unit / module is implemented in hardware, the hardware may be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor may be any appropriate hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC. Unless otherwise specified, the storage unit may be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.
[0164] If the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0165] In the above embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any way. To keep the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0166] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only.
[0167] It will be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.
[0168] Finally, it should be noted that other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structure described above and shown in the drawings. Various modifications and variations may be made without departing from the scope of the present invention.
Claims
1. A training data generation method, characterized in that: include: Obtaining a page and page components, dividing the page into multiple page sections, and determining a page component corresponding to each page section; Filling the page components corresponding to each page section into the corresponding page section according to different layout attributes, and determining the page cost value of the filled page obtained after each filling; The page cost value represents the layout matching degree of the page component under the filled page; The layout attributes represent the layout position and layout size of the page component; A filling page corresponding to the minimum page cost is selected, and the filling page is rendered to obtain a rendered page image; and training data is obtained based on the rendered page image.
2. The method according to claim 1, characterized in that Determining the page cost of the filled page obtained after each filling includes: Determining the cost of each page component in the filled page according to the position information and size information of each page component filled in the filled page; The cost values of the various page components are summed to obtain the page cost value of the filled page.
3. The method according to claim 1, characterized in that The dividing the page into a plurality of page sections and determining a page component corresponding to each page section includes: Splitting the page into a preset number of page sections based on a block segmentation algorithm; Determine the area type of each page section based on the proportion of each page section in the page; According to the area type of the page section, based on the component style and component category of the page component, a page component matching the area type of the page section is determined as the page component corresponding to the page section.
4. The method according to claim 3, characterized in that The method further comprises: Determine the range of page components in each page section; According to the number range of page components in the page section, the page components corresponding to the page section are screened and obtained, and filled into the corresponding page section.
5. The method according to claim 4, characterized in that Determining the cost value of each page component in the filled page includes: Calculating boundary constraints corresponding to each page section based on the center point position coordinates of each page section in the filled page and the page range of the page; the boundary constraints represent the layout boundaries of each page section; Based on the boundary constraints and weight coefficients, a weighted summation calculation is performed according to the size of each page component in the filling page, the position coordinates of the center point of each page component, and the minimum distance between the center point of each page component and the layout boundary under the page block where it is located, to obtain the cost value of each page component in the filling page; the weight coefficient includes the spacing weight of the page component, the size weight of the page component, the distance weight from the page component to the boundary of the page block where it is located, and the influence weight of the size of the page component on the boundary of the page block where it is located.
6. The method according to claim 5, characterized in that Before rendering the filled page to obtain the rendered page image, the method further includes: Comparing the page cost of the filled page with a preset threshold; Rendering the filled page to obtain a rendered page image specifically includes: If the page cost value is not greater than a preset threshold, the filled page is rendered to obtain a rendered page image.
7. The method according to claim 6, characterized in that The method further comprises: If the page cost value is greater than the preset threshold, page rendering of the filled page is stopped.
8. The method according to claim 7, characterized in that Rendering the filled page to obtain a rendered page image includes: Defining the component style of each page component in the filled page; the component style represents the color, shape, effect, and font of the page component; Saving the component style and component position of each page component in the populated page to a configuration file; Generate a type script based on the configuration file; and generate a renderable page according to the type script; The renderable page is subjected to image rendering to generate the rendered page image.
9. The method according to claim 8, characterized in that The obtaining of training data according to the rendered page image includes: Obtaining pixel coordinates of each page component according to the rendered page image; Annotating the rendered image according to the pixel coordinates of each page component to obtain annotated data corresponding to the rendered image; Data enhancement is performed based on the rendered image and the labeled data corresponding to the rendered image to obtain the training data.
10. The method according to claim 9, characterized in that Obtaining pixel coordinates of each page component according to the rendered page image includes: According to the position coordinates and size of each page component in the rendered page image, based on the scaling ratio, the pixel coordinates of the page component in the actual saved image are generated; the scaling ratio represents the scaling ratio between the rendered image and the actual saved image.
11. The method according to any one of claims 1 to 10, characterized in that The training data is used to train the user interface parsing model; the user interface parsing model is used to parse the components and layout in the user's page design draft and convert them into front-end code and / or structured data, so that the user can develop the user interface based on the front-end code and / or structured data.
12. A training data generating device, characterized in that: include: An acquisition module, configured to acquire a page and page components, divide the page into a plurality of page sections, and determine a page component corresponding to each page section; A filling module, configured to fill the page components corresponding to each page area into the corresponding page area according to different layout attributes, and determine a page cost value of the filled page obtained after each filling; The page cost value represents the layout matching degree of the page component under the filled page; The layout attributes represent the layout position and layout size of the page component; The generation module is used to select a filling page corresponding to the minimum page cost value, render the filling page to obtain a rendered page image; and obtain training data based on the rendered page image.
13. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 11 when executed by a processor.
15. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 11 when being executed by a processor.
Citation Information
Patent Citations
Page rendering method and device and computer storage medium
CN110489116A
User interface prototype code generation method and device, equipment and medium
CN113377356A
Webpage generation method and device, electronic equipment and storage medium
CN116974525A
Cross-platform page rendering system, electronic equipment and storage medium
CN117093316A
Method and device for building simulation test management platform and electronic equipment
CN118363835A