Front-end development method and system based on image recognition technology

Declarative text is generated through image recognition technology and large-model analysis technology, combined with WebSocket real-time update and editable environment, the problem of difficult code reading and maintenance in the existing technology is solved, and efficient and standardized front-end development is achieved.

CN120335787APending Publication Date: 2025-07-18SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510373889.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing picture-to-code scheme has problems such as difficult to read and maintain the code, cannot be engineered, and cannot adapt to multiple frameworks and component libraries, resulting in inefficient front-end development.

Method used

Image elements are marked and trained through image recognition technology, declarative text is generated, combined with WebSocket real-time updates and editable environments, and used big models to generate code based on specified frameworks and component libraries to achieve engineering development.

Benefits of technology

It simplifies the development process, reduces the operation difficulty of developers, improves development efficiency and code maintenance, and realizes the adaptation of multiple frameworks and component libraries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335787A_ABST
    Figure CN120335787A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of Internet, in particular to a front-end development method and system based on an image recognition technology, and the method comprises the following steps: picture marking, model training, image recognition, code generation and engineering. The method has the beneficial effects that a design drawing or a file screenshot is converted into an available code through an image recognition technology and a large model analysis technology, so that the development process is simplified; through targeted training of the large model, the code based on the specified framework and the component library can be generated, and the later maintenance difficulty of developers is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technologies, and particularly to a front-end development method and system based on image recognition technology. Background Art

[0002] Large models refer to machine learning models with large-scale parameters and complex computational structures. These models are usually constructed by deep neural networks and have billions or even hundreds of billions of parameters. The design purpose of large models is to improve the expressive power and prediction performance of the models, enabling them to handle more complex tasks and data. Intelligent prediction models, on the basis of large models, perform data analysis and prediction based on machine learning technologies and are models used to guide decision-making and optimize business processes. Such models are usually constructed based on a large amount of historical data and algorithms for future prediction.

[0003] Image recognition technology is one of the most remarkable breakthroughs in artificial intelligence in recent years. Image recognition technology uses a computer to deeply analyze and understand images, thereby identifying various different patterns and objects. Image recognition mainly relies on computer vision technology. By analyzing features such as pixel values, colors, edges, and textures of images, useful information is extracted. These features are then compared through a classifier or a deep learning model to determine the object or scene represented by the image. In practice, in order to improve the recognition accuracy, it is usually necessary to use a large amount of labeled data to train the model.

[0004] Based on the above technologies, the present solution proposes a front-end development solution based on image recognition technology to solve the deficiencies of existing picture-to-code solutions and improve the efficiency and standardization of front-end application development. Summary of the Invention

[0005] The purpose of the present invention is to provide a front-end development method and system based on image recognition technology to solve the deficiencies of existing picture-to-code solutions, such as difficult-to-read and maintain code, inability to be engineered, and inability to adapt to multiple frameworks and component libraries, so as to meet the requirements of high efficiency, convenience, and standardization in modern front-end development.

[0006] To achieve the above purpose, the present invention provides the following technical solution: A front-end development method based on image recognition technology, the method comprising the following steps:

[0007] Picture Marking: Mark the status of each element in the image, including the tags used by the elements, component types, and the layout method between the elements; and use data augmentation technology to generate new training samples to improve the robustness of the model;

[0008] Model Training: Select high-quality labeled data for training a pre-trained large model to enhance the model's ability to recognize various elements in a web page;

[0009] Image Recognition: Load the trained model, perform element recognition and structural analysis on the target image, and generate a declarative text that details each element in the image, including the element's label, attributes, style, text content, and hierarchical relationship.

[0010] Preferably, the dynamic update step further includes:

[0011] Initialize the page: Based on the declarative text generated by image recognition, construct an initial version of the page on the front end;

[0012] Establish a WebSocket connection: Establish a WebSocket connection between the front-end page and the back-end server to achieve real-time two-way communication;

[0013] Create an editable environment: Allow users to directly edit the attributes, style, text, and position of elements on the page and send them to the back-end server in real time via WebSocket;

[0014] Update the page in real time: After the back-end server receives the change request, update the stored declarative text and send it back to the front end, and the front end immediately updates the elements on the page.

[0015] Preferably, the method further includes a code generation step:

[0016] Parse the declarative file: Extract the label, attributes, style, and text information of each element to construct an element object structure;

[0017] Generate an HTML code string: Based on the inclusion relationship between elements, use methods such as loops and recursion to construct a complete HTML code string.

[0018] Preferably, the method further includes an engineering step:

[0019] Built-in mainstream framework templates: Provide project initialization templates for mainstream frameworks such as React, Vue, and Angular;

[0020] Function module integration: Allow users to select function modules through a visual interface, such as route management, state storage, and network requests, and automatically configure the corresponding libraries and structures according to the user's selection;

[0021] Page code generator: Generate corresponding page code based on the function modules and page layouts selected by the user;

[0022] Project packaging: Integrate function templates, project templates, and the generated page code, automatically generate configuration files, and organize files according to the standard directory structure of the selected framework;

[0023] Output to the local file system: Output the files of the entire project to the directory specified by the user.

[0024] Preferably, the structural analysis in the image recognition step further includes:

[0025] Identify element relationships: Analyze the hierarchical and sequential relationships between elements in the image;

[0026] Generate hierarchical structure text: Describe in detail the wrapping and sequential relationships between each element and other elements in the declarative text, ensuring that the generated text can accurately reflect the page layout represented by the image, so that the subsequent code generation and engineering steps can accurately reproduce the page design.

[0027] A system for a front-end development method based on image recognition technology, the system includes:

[0028] Image marking module: Used to mark the status of each element in the input image, including the tags used by the elements, component types, and the layout methods between elements; and apply data augmentation technology to generate new training samples to improve the robustness of the subsequent model;

[0029] Model training module: Used to receive the marked image data, select high-quality annotation data to train a pre-trained large-scale model to improve the model's ability to recognize various elements in the web page, and optimize the performance of the model as the training process iterates, achieving an efficient and accurate page element recognition function;

[0030] Image recognition module: Used to load the trained model, perform element recognition and structural analysis on the target image, and generate a declarative text that details each element in the image, including the element's tags, attributes, styles, text content, and hierarchical relationships. This text format is similar to the virtual DOM nodes in modern front-end frameworks.

[0031] Preferably, the system further includes a dynamic update module, which includes:

[0032] Page initialization unit: Build an initial version of the page at the front end according to the declarative text generated by the image recognition module;

[0033] WebSocket connection unit: Establish a WebSocket connection between the front-end page and the back-end server to achieve real-time two-way communication;

[0034] Editable environment unit: Provide an editable environment that enables users to directly edit the attributes, styles, text, and positions of elements on the page and send the changes to the back-end server in real time via WebSocket;

[0035] Real-time update unit: After the backend server receives a change request, it updates the stored declarative text and sends it back to the frontend via WebSocket. The frontend immediately updates the elements on the page according to the new declarative text.

[0036] Preferably, the system further includes a code generation module, which is used for:

[0037] Declarative file parsing unit: Parse the declarative file generated by the image recognition module, extract the tags, attributes, styles, and text information of each element, and construct an element object structure;

[0038] HTML code generation unit: According to the inclusion relationship between elements, use loop and recursive methods to convert the element object structure into a complete HTML code string.

[0039] Preferably, the system further includes an engineering module, which is used for:

[0040] Built-in framework template unit: Provide project initialization templates for mainstream frameworks such as React, Vue, and Angular;

[0041] Functional module integration unit: Provide a visual interface that allows users to select the required functional modules, such as route management, state storage, network requests, and automatically configure the corresponding libraries and structures according to the user's selection;

[0042] Page code generator unit: Generate corresponding page code according to the functional modules and page layouts selected by the user;

[0043] Project packaging unit: Integrate functional templates, project templates, and the generated page code, automatically generate configuration files, and organize files according to the standard directory structure of the selected framework;

[0044] Output unit: Output the files of the entire project to the directory specified by the user.

[0045] Preferably, the system further includes:

[0046] Structure analysis sub-module: In the image recognition module, there is a structure analysis sub-module, which is used to build the hierarchical relationship between these elements, including the wrapping relationship and the order relationship, after all elements are recognized;

[0047] Optimization of declarative text generation: The generated declarative text not only contains the basic information of the elements, but also details the hierarchical relationship and order relationship between the elements, ensuring that subsequent code generation and engineering steps can accurately reproduce the page design;

[0048] User Interaction Interface: The system provides a user-friendly interaction interface, enabling users to conveniently perform operations such as image uploading, marking, model training, page editing, code generation, and project packaging.

[0049] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0050] The front-end development method and system based on image recognition technology proposed by the present invention convert design drawings or file screenshots into available code through image recognition technology and large model analysis technology, simplifying the development process; through targeted training of the large model, it is possible to generate code based on a specified framework and component library, reducing the later maintenance difficulty for developers; by generating an editable page and implementing real-time update of page changes to the code through WebSocket connection; and realizing the generation of complete engineering code from a single image or multiple images, which simplifies the operation process and reduces the operation difficulty for developers. Description of the Drawings

[0051] Figure 1 It is a flow chart of the method of the present invention. Detailed Embodiments

[0052] In order to clearly and completely describe the objectives, technical solutions, and advantages of the present invention, the following further details the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are some, but not all, of the embodiments of the present invention, and are merely used to explain the embodiments of the present invention, rather than to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0053] Embodiment 1, please refer to Figure 1 The present invention provides a technical solution: A front-end development method based on image recognition technology, the method comprising the following steps:

[0054] 1) Image Marking: In the data preparation stage, the annotation of images is an essential step. The accuracy and comprehensiveness of annotation directly affect the training effect of the model. Usually, developers need to mark the status of each element in the image, including the labels used for the elements, component types, and the previous layout methods of the elements, etc. This step requires not only meticulous work but also rich domain knowledge to ensure the accuracy of annotation. For example, in a web design diagram, developers need to identify different types of elements such as buttons, input boxes, text areas, etc., and assign the correct labels to them. In addition, to improve the robustness of the model, developers also need to consider data augmentation techniques. Data augmentation refers to generating new training samples by transforming the original images, such as rotation, scaling, cropping, etc. This process can not only effectively increase the quantity of training data but also improve the adaptability of the model in the face of different situations. In this way, the model can learn more features during the training process, thereby improving the accuracy of recognition.

[0055] 2) Model Training: After completing the first step of marking a large amount of image data, the next step is to select those high-quality annotated data for further machine learning training. These carefully selected data sets will be used to "feed" a pre-trained large-scale model, aiming to improve the model's ability to recognize various elements in the web page, such as buttons, text boxes, links, etc. In this way, we can significantly enhance the performance of the model, enabling it not only to recognize these elements faster but also to classify and locate them more accurately. As the training process iterates, the performance of the model will be gradually optimized, and finally, it can achieve efficient and accurate page element recognition.

[0056] 3) Image Recognition: After completing the model training, the next step is to apply this carefully tuned large model to the actual scenario - namely image recognition. The goal of this stage is to conduct a meticulous analysis of the target image and generate a declarative text that can accurately describe each element in the image. This text needs to record in detail the label, attributes, style, text content of each element, as well as the hierarchical relationship between these elements. This format is very similar to the virtual DOM nodes in modern front-end frameworks such as Vue and React, so that the output text can clearly reflect the page layout represented by the image.

[0057] Specifically, the process can be carried out as follows:

[0058] 1. Load the Model: First, load the already trained model and prepare to start the recognition task.

[0059] 2. Image Input: Take the target image as the input. This image may be a complete web page screenshot or any other form of image file.

[0060] 3. Element recognition: The model starts to recognize elements in the image one by one, such as buttons, text boxes, links, etc., and analyzes their features such as position, size, and style.

[0061] 4. Structure analysis: After recognizing all elements, the model also needs to construct the hierarchical relationships between these elements. For example, whether an element is wrapped by another element, or whether there is a specific sequential relationship between two elements.

[0062] 5. Generate declarative text: The last step is to integrate all the information to generate a declarative text that is easy to understand and process. This text should be able to fully reflect the page represented by the image, including but not limited to the tag names, IDs, class names, style attributes (such as color, font size), and text content of the elements.

[0063] 4) Dynamic update: To achieve dynamic update, a fully editable page needs to be generated based on the image recognition results, and this page should have the ability to be updated in real time. This solution designs a system based on the WebSocket communication mechanism. The core function of this system is to enable users to directly edit each element on the page in the browser, including styles, attributes, text, and positions, and to be able to see the effects of these changes in real time. This can be achieved through the following steps:

[0064] 1. Initialize the page: Based on the declarative text generated by image recognition, construct an initial version of the page on the front end.

[0065] 2. Establish a WebSocket connection: Establish a WebSocket connection between the front-end page and the back-end server to achieve real-time two-way communication, so as to be able to update the state of the page immediately.

[0066] 3. Create an editable environment: Build an editable environment that allows users to directly edit the attributes, styles, text, and positions of elements on the page. When users make changes, these changes will be immediately sent to the back-end server via WebSocket.

[0067] 4. Update the page in real time: After the back-end server receives the change request, it updates the stored declarative text. At the same time, the back-end server will also send the updated declarative text back to the front end via WebSocket. After the front end receives the new declarative text, it immediately updates the elements on the page to achieve the hot update effect.

[0068] 5) Code generation: To achieve the conversion from a declarative file to a page code string, we can adopt a step-by-step approach. First, we need to parse the data in the declarative file, extract information such as tags, attributes, styles, and text for each element, and construct an element object structure. Then, based on the inclusion relationships between elements, we use methods like loops and recursion to construct a complete HTML code string.

[0069] 6) Engineering: To achieve engineering, this solution designs an infrastructure that can generate the basic structure of a specific framework project and allows users to select the functional modules they want through a visual interface, such as route management, state storage, and network requests. Once the user has completed the selection of all options, our tool will automatically integrate these functional modules, project templates, and the user-defined page code, finally generating an engineering project that meets expectations and outputting it to the user's local file system.

[0070] The specific steps are as follows:

[0071] 1. Built-in mainstream framework templates: React: including the default configuration of create-react-app, Vue: creating a project using vue-cli, Angular: initializing a project using the ng new command.

[0072] 2. Functional module integration: Provide a dropdown list for users to select the functional modules they use: route management (such as React Router, Vue Router, etc.), state storage (such as Redux, Vuex, MobX, etc.), network requests (such as Axios, Fetch API, etc.). According to the selected framework, configure the corresponding routing library. According to the selected state management library, set up the basic store structure, and configure the basic template for network requests, including interceptors, error handling, etc.

[0073] 3. Page code generator: Based on the method of generating a page code string in the previous step, further expand the functionality to adapt to more complex page structures, and generate the corresponding page code according to the selected functional modules and page layout.

[0074] 4. Project packaging: Integrate the functional templates, project templates, and the generated page code together, automatically generate configuration files such as.gitignore and package.json, and finally organize the files according to the standard directory structure of the selected framework.

[0075] 5. Output to the local file system: Output the files of the entire project to the specified directory of the user.

[0076] Embodiment 2, based on Embodiment 1, proposes a system for a front-end development method based on image recognition technology. The system includes:

[0077] Image marking module: used to mark the status of each element in the input image, including the tags used by the elements, component types, and the layout methods between the elements; and apply data augmentation technology to generate new training samples to improve the robustness of subsequent models.

[0078] Model training module: used to receive the marked image data, select high-quality annotation data to train a pre-trained large-scale model to enhance the model's ability to recognize various elements in the web page, and optimize the performance of the model with the continuous iteration of the training process to achieve efficient and accurate page element recognition.

[0079] Image recognition module: used to load the trained model, perform element recognition and structure analysis on the target image, and generate a declarative text that details each element in the image, including the element's tags, attributes, styles, text content, and hierarchical relationship. This text format is similar to the virtual DOM nodes in modern front-end frameworks.

[0080] It also includes a dynamic update module, which includes: Page initialization unit: construct an initial version of the page at the front end according to the declarative text generated by the image recognition module; WebSocket connection unit: establish a WebSocket connection between the front-end page and the back-end server to achieve real-time two-way communication; Editable environment unit: provide an editable environment that enables users to directly edit the attributes, styles, text, and positions of elements on the page and send the changes to the back-end server in real time through WebSocket; Real-time update unit: after the back-end server receives the change request, update the stored declarative text and send it back to the front end through WebSocket, and the front end immediately updates the elements on the page according to the new declarative text.

[0081] The system also includes a code generation module, which is used for: Declarative file parsing unit: parse the declarative file generated by the image recognition module, extract the tags, attributes, styles, and text information of each element, and construct an element object structure; HTML code generation unit: according to the inclusion relationship between the elements, use loop and recursive methods to convert the element object structure into a complete HTML code string.

[0082] The system also includes an engineering module, which is used for: Built-in framework template unit: providing project initialization templates for mainstream frameworks such as React, Vue, and Angular; Functional module integration unit: providing a visual interface that allows users to select the required functional modules, such as routing management, state storage, and network requests, and automatically configuring the corresponding libraries and structures according to the user's selection; Page code generator unit: generating corresponding page code according to the functional modules and page layouts selected by the user; Project packaging unit: integrating functional templates, project templates, and the generated page code, automatically generating configuration files, and organizing files according to the standard directory structure of the selected framework; Output unit: outputting the files of the entire project to the directory specified by the user.

[0083] The system also includes: Structural analysis sub-module: In the image recognition module, there is a structural analysis sub-module, which is used to build the hierarchical relationships between these elements, including wrapping relationships and sequential relationships, after all elements have been recognized; Declarative text generation optimization: The generated declarative text not only includes the basic information of the elements, but also details the hierarchical relationships and sequential relationships between the elements, ensuring that subsequent code generation and engineering steps can accurately reproduce the page design; User interaction interface: The system provides a friendly user interaction interface, enabling users to conveniently perform operations such as image uploading, marking, model training, page editing, code generation, and project packaging.

[0084] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A front-end development method based on image recognition technology, characterized in that: The method includes the following steps: Image marking: Mark the status of each element in the image, including the labels used by the elements, component types, and the layout method between the elements; and use data augmentation technology to generate new training samples to improve the robustness of the model; Model training: Select high-quality labeled data to train a pre-trained large-scale model to improve the model's ability to recognize various elements in the web page; Image recognition: Load the trained model, perform element recognition and structural analysis on the target image, and generate a declarative text that details each element in the image, including the element's label, attributes, styles, text content, and hierarchical relationship.

2. The front-end development method based on image recognition technology according to claim 1, characterized in that: The dynamic update step further includes: Initializing the page: Build an initial version of the page on the front end according to the declarative text generated by image recognition; Establishing a WebSocket connection: Establish a WebSocket connection between the front-end page and the back-end server to achieve real-time two-way communication; Creating an editable environment: Allow users to directly edit the attributes, styles, text, and positions of elements on the page and send them to the back-end server in real time via WebSocket; Real-time page update: After the back-end server receives the change request, update the stored declarative text and send it back to the front end, and the front end immediately updates the elements on the page.

3. A front-end development method based on image recognition technology according to claim 3, characterized in that: The method further includes a code generation step: Parsing the declarative file: Extract the labels, attributes, styles, and text information of each element to build an element object structure; Generating an HTML code string: Build a complete HTML code string using methods such as loops and recursion according to the inclusion relationship between the elements.

4. A front-end development method based on image recognition technology according to claim 3, characterized in that: The method further includes an engineering step: Built-in mainstream framework templates: Provide project initialization templates for the React, Vue, and Angular mainstream frameworks; Function module integration: Allow users to select function modules through a visual interface, such as route management, state storage, and network requests, and automatically configure the corresponding libraries and structures according to the user's selection; Page code generator: Generate corresponding page code according to the function modules and page layouts selected by the user; Project packaging: Integrate function templates, project templates, and the generated page code, automatically generate configuration files, and organize files according to the standard directory structure of the selected framework; Output to the local file system: Output the files of the entire project to the directory specified by the user.

5. A front-end development method based on image recognition technology according to claim 4, characterized in that: The structural analysis in the image recognition step further includes: Identifying element relationships: Analyze the hierarchical relationship and sequential relationship between elements in the image; Generating hierarchical structure text: Describe in detail the wrapping relationship and sequential relationship between each element and other elements in the declarative text, ensuring that the generated text can accurately reflect the page layout represented by the image, so that the subsequent code generation and engineering steps can accurately reproduce the page design.

6. A system for a front-end development method based on image recognition technology according to any one of claims 1-5, characterized in that: The system includes: Image marking module: Used to mark the status of each element in the input image, including the labels used by the elements, component types, and the layout method between the elements; and apply data augmentation technology to generate new training samples to improve the robustness of the subsequent model; Model training module: It is used to receive the labeled image data, select the high-quality annotation data to train the pre-trained large-scale model, so as to improve the model's ability to recognize various elements in the web page, and optimize the performance of the model with the continuous iteration of the training process, realizing the efficient and accurate page element recognition function; Image recognition module: It is used to load the trained model, perform element recognition and structure analysis on the target image, and generate a declarative text that details each element in the image, including the element's tag, attributes, styles, text content, and hierarchical relationship. This text format is similar to the virtual DOM nodes in modern front-end frameworks.

7. A system according to claim 6, characterized in that: The system also includes a dynamic update module, which includes: Page initialization unit: According to the declarative text generated by the image recognition module, construct an initial version of the page on the front end; WebSocket connection unit: Establish a WebSocket connection between the front-end page and the back-end server to achieve real-time two-way communication; Editable environment unit: Provide an editable environment that allows users to directly edit the attributes, styles, text, and positions of elements on the page, and send the changes to the back-end server in real time through WebSocket; Real-time update unit: After the back-end server receives the change request, update the stored declarative text and send it back to the front end through WebSocket. The front end immediately updates the elements on the page according to the new declarative text.

8. A system according to claim 7, characterized in that: The system also includes a code generation module, which is used for: Declarative file parsing unit: Parse the declarative file generated by the image recognition module, extract the tag, attributes, styles, and text information of each element, and construct an element object structure; HTML code generation unit: According to the inclusion relationship between elements, use loop and recursive methods to convert the element object structure into a complete HTML code string.

9. A system according to claim 8, wherein: The system also includes an engineering module, which is used for: Built-in framework template unit: Provide project initialization templates for mainstream frameworks such as React, Vue, and Angular; Functional module integration unit: Provide a visual interface that allows users to select the required functional modules, such as routing management, state storage, and network requests, and automatically configure the corresponding libraries and structures according to the user's selection; Page code generator unit: Generate the corresponding page code according to the functional modules and page layouts selected by the user; Project packaging unit: Integrate the functional templates, project templates, and generated page codes, automatically generate configuration files, and organize the files according to the standard directory structure of the selected framework; Output unit: Output the files of the entire project to the directory specified by the user.

10. A system according to claim 9, wherein: The system also includes: Structure analysis sub-module: In the image recognition module, there is a structure analysis sub-module, which is used to build the hierarchical relationship between these elements, including the wrapping relationship and the sequential relationship, after all elements are recognized; Optimization of declarative text generation: The generated declarative text not only contains the basic information of the elements, but also details the hierarchical relationship and sequential relationship between the elements, ensuring that the subsequent code generation and engineering steps can accurately reproduce the page design; User Interaction Interface: The system provides a user-friendly interaction interface, enabling users to conveniently perform operations such as image uploading, marking, model training, page editing, code generation, and project packaging.