Method and system for generating AI graph recognition front-end code and computer storage medium
Through the configuration-based AI graph recognition front-end code generation method, the convolutional neural network and object detection algorithm are used to identify and extract interface elements, and the front-end code is generated in combination with greedy search or bundled search strategies. The problems of insufficient support for complex interface elements and inaccurate generation results in the existing technology are solved, and high-quality and flexible code generation is achieved.
Patent Information
- Application Number
- CN202510127158.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-31
AI Technical Summary
In the prior art, front-end code automatic generation lacks support for complex interface elements, the generation results are inaccurate, and the auxiliary information provided by the designer cannot be effectively utilized.
The configuration-based AI graph recognition front-end code generation method is adopted, and image recognition and feature extraction is used to use convolutional neural networks, combined with optical character recognition technology and object detection algorithms, the final rendering tree is generated, and high-quality and diverse front-end code is generated through greedy search or cluster search strategies.
It significantly improves the quality and flexibility of code generation, improves the accuracy of generated results, can effectively utilize the auxiliary information provided by designers, reduces the communication cost between designers and developers, and shortens the time period from design to implementation.
Smart Images

Figure CN120122934A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, system, and computer storage medium for generating AI image recognition front-end code. Background Art
[0002] With the development of the Internet and the improvement of user experience requirements, front-end development plays an increasingly important role in software engineering. Traditionally, web design and development is a highly manual process. Designers first create visual prototype diagrams, and then front-end developers convert these static images into interactive web code (HTML, CSS, and JavaScript). However, in the practical process, the inventor found some problems. First, errors are easily introduced during the conversion process from the design draft to the executable code. Second, developers must have a deep understanding of the design intent to accurately achieve the expected effect, which often requires a large amount of communication costs. Finally, in the face of rapidly changing design trends and technology stacks, traditional methods are difficult to respond efficiently.
[0003] In recent years, the progress of automated tools and AI technology has made it possible to simplify this process. Automated front-end code generation tools can reduce manual intervention, improve efficiency, and reduce the probability of errors. However, existing solutions are usually limited to simple layout and style generation, lack of understanding and support for complex interface elements, and lack flexibility, making it difficult to adapt well to diverse user needs and specific design styles. In addition, they often ignore the importance of human-computer collaboration and fail to make full use of the additional information provided by designers to guide the code generation process, resulting in inaccurate or unexpected generation results.
[0004] Therefore, the inventor designed a configuration-based AI image recognition front-end code solution to solve the above problems. Summary of the Invention
[0005] Aiming at the defects of the above-mentioned prior art, the present invention provides a configuration-based AI image recognition front-end code solution, aiming to solve the problems in the prior art that the automatic generation of front-end code lacks support for complex interface elements, the generation results are inaccurate, and it is impossible to effectively utilize the auxiliary information provided by designers.
[0006] To achieve the above object, the technical solution adopted by the present invention is: a method for generating AI image recognition front-end code, including the following steps:
[0007] S1: Use a convolutional neural network for image recognition, and the image recognition includes the following steps:
[0008] S11: Input the web prototype diagram and its configuration information;
[0009] S12: Preprocess the prototype diagram;
[0010] S13: Extract features from the preprocessed prototype diagram, identify interface elements, determine the hierarchical structure, and generate a preliminary hierarchical structure tree;
[0011] S14: The classifier identifies and marks user interface elements based on the text information on the prototype diagram obtained through optical character recognition technology and the text descriptions in the input configuration information;
[0012] S15: Generate a final preliminary hierarchical structure tree based on the marked user interface elements;
[0013] S2: Layout analysis, and the layout analysis includes the following steps:
[0014] S21: Detect the boundary information of the prototype diagram using an object detection algorithm;
[0015] S22: Extract the attribute information of the interface elements, including color, size, and position;
[0016] S23: Ensure that all attributes are correctly extracted and processed by looping and searching, and repeat steps S21 and S22 if necessary until completion;
[0017] S24: Generate a final rendering tree based on the extracted attribute information;
[0018] S3: Code generation, and the code generation includes the following steps:
[0019] S31: Tokenize the data generated in steps S15 and S24 through data preprocessing and building a dataset, and map it into HTML, CSS, or JavaScript code;
[0020] S32: Split and tokenize the HTML, CSS, or JavaScript code according to syntax units to form a dataset pair of input sequence and output sequence. During the decoder decoding process, use greedy search or beam search to gradually generate a front-end code token sequence by comparing the syntax tokens generated by each sequence;
[0021] S33: Concatenate and parse the generated code token sequence according to the syntax rules of the front-end code to obtain the final front-end code.
[0022] Based on the above, the beneficial effects of an AI image recognition front-end code solution based on configuration are to solve the problems in the prior art that the automatic generation of front-end code lacks support for complex interface elements, the generation results are inaccurate, and the auxiliary information provided by designers cannot be effectively utilized; mainly reflected in:
[0023] 1. In the decoding process of the decoder, the present invention flexibly selects the greedy search or beam search strategy according to the clarity of user requirements. When the requirements are clear, the greedy search can be used to quickly generate the basic sequence, thereby reducing the computational cost; when the requirements are unclear, the beam search can maintain a candidate list of size k, select the k sequences with the highest probability to continue generation, and finally select the optimal solution from them. In summary, it not only speeds up the generation speed of the basic sequence, but also ensures that high-quality and diverse outputs can be produced even in the face of ambiguous requirements, significantly improving the quality and flexibility of code generation, enabling effective utilization of the auxiliary information provided by the designer, and improving the accuracy of the generation result;
[0024] 2. The present invention extracts features through a convolutional neural network and combines optical character recognition technology and object detection algorithms, and can accurately identify and process complex visual and text information, thereby supporting the accurate parsing of more diverse and complex user interface elements;
[0025] 3. The process of automatically generating front-end code in the present invention greatly shortens the time cycle from design to implementation, reduces the communication cost between designers and developers, and improves the efficiency of the entire development process. Especially through the intelligent selection search strategy of greedy search or beam search, the system can generate code that meets requirements more efficiently, further accelerating the development speed.
[0026] Further, in step S11, the ways of inputting the web page prototype diagram include uploading the design draft or temporarily hand-drawing the electronic draft. The input configuration information includes text input control information and selection of the syntax template. When the syntax template is not selected, the default template is automatically selected to assist code generation and enhance the recognition accuracy and code generation accuracy.
[0027] Based on the above, by inputting the web page prototype diagram and inputting the configuration information, the function of allowing users to mark control information and output descriptions on the prototype diagram is provided, thereby enhancing the recognition accuracy of the system, ensuring that the finally generated code is closer to the original intention of the designer, reducing errors caused by misunderstanding the design draft, enabling the system to more accurately understand the user's intention, and thus generating front-end code that is more in line with expectations.
[0028] Further, in step S15, it is judged whether the user has selected the syntax template. If so, it jumps to steps S21 and S22. If not, it continues to judge whether the user has marked control information on the prototype diagram. If not, the default syntax template is called, and it is re-judged whether the user has selected the syntax template. If so, it is retrieved from the library to judge whether the control information is an unknown control. If so, the default syntax template is called, and it is re-judged whether the user has selected the syntax template. If not, it jumps to steps S21 and S22.
[0029] Further, in step S11, the input configuration information further includes the code generated by the user's description through the input interface to assist in code generation for further customization.
[0030] Based on the above, through the user's output description, the system can adjust the code generation logic according to the user's special requirements to generate more personalized and customized front-end code. For example, the user can specify the behavior mode, style change rules, or responsive layout requirements of certain components through the description, so that the generated code better adapts to the unique needs of the project. For the special requirements or technical limitations existing in different projects, the user can inform the system through the description, such as information on the usage habits of specific front-end frameworks, libraries, or internal company technical specifications, etc., so that the generated code not only meets the general standards but also satisfies the special conditions in a specific environment, enhancing the adaptability and practicality of the system.
[0031] Further, in step S32, when the system detects that the amount of information provided by the user meets the preset threshold and the complexity of the design draft is moderate, a greedy search is used to quickly generate the basic sequence; when the amount of information is below the threshold or the complexity of the design draft is higher than the set standard, beam search is used to increase the possibility and quality of the generated code, where the complexity is evaluated by an algorithm predefined by the system.
[0032] Further, in step S32, the greedy search strategy selects the single token with the highest current probability at each iteration to construct the code token sequence; while the beam search strategy maintains a candidate list with a fixed size of k, where k is jointly determined by the system's performance parameters and the accuracy requirements specified by the user.
[0033] The present invention also provides a system for implementing the method for generating the AI image recognition front-end code, including:
[0034] An image processing module for performing steps S11 and S12;
[0035] A feature extraction module for performing steps S13 - S15;
[0036] A text parsing module for performing step S14;
[0037] A target detection module for performing steps S21 and S22;
[0038] A loop search module for performing step S23;
[0039] A rendering tree generation module for performing step S24;
[0040] And a code generation module for performing steps S31 - S33.
[0041] Based on the above, each module is responsible for a specific task and has clear functional boundaries, enabling developers to focus on optimizing the performance of a single module without affecting the work of other parts. When a certain function needs to be updated or repaired, only the corresponding module needs to be operated on, reducing the maintenance cost of the entire system. Thus, with the progress of technology, even if new algorithms and technologies can be easily introduced into the corresponding modules, there is no need to comprehensively transform the entire system;
[0042] Modular design also separates different functions. Even if a problem occurs in a certain module, it can limit the scope of the fault and does not affect the normal operation of other modules. In summary, it can ensure that the system can still maintain its basic functions when some components are abnormal, improving the fault tolerance and stability of the system.
[0043] Furthermore, the system of the method for generating the AI image recognition front-end code further includes a user interaction interface, which provides functions for users to upload prototype diagrams, select syntax templates, input the resolution of prototype diagrams, add additional annotations or descriptions, and view the generated code.
[0044] Based on the above, the beneficial effect that the user interaction interface allows designers to directly add annotations or descriptions when uploading prototype diagrams is that it can be directly used to guide code generation, thus greatly reducing the possibility of misunderstandings and rework; the beneficial effect that users can select different syntax templates is that it can generate front-end code that conforms to specific specifications according to the technical stack requirements of the project, and the generated code can be directly applied to actual projects, reducing the workload of subsequent modifications.
[0045] Furthermore, the system of the method for generating the AI image recognition front-end code supports a variety of popular front-end frameworks and component libraries, including Element UI and Ant Design.
[0046] The present invention also provides a computer storage medium, on which instructions are stored. When the instructions are executed by a computer, the computer is made to execute the method for generating the front-end code based on configured AI image recognition.
[0047] To more clearly elaborate the above features of the present invention and the purposes to be achieved, the following further describes the present invention in combination with the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 : is the flowchart of the front-end code steps of the present invention;
[0049] Figure 2 : is the flowchart of the method for generating the front-end code of the present invention;
[0050] Figure 3 : is the module connection diagram of the system of the method for generating the front-end code of the present invention. Detailed implementation manners
[0051] Refer to Figures 1 - 3 ;
[0052] This embodiment discloses a method for generating the front - end code of AI image recognition, including the following steps:
[0053] S1: Use a convolutional neural network for image recognition, and the image recognition includes the following steps:
[0054] S11: Input the web page prototype diagram and its configuration information;
[0055] S12: Pre - process the prototype diagram;
[0056] S13: Extract features from the pre - processed prototype diagram, identify interface elements, determine the hierarchical structure, and generate a preliminary hierarchical structure tree;
[0057] S14: The classifier identifies and marks user interface elements through the text information on the prototype diagram obtained by optical character recognition technology and the text descriptions in the input configuration information;
[0058] S15: Generate a final preliminary hierarchical structure tree according to the marked user interface elements;
[0059] S2: Layout analysis, and the layout analysis includes the following steps:
[0060] S21: Use an object detection algorithm to detect the boundary information of the prototype diagram;
[0061] S22: Extract the attribute information of the interface elements, including color, size, and position;
[0062] S23: Ensure that all attributes are correctly extracted and processed through loop search, and repeat steps S21 and S22 if necessary until completion;
[0063] S24: Generate a final rendering tree according to the extracted attribute information;
[0064] S3: Code generation, and the code generation includes the following steps:
[0065] S31: Tokenize the data generated in steps S15 and S24 through data pre - processing and building a data set, and map it into HTML, CSS, or JavaScript code;
[0066] S32: Split and tag the HTML, CSS, or JavaScript code according to syntax units to form a dataset pair of input sequence and output sequence. During the decoder decoding process, use greedy search or beam search to gradually generate the front-end code token sequence by comparing the syntax tokens generated by each sequence.
[0067] S33: Concatenate and parse the generated code token sequence according to the syntax rules of the front-end code to obtain the final front-end code.
[0068] In this embodiment, in step S11, the ways of inputting the web page prototype diagram include uploading a design draft or a temporary hand-drawn electronic draft. The input configuration information includes text input control information and a selected syntax template. When no syntax template is selected, the default template is automatically selected to assist in code generation and enhance the recognition accuracy and code generation accuracy.
[0069] In this embodiment, in step S15, it is judged whether the user has selected a syntax template. If so, jump to steps S21 and S22. If not, continue to judge whether the user has marked control information on the prototype diagram. If not, call the default syntax template and re-judge whether the user has selected a syntax template. If so, retrieve the library to judge whether the control information is an unknown control. If so, call the default syntax template and re-judge whether the user has selected a syntax template. If not, jump to steps S21 and S22.
[0070] In this embodiment, in step S11, the input configuration information further includes that the user describes through the input interface to assist in code generation and further customize the generated code.
[0071] In this embodiment, in step S32, when the system detects that the amount of information provided by the user meets the preset threshold and the complexity of the design draft is moderate, use greedy search to quickly generate the basic sequence; when the amount of information is lower than the threshold or the complexity of the design draft is higher than the set standard, use beam search to increase the possibility and quality of the generated code, where the complexity is evaluated by an algorithm predefined by the system.
[0072] In this embodiment, in step S32, the greedy search strategy selects the single token with the highest current probability in each iteration to construct the code token sequence; while the beam search strategy maintains a candidate list with a fixed size of k, where k is jointly determined by the performance parameters of the system and the accuracy requirements specified by the user.
[0073] The present invention also discloses a system for implementing the front-end code generation method of the above-mentioned configuration-based AI image recognition, which is characterized by including:
[0074] An image processing module, used to execute steps S11 and S12;
[0075] A feature extraction module for performing steps S13 - S15;
[0076] A text parsing module for performing step S14;
[0077] An object detection module for performing steps S21 and S22;
[0078] A loop search module for performing step S23;
[0079] A rendering tree generation module for performing step S24;
[0080] And a code generation module for performing steps S31 - S33.
[0081] In this embodiment, the system for the method of generating the AI image recognition front - end code further includes a user interface, which provides functions for the user to upload a prototype diagram, select a syntax template, input the resolution of the prototype diagram, add additional notes or descriptions, and view the generated code.
[0082] In this embodiment, the system for the method of generating the AI image recognition front - end code supports a variety of popular front - end frameworks and component libraries, including Element UI and Ant Design, facilitating project maintenance and iteration.
[0083] The present invention also discloses a computer storage medium, on which instructions are stored. When the instructions are executed by a computer, the computer is made to execute the front - end code generation method based on configured AI image recognition.
[0084] In summary, the specific implementation of the present invention is as follows:
[0085] The pre - processing of the prototype diagram includes image scaling, cropping, filling, and normalization, specifically:
[0086] 1. Uniformly scale the uploaded prototype diagram to 224x224 pixels;
[0087] 2. Cropping and filling: If the aspect ratio of the image is greater than 1 (i.e., the width is greater than the height), black filling bars can be added at the top and bottom to make the image height reach 224 pixels; if the aspect ratio is less than 1 (i.e., the height is greater than the width), black filling bars can be added on the left and right sides to make the image width reach 224 pixels, ensuring that the image maintains the correct aspect ratio and avoiding recognition errors caused by deformation;
[0088] 3. Normalization: Convert the pixel values of the image from 0 - 255 to floating - point numbers between 0 - 1 by dividing by 255 to improve the convergence speed and performance of the model;
[0089] Feature extraction uses a convolutional neural network called VGGNet (Visual Geometry Group Net) to extract features of an image. When recognizing an image, if there are clearly marked attributes on the image, they are directly written into the code of the control; if there are no clearly marked attributes, the div tag is used, and then the style is extracted to achieve the same visual effect;
[0090] When extracting the attribute information of interface elements, if there is no control layout with marked attributes, default attributes are used to generate the control. Among them, the recognized control structure is:
[0091] {'type': 'button',
[0092] 'text': 'demo',
[0093] 'id': 'demo',
[0094] 'class': 'demo',
[0095] 'position': {'x': 100, 'y': 100},
[0096] 'size': {'width': 100, 'height': 50},
[0097] 'color': {'background': '#0000FF', 'text': '#FFFFFF'}
[0098] 'children': {} ....
[0100] }};
[0101] In steps S32 and S33, it should be understood that the input sequence in step S32 is input into the encoder of the trained Transformer model. Then, at the decoder end, starting from the start token, according to the token probability distribution output by the model, greedy search or beam search is used to sequentially select the code tokens with high weights. By comparing the syntax tokens generated by each sequence, the front-end code token sequence is gradually generated. Finally, these token sequences are spliced and parsed according to the syntax rules of the front-end code to obtain the final front-end code, and then it is converted into the syntax of a specific front-end framework through a plugin.
[0102] The above is only the optimal solution embodiment of the present invention and is not used to limit the present invention. Various modifications or substitutions made by those skilled in the art to the present invention without departing from the essence and protection scope of the present invention should also be within the protection scope of the present invention.
Claims
1. A method for generating AI image recognition front-end code, characterized in that: The following steps are involved: S1: Use a convolutional neural network to perform image recognition, the image recognition comprising the following steps: S11: Input the webpage prototype image and its configuration information; S12: preprocessing the prototype image; S13: extracting features from the preprocessed prototype image, identifying interface elements, determining the hierarchical structure, and generating a preliminary hierarchical structure tree; S14: The classifier identifies and marks user interface elements using text information on the prototype image obtained by optical character recognition technology and text descriptions in the input configuration information; S15: generating a final preliminary hierarchical structure tree according to the marked user interface elements; S2: Layout analysis, the layout analysis includes the following steps: S21: Use the target detection algorithm to detect the boundary information of the prototype image; S22: extracting attribute information of interface elements, including color, size and position; S23: Ensure that all attributes are correctly extracted and processed by loop search, and repeat steps S21 and S22 until completion if necessary; S24: generating a final rendering tree according to the extracted attribute information; S3: Code generation, the code generation includes the following steps: S31: tokenizing the data generated in steps S15 and S24 by means of data preprocessing and data set construction, and mapping them into HTML, CSS or JavaScript codes; S32: Split and mark the HTML, CSS or JavaScript code according to grammatical units to form a data set pair of input sequence and output sequence. In the decoding process of the decoder, greedy search or beam search is used to gradually generate a front-end code mark sequence by comparing the grammatical marks generated by each sequence. S33: The generated code token sequence is concatenated and parsed according to the grammatical rules of the front-end code to obtain the final front-end code.
2. The method for generating AI image recognition front-end code according to claim 1, characterized in that: In step S11, the method of inputting the web page prototype includes uploading a design draft or a temporary hand-drawn electronic draft, and inputting configuration information includes text input control information and selecting a syntax template. When the syntax template is not selected, the default template is automatically selected to assist code generation and enhance recognition accuracy and code generation accuracy.
3. The method for generating AI image recognition front-end code according to claim 2, characterized in that: In step S15, determine whether the user has selected a grammar template. If so, jump to steps S21 and S22. If not, continue to determine whether the user has marked the control information on the prototype diagram. If not, call the default grammar template and re-determine whether the user has selected a grammar template. If so, search the library to determine whether the control information is an unknown control. If so, call the default grammar template and re-determine whether the user has selected a grammar template. If not, jump to steps S21 and S22.
4. The method for generating AI image recognition front-end code according to claim 3, characterized in that: In step S11, inputting configuration information also includes a user inputting an interface description to assist in code generation, thereby further customizing the generated code.
5. The method for generating AI image recognition front-end code according to claim 4, characterized in that: In step S32, when the system detects that the amount of information provided by the user meets the preset threshold and the complexity of the prototype is moderate, greedy search is used to quickly generate a basic sequence; when the amount of information is lower than the threshold or the complexity of the design is higher than the set standard, beam search is used to increase the possibility and quality of code generation, so that each time the user uploads the same prototype, the result is not unique and unchanging.
6. The method for generating AI image recognition front-end code according to claim 5, characterized in that: In step S32, the greedy search strategy selects the single word with the highest current probability in each iteration to construct the code token sequence; while the beam search strategy maintains a candidate list of a fixed size of k, where k is determined by the system's performance parameters and the accuracy requirements specified by the user.
7. A system for implementing the method for generating the AI image recognition front-end code according to claim 6, characterized in that: The system comprises: An image processing module, used for executing steps S11 and S12; A feature extraction module, used to execute steps S13-S15; A text parsing module, used to execute step S14; A target detection module, used to execute steps S21 and S22; A loop search module, used for executing step S23; A rendering tree generation module, used to execute step S24; And a code generation module, used to execute steps S31-S33.
8. The system according to claim 7, characterized in that The system of the method for generating the AI image recognition front-end code also includes a user interaction interface, which provides users with the functions of uploading prototype images, selecting syntax templates, inputting prototype image resolution, adding additional comments or descriptions, and viewing generated code.
9. The system according to claim 8, characterized in that The system of the method for generating the AI image recognition front-end code supports a variety of popular front-end frameworks and component libraries, including Element UI and Ant Design.
10. A computer storage medium having instructions stored thereon, which, when executed by a computer, causes the computer to execute the method for generating an AI image recognition front-end code as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Front end interface code generation method and device, electronic equipment and storage medium
CN108228183A
Method and system for generating front-end webpage code based on PSD file and storage medium
CN111562919A
Code generation method and system based on natural semantic understanding
CN115202640A
Method for generating image descriptor and related device
CN118484557A
Public code library management method, system, equipment and medium
CN118897668A