Training method of neural network model for generating codes and method and device for generating codes based on neural network model
By combining visual encoder, adapter and language model neural network models, the problem that complex design concepts in front-end development are difficult to accurately describe text, and automatic image-to-code conversion is realized, improving the efficiency and quality of code generation.
Patent Information
- Application Number
- CN202411978111.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
In front-end development, many complex design concepts are difficult to accurately describe through plain text, especially when designing user interfaces, text requires a lot of space to fully convey the needs, resulting in limited code generation efficiency and quality.
A training method of a neural network model is adopted, which includes a visual encoder, an adapter and a language model. Through the main training stage, images are acquired, codes are marked, and model parameters are adjusted to realize automatic image-to-code conversion.
This method gives the code generation model the ability to interpret and understand images, so that the model can generate code based on design drawings or screenshots, thereby improving the accuracy and efficiency of code generation and reducing the complexity of manual encoding.
Smart Images

Figure CN119938024A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and in particular to a training method and device for a neural network model for generating code, a method and device for generating code based on a neural network model, as well as a computer device, a computer-readable storage medium, and a computer program product. Background Art
[0002] In the field of modern front-end development, as users' requirements for application interaction experience continue to increase, the complexity and efficiency requirements of front-end page development are also increasing. Developers need to face the demand for fast iteration of products, and how to improve the efficiency and quality of front-end code generation has become an important technical problem. With the rapid development of large language models (LLMs) and deep learning technology, automatic code generation technology has brought great potential for simplifying front-end development work. For example, tools such as GitHub Copilot have been widely used in software development, which can generate high-quality code based on text instructions provided by users, thereby greatly improving development efficiency. However, there are still limitations to achieving front-end code generation only through natural language instructions. Many complex design concepts are difficult to accurately describe in plain text, especially when it comes to user interface design, the text often takes a lot of space to fully convey the requirements. It would be beneficial to provide a mechanism to alleviate, mitigate or even eliminate one or more of the above problems. Summary of the invention
[0003] According to one aspect of the present disclosure, a training method for a neural network model for generating code is provided, wherein the neural network model includes a visual encoder, an adapter and a language model, and the training method includes a main training stage, which includes: acquiring a first image, wherein the first image includes a first page layout; annotating a first label value of the first image, wherein the first label value is used to represent the code for generating the first page layout, and wherein the first label value includes the index corresponding to each first word of multiple first words in the code in a preset dictionary; inputting the first image to the visual encoder to obtain a first feature vector of the first image output by the visual encoder; inputting the first feature vector to the adapter to obtain a second feature vector output by the adapter, wherein the dimension of the second feature vector is adapted to the language model; inputting the second feature vector to the language model to obtain a first code representation output by the language model for generating the first page layout, wherein the first code representation includes the index corresponding to each second word of multiple second words in the preset dictionary; calculating a first loss value based on the first code representation and the first label value; and adjusting parameters of the visual encoder, the adapter and the language model based on the first loss value.
[0004] According to another aspect of the present disclosure, there is provided a method for generating code based on a neural network model, comprising: obtaining a reference image, wherein the reference image contains a target page layout to be generated; inputting the reference image into a neural network model, obtaining a code representation output by the neural network model, wherein the code representation includes an index corresponding to each word in at least one word in a preset dictionary; based on the preset dictionary, determining a code corresponding to the code representation; and generating a target page based on the code, wherein the neural network model is obtained based on the aforementioned training method for a neural network model for generating code.
[0005] According to another aspect of the present disclosure, a training device for a neural network model for generating code is provided, wherein the neural network model includes a visual encoder, an adapter and a language model, and the training device includes a main training module, and the main training module includes: a first acquisition module, configured to acquire a first image, wherein the first image includes a first page layout; a first annotation module, configured to annotate a first label value of the first image, wherein the first label value is used to represent a code for generating the first page layout, and wherein the first label value includes an index corresponding to each first word of a plurality of first words in a preset dictionary; a second acquisition module, configured to input the first image into the visual encoder, and acquire the first image output by the visual encoder a first feature vector of the first page layout; a third acquisition module, configured to input the first feature vector into the adapter, and obtain the second feature vector output by the adapter, wherein the dimension of the second feature vector is adapted to the language model; a fourth acquisition module, configured to input the second feature vector into the language model, and obtain the first code representation output by the language model for generating the first page layout, wherein the first code representation includes the index corresponding to each second word of a plurality of second words in the preset dictionary; a calculation module, configured to calculate a first loss value based on the first code representation and the first label value; and an adjustment module, configured to adjust the parameters of the visual encoder, the adapter and the language model based on the first loss value.
[0006] According to another aspect of the present disclosure, there is provided an apparatus for generating code based on a neural network model, comprising: a fifth acquisition module, configured to acquire a reference image, wherein the reference image contains a target page layout to be generated; a sixth acquisition module, configured to input the reference image into the neural network model, and acquire a code representation output by the neural network model, wherein the code representation includes an index corresponding to each word in at least one word in a preset dictionary; a determination module, configured to determine a code corresponding to the code representation based on the preset dictionary; and a generation module, configured to generate a target page based on the code, wherein the neural network model is obtained according to the aforementioned training method for a neural network model for generating code.
[0007] According to another aspect of the present disclosure, a computer device is provided, comprising: at least one processor; and at least one memory on which a computer program is stored, wherein when the computer program is executed by the at least one processor, the at least one processor executes any one of the above methods.
[0008] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor is caused to perform any of the above methods.
[0009] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the processor is caused to perform any one of the above methods.
[0010] According to the embodiments of the present disclosure, a visual encoder, a visual-to-text adapter, and a language model are combined to construct a code generation model capable of processing images. This architecture can give the code generation model the ability to interpret and understand images, so that the model can generate code based on design drawings, screenshots, etc., for rendering and generating corresponding pages.
[0011] These and other aspects of the disclosure will be apparent from and elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Further details, features and advantages of the present disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0013] Figure 1 is a schematic diagram illustrating an example system in which the various methods described herein may be implemented according to an exemplary embodiment;
[0014] Figure 2 is a flow chart illustrating a method for training a neural network model for generating code according to an exemplary embodiment;
[0015] Figure 3 is a flow chart illustrating a partial process of a training method of a neural network model for generating code according to an exemplary embodiment;
[0016] Figure 4 is a flow chart illustrating a partial process of a training method of a neural network model for generating code according to an exemplary embodiment;
[0017] Figure 5 is a flow chart illustrating a partial process of a training method of a neural network model for generating code according to an exemplary embodiment;
[0018] Figure 6 is a schematic diagram illustrating an image for training a neural network model according to an exemplary embodiment;
[0019] Figure 7 is a flow chart illustrating a method of generating code based on a neural network model according to an exemplary embodiment;
[0020] Figure 8 is a schematic block diagram illustrating a training apparatus for a neural network model for generating code according to an exemplary embodiment;
[0021] Fig. 9 is a schematic block diagram illustrating an apparatus for generating code based on a neural network model according to an exemplary embodiment; and
[0022] Fig.10 is a block diagram illustrating an exemplary computer device that can be applied to the exemplary embodiments. DETAILED DESCRIPTION
[0023] In the present disclosure, unless otherwise specified, the use of the terms "first", "second", etc. to describe various elements is not intended to limit the positional relationship, timing relationship, or importance relationship of these elements, and such terms are only used to distinguish one element from another element. In some examples, the first element and the second element may refer to the same instance of the element, and in some cases, based on the description of the context, they may also refer to different instances.
[0024] The terms used in the description of various examples described in this disclosure are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element can be one or more. As used herein, the term "plurality" means two or more, and the term "based on" should be interpreted as "based at least in part on". In addition, the terms "and / or" and "at least one of..." cover any one of the listed items and all possible combinations.
[0025] In related technologies, neural network models can be used through natural language instructions to generate front-end code. However, many complex design concepts are difficult to accurately describe through plain text, especially when it comes to user interface design. Text often takes a lot of space to fully convey the requirements, and the accuracy of text description is required to be high.
[0026] In view of the above situation, the present disclosure proposes a training method and device for a neural network model for generating code, a method and device for generating code based on a neural network model, as well as a computer device, a computer-readable storage medium, and a computer program product.
[0027] In addition, in the technical solution of the present disclosure, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0028] Exemplary embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0029] Figure 1 is a schematic diagram illustrating an example system 100 in which the various methods described herein may be implemented, according to an example embodiment.
[0030] refer to Figure 1 The system 100 includes a client device 110 , a server 120 , and a network 130 that communicatively couples the client device 110 and the server 120 .
[0031] The client device 110 includes a display 114 and a client application (APP) 112 that can be displayed via the display 114. The client application 112 can be an application that needs to be downloaded and installed before running or a small program (liteapp) as a lightweight application. In the case where the client application 112 is an application that needs to be downloaded and installed before running, the client application 112 can be pre-installed on the client device 110 and activated. In the case where the client application 112 is a small program, the user 102 can directly run the client application 112 on the client device 110 by searching for the client application 112 in the host application (for example, by the name of the client application 112, etc.) or scanning the graphic code of the client application 112 (for example, a bar code, a QR code, etc.), without installing the client application 112. In some embodiments, the client device 110 can be any type of mobile computer device, including a mobile computer, a mobile phone, a wearable computer device (for example, a smart watch, a head-mounted device, including smart glasses, etc.) or other types of mobile devices. In some embodiments, client device 110 may alternatively be a stationary computer device, such as a desktop computer, a server computer, or other type of stationary computer device.
[0032] The server 120 is typically a server deployed by an Internet Service Provider (ISP) or an Internet Content Provider (ICP). The server 120 may represent a single server, a cluster of multiple servers, a distributed system, or a cloud server that provides basic cloud services (such as cloud databases, cloud computing, cloud storage, and cloud communications). It will be understood that although Figure 1 The server 120 is shown in FIG. 1 in communicating with only one client device 110 , but the server 120 may provide background services for multiple client devices simultaneously.
[0033] Examples of network 130 include a local area network (LAN), a wide area network (WAN), a personal area network (PAN), and / or a combination of communication networks such as the Internet. Network 130 can be a wired or wireless network. In some embodiments, the data exchanged through network 130 is processed using technologies and / or formats including hypertext markup language (HTML), extensible markup language (XML), etc. In addition, encryption technologies such as secure socket layer (SSL), transport layer security (TLS), virtual private network (VPN), Internet protocol security (IPsec) can also be used to encrypt all or some links. In some embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.
[0034] For the purpose of the embodiments of this disclosure, Figure 1 In the example of , the client application 112 may be an application for providing communication, decision-making, etc. for all parties involved in the real estate sale or rental process. Accordingly, the server 120 may be a server used together with such an application.
[0035] Figure 2 is a flow chart illustrating a method 200 of training a neural network model for generating code according to an exemplary embodiment.
[0036] The method 200 may be performed on a client device (e.g., Figure 1 The execution of each step of the method 200 may be performed at the client device 110 shown in FIG. Figure 1 In some embodiments, the method 200 may be performed on a server (e.g., Figure 1 In some embodiments, method 200 may be performed by a client device (e.g., client device 110) and a server (e.g., server 120) in combination. Hereinafter, each step of method 200 is described in detail by taking server 120 as an example.
[0037] refer to Figure 2 , the training method 200 of the neural network model for generating code includes steps 210 to 270.
[0038] Step 210: Acquire a first image, wherein the first image includes a first page layout;
[0039] Step 220: annotate the first image with a first tag value, wherein the first tag value is used to represent a code for generating the first page layout, and wherein the first tag value includes an index corresponding to each first word in a plurality of first words in the code in a preset dictionary;
[0040] Step 230: input the first image to the visual encoder, and obtain a first feature vector of the first image output by the visual encoder;
[0041] Step 240: Input the first feature vector to the adapter to obtain a second feature vector output by the adapter, wherein the dimension of the second feature vector is adapted to the language model;
[0042] Step 250: input the second feature vector into the language model, and obtain a first code representation output by the language model for generating the first page layout, wherein the first code representation includes an index corresponding to each second word in the preset dictionary among a plurality of second words;
[0043] Step 260: Calculate a first loss value based on the first code representation and the first label value; and
[0044] Step 270: Adjust parameters of the visual encoder, the adapter, and the language model based on the first loss value.
[0045] By combining the visual encoder, the vision-to-text adapter, and the language model, we built a code generation model that can process images. This architecture enables the code generation model to interpret and understand images, so that the model can generate code based on design drawings, screenshots, and other content for rendering and generating corresponding pages.
[0046] It is understandable that the first image acquired in step 210 may be a rendered page image, which includes a first page layout, so that the neural network model is trained using the first image to give the neural network model the ability to interpret and understand the page layout included in the image, and then generate a code that can be used to generate the first page layout. Exemplarily, the first page layout may include layout information such as a title, a picture, and text.
[0047] In step 220, a first tag value is annotated for the first image. Specifically, the first tag value is an index sequence of a code for generating a first page layout. In one example, the code is <div class="header"> , the mapping relationship between the index and vocabulary defined in the preset dictionary is as follows:
[0048] :1, <h1> :2,class:3,"header":4,< / h1> :5...
[0049] Therefore, the code <div class="header">The corresponding index sequence is [1,3,4,5], so that the first label value can be obtained. It can be understood that the dictionary index sequence of text tokens is to map each token in the natural language text to a numerical index in the vocabulary. This method helps to convert text or code into a machine-processable numerical form and is widely used in natural language processing tasks. The above mapping relationship is only an example, and the preset dictionary can be determined according to the specific word segmentation method and application scenario.
[0050] In one example, the first tag value is used to represent the React code for generating the first page layout, thereby training the neural network model to generate the corresponding React code.
[0051] In step 230, the first image is input to a visual encoder, so that the visual encoder is used to learn and understand the first image to output a first feature vector of the first image.
[0052] According to some embodiments, the visual encoder is used to decompose the input image into multiple image blocks and output a feature vector corresponding to each image block in the multiple image blocks.
[0053] Exemplarily, the visual encoder can decompose the input first image into multiple image blocks (e.g., image slices of fixed size), and generate a feature vector for each image block, which comprehensively reflects the local and global image features. The decomposition of image blocks helps to capture the local features of the image and enhance the detail expression ability of the visual encoder. The block method optimizes the visual encoder's understanding of the complex page layout in the first image.
[0054] According to some embodiments, the visual encoder is a visual encoder based on a Transformer structure. It is understandable that the Transformer structure can better model the global and local associations and improve the accuracy of feature extraction by virtue of the multi-head self-attention mechanism.
[0055] Step 240 uses an adapter to adjust the dimension of the first feature vector to adapt it to the input format of the language model and generate a second feature vector. Different models can usually process and receive vectors of different dimensions. For example, in one example, the dimension of the feature vector obtained after the first image is processed by the visual encoder is 768, while the language model can only receive vectors of 2048 dimensions. At this time, the adapter is required to adjust the dimension of the feature vector so that different models can be connected to each other.
[0056] Some models (such as visual encoders or language models) have very large structures, and the dimensions of the feature vectors obtained after pre-training may not match the task requirements. The adapter can adjust the dimensions of the features according to the requirements of the target task. For example, in an image-to-text task, the image features may need to be converted into dimensions that the text language model can understand. The adapter uses this conversion to ensure that the data can pass smoothly through different modules. By adjusting the dimensions, the adapter maintains the integrity of the feature information while reducing the burden of computation and storage, which helps improve the overall performance of the model.
[0057] According to some embodiments, the adapter can also be used to perform operations such as compressing and transforming the length of the input token. When processing natural language processing (NLP) and multimodal (e.g., image-text) tasks, the length of the input token sequence directly affects the computational cost of the model. In addition, the length of the token sequence that some language models can process also has some limitations. The adapter can reduce the number of tokens by aggregating local context information to adapt to the needs and performance of the language model. The adapter can flexibly adjust the degree of compression of the number of tokens according to different tasks and input data. For example, in multimodal tasks, the adapter can automatically determine the compression ratio based on the characteristics of the image and text, thereby improving the adaptability of the model under different tasks and data conditions.
[0058] According to some embodiments, the adapter comprises any of the following structures: one or more linear layer neural networks, a resampling perceptron-resampler and a query generator q-former.
[0059] In step 250, the language model outputs a first code representation based on the second feature vector for generating a first page layout. In step 260, a first loss value (such as a cross entropy loss) is calculated based on the generated first code representation and the first label value. Exemplarily, in step 270, the parameters of the visual encoder, adapter, and language model can be adjusted based on the first loss value through a back propagation algorithm to optimize the parameters of the visual encoder, adapter, and language model, respectively, thereby enhancing the collaborative capabilities of image feature extraction, feature conversion, and code generation, improving the accuracy and efficiency of image-based code generation, and realizing automatic conversion from vision to code.
[0060] According to some embodiments, the language model is a large language model based on a Transformer structure. The large language model is good at modeling natural language and code, and can significantly improve the grammatical correctness and semantic rationality of code generation.
[0061] Figure 3is a flow chart illustrating a partial process of a training method for a neural network model for generating code according to an exemplary embodiment. Figure 3 As shown, based on steps 210 to 270, the training method 200 for a neural network model for generating code further includes a pre-training stage, and the pre-training stage 300 includes:
[0062] Step 310: Acquire a second image, wherein the second image includes a second page layout;
[0063] Step 320: annotate the second image with a second tag value, wherein the second tag value is used to describe the second page layout;
[0064] Step 330: input the second image to the visual encoder, and obtain a third feature vector of the second image output by the visual encoder;
[0065] Step 340: Input the third feature vector to the adapter to obtain a fourth feature vector output by the adapter, wherein the dimension of the fourth feature vector is adapted to the language model;
[0066] Step 350: input the fourth feature vector into the language model to obtain a first layout representation output by the language model for describing the second page layout;
[0067] Step 360: Calculate a second loss value based on the first layout representation and the second label value; and
[0068] Step 370: Keep the parameters of the video encoder and the parameters of the language model unchanged, and adjust the parameters of the adapter based on the second loss value.
[0069] The pre-training phase is intended to initialize the vision-to-text adapter. The parameters of the code model and the visual encoder are locked, and the parameters of the vision-to-text adapter are only trained on the pre-trained dataset. Exemplarily, each data in the pre-trained dataset includes a set of image-text pairs, i.e., a second image and its corresponding second label value, and the second label value is used to describe the second page layout in the second image. The description label does not point to a specific code, but to the semantic information of the page layout.
[0070] Exemplarily, the first layout representation and the second tag value are both vector representations composed of index values corresponding to each word in a natural language used to describe the page layout.
[0071] Therefore, during the pre-training phase, the parameters of the visual encoder and language model are frozen, and only the parameters of the adapter are adjusted to ensure that the adapter's dimensional mapping of the feature vector is more efficient.
[0072] Figure 4 is a flow chart illustrating a partial process of a training method for a neural network model for generating code according to an exemplary embodiment. Figure 4 As shown, based on steps 210 to 270 and steps 310 to 370, the training method 200 for a neural network model for generating code further includes an image understanding training phase, and the image understanding training phase 400 includes:
[0073] Step 410: Acquire a third image, wherein the third image includes a third page layout;
[0074] Step 420: annotate a third tag value of the third image, wherein the third tag value is used to describe the third page layout;
[0075] Step 430: input the third image into the visual encoder, and obtain a fifth eigenvector of the third image output by the visual encoder;
[0076] Step 440: Input the fifth feature vector to the adapter to obtain a sixth feature vector output by the adapter, wherein the dimension of the sixth feature vector is adapted to the language model;
[0077] Step 450: input the sixth feature vector into the language model to obtain a second layout representation output by the language model for describing the layout of the third page;
[0078] Step 460: Calculate a third loss value based on the second layout representation and the third label value; and
[0079] Step 470: Based on the third loss value, at least adjust the parameters of the visual encoder and the adapter in the neural network model.
[0080] The image understanding training phase aims to enhance the model's ability to understand the front-end development page. The neural network model is trained on the fine-tuning dataset, and all model structure parameters (including the code model, the visual-to-text adapter, and the visual encoder) are involved in the training. For example, each piece of data in the fine-tuning dataset contains one or more page rendering images, corresponding to the layout and appearance description of the image, which is used to generate the code snippet of the corresponding page.
[0081] In one example, the visual encoder and adapter parameters in the neural network model are adjusted based on the third loss value. In another example, the parameters of the visual encoder, adapter, and language model in the neural network model are adjusted based on the third loss value.
[0082] Therefore, by fine-tuning the dataset, we can adjust the parameters of at least the visual encoder and adapter, further enhancing the overall capabilities of the model. We can establish a stronger connection between semantic description and code generation, and improve the model's ability to understand page layout information. Through joint parameter optimization, we can enhance the synergy between the visual encoder, adapter, and language model.
[0083] Figure 5 is a flow chart illustrating a partial process of a training method for a neural network model for generating code according to an exemplary embodiment. Figure 5 As shown, based on steps 210 to 270, the training method 200 for a neural network model for generating code further includes a preference alignment training phase, and the preference alignment training phase 500 includes:
[0084] Step 510: Acquire a fourth image, wherein the fourth image includes a fourth page layout;
[0085] Step 520: marking a preference response and a non-preference response of the fourth image, wherein the preference response is used to indicate a preference code for generating the fourth page layout, and the non-preference response is used to indicate a non-preference code for generating the fourth page layout;
[0086] Step 530: input the fourth image to the visual encoder, and obtain a seventh eigenvector of the fourth image output by the visual encoder;
[0087] Step 540: Input the seventh feature vector to the adapter to obtain an eighth feature vector output by the adapter, wherein the dimension of the eighth feature vector is adapted to the language model;
[0088] Step 550: input the eighth feature vector into the language model, and obtain a second code representation output by the language model for generating the fourth page layout;
[0089] Step 560, calculating a fourth loss value based on the second code representation, the preferred response and the non-preferred response; and
[0090] Step 570: Adjust parameters of the visual encoder, the adapter, and the language model based on the fourth loss value.
[0091] The preference alignment training phase aims to enable the neural network model to generate codes that meet user preferences based on image input. The model is trained on the preference dataset using the DPO (Direct Preference Optimization) method, and all model structures (including the code model, the vision-to-text adapter, and the visual encoder) parameters are involved in the training. The training data in the preference dataset includes one or more page images, the preferred response, i.e., the implementation code of the preference, and the non-preference response, i.e., the implementation code of the non-preference.
[0092] Therefore, the model parameters are further adjusted through the preference alignment training phase to make it more inclined to generate preference codes. The introduction of user preference information improves the quality of generated codes and user satisfaction. Through comparative learning, the model's ability to prioritize the generation of preference codes is significantly improved.
[0093] Figure 6 is a schematic diagram illustrating an image used for training a neural network model according to an exemplary embodiment. Figure 6 An exemplary image 600 is shown for training a neural network model. Figure 6 As shown, image 600 shows the actual rendering effect of the input box, that is, a page layout of a text input box with a placeholder text "Enter text". Exemplarily, the page description of image 600 may be "the page content is a single input box displayed in the center, with a placeholder text 'Enter text' in the input box". The exemplary code for generating this page layout is as follows:
[0094] CSS style code:
[0095] html{
[0096] font-size:16px;
[0097] }
[0098] .comp{
[0099] color:red;
[0100] }
[0101] React component code
[0102] import React from 'react';
[0103] import{Button,Form,Input,Message}from'semantic-ui-react';
[0104] const SubComponent=({handleInputChange,inputValue})=>(
[0105] <form>
[0106] <Input
[0107] type="text"
[0108] placeholder="Enter text"
[0109] onChange={handleInputChange}
[0110] value={inputValue}
[0111] / >
[0112] < / form> );
[0114] const MainComponent=()=>{
[0115] const[inputValue,setInputValue]=React.useState(”);
[0116] const handleInputChange=(e)=>{
[0117] setInputValue(e.target.value);
[0118] };
[0119] return(
[0120]
[0121] <SubComponent handleInputChange={handleInputChange}inputValue={inputValue} / >
[0122] {inputValue&&<Message content={`You entered:${inputValue}`} / >}
[0123] );
[0125] };
[0126] export default MainComponent;
[0127] In one example, during the training of a neural network model, image 600 and a page layout description of image 600 "the page content is a single input box displayed in the center, and there is a placeholder text 'Enter text' in the input box" can be input into a visual encoder to obtain the page layout description of image 600 output by the language model to pre-train the model or perform image understanding training.
[0128] In one example, during the training of the neural network model, the image 600 and the code representation corresponding to the above code, that is, the label value of the image 600, can be input into the visual encoder to obtain the code representation of the code output by the language model for generating the page layout of the image 600, so as to perform main training on the model, so that the model can generate code based on design drawings, screenshots and other content for rendering and generating corresponding pages.
[0129] According to another aspect of the present disclosure, a method for generating code based on a neural network model is provided. Figure 7 As shown, the method of generating code based on the neural network model includes:
[0130] Step 710: Acquire a reference image, wherein the reference image includes a target page layout to be generated;
[0131] Step 720: input the reference image into a neural network model, and obtain a code representation output by the neural network model, wherein the code representation includes an index corresponding to each word in at least one word in a preset dictionary;
[0132] Step 730: Determine the code corresponding to the code representation based on the preset dictionary; and
[0133] Step 740: Generate a target page based on the code, wherein the neural network model is trained according to a training method for a neural network model used to generate code.
[0134] The reference image is input into the neural network model trained by the aforementioned training method to generate a code representation (index sequence), and then the code representation is converted into code based on the preset dictionary to generate code using the neural network model. In step 740, the target page layout is restored using the generated code. Thus, based on the trained neural network model, the image can be efficiently converted into code to achieve automatic code generation. The accuracy and efficiency of code generation are improved, and the complexity of manual coding is reduced.
[0135] According to another aspect of the present disclosure, a training device for a neural network model for generating code is provided, wherein the neural network model includes a visual encoder, an adapter, and a language model.
[0136] like Figure 8As shown, the training device 800 for the neural network model for generating code includes a main training module 801, and the main training module 801 includes: a first acquisition module 802, configured to acquire a first image, wherein the first image includes a first page layout; a first annotation module 803, configured to annotate a first label value of the first image, wherein the first label value is used to represent the code for generating the first page layout, and wherein the first label value includes the index corresponding to each first word in a plurality of first words in a preset dictionary; a second acquisition module 804, configured to input the first image into the visual encoder, and acquire a first feature vector of the first image output by the visual encoder; a third acquisition module 805, configured To input the first feature vector into the adapter, obtain the second feature vector output by the adapter, wherein the dimension of the second feature vector is adapted to the language model; a fourth acquisition module 806 is configured to input the second feature vector into the language model, obtain the first code representation output by the language model for generating the first page layout, wherein the first code representation includes the index corresponding to each second word of a plurality of second words in the preset dictionary; a calculation module 807 is configured to calculate a first loss value based on the first code representation and the first label value; and an adjustment module 808 is configured to adjust the parameters of the visual encoder, the adapter and the language model based on the first loss value.
[0137] It should be understood that Figure 8 The various modules of the apparatus 800 shown in FIG. 8 can be used in conjunction with the reference Figure 2 The steps in the method 200 described above correspond to each other. Therefore, the operations, features and advantages described above for the method 200 are also applicable to the apparatus 800 and the modules included therein. For the sake of brevity, some operations, features and advantages are not described in detail here.
[0138] According to another aspect of the present disclosure, a device for generating code based on a neural network model is provided. Fig. 9 As shown, the device 900 for generating code based on a neural network model includes: a fifth acquisition module 901, configured to acquire a reference image, wherein the reference image contains a target page layout to be generated; a sixth acquisition module 902, configured to input the reference image into the neural network model, and acquire a code representation output by the neural network model, wherein the code representation includes an index corresponding to each word in at least one word in a preset dictionary; a determination module 903, configured to determine a code corresponding to the code representation based on the preset dictionary; and a generation module 904, configured to generate a target page based on the code, wherein the neural network model is trained according to a training method for a neural network model for generating code.
[0139] It should be understood that Fig. 9 The various modules of the apparatus 900 shown in FIG. Figure 7 The steps in the method 700 described above correspond to each other. Therefore, the operations, features and advantages described above for the method 700 are also applicable to the apparatus 900 and the modules included therein. For the sake of brevity, some operations, features and advantages are not described in detail here.
[0140] Although specific functions are discussed above with reference to specific modules, it should be noted that the functions of the various modules discussed herein may be divided into multiple modules, and / or at least some functions of multiple modules may be combined into a single module. The specific module discussed herein performing an action includes the specific module itself performing the action, or alternatively the specific module calling or otherwise accessing another component or module that performs the action (or performs the action in conjunction with the specific module). Therefore, the specific module that performs an action may include the specific module itself that performs the action and / or another module that the specific module calls or otherwise accesses to perform the action. As used herein, the phrase "entity A initiates action B" may mean that entity A issues an instruction to perform action B, but entity A itself does not necessarily perform the action B.
[0141] It should also be understood that various techniques may be described herein in the general context of software hardware elements or program modules. Figure 8 and Fig. 9 The various modules described can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these modules can be implemented as computer program code / instructions, which are configured to be executed in one or more processors and stored in a computer-readable storage medium. Alternatively, these modules can be implemented as hardware logic / circuits. For example, in some embodiments, one or more of the above modules can be implemented together in a system on chip (System on Chip, SoC). SoC may include an integrated circuit chip (which includes a processor (e.g., a central processing unit (CPU), a microcontroller, a microprocessor, a digital signal processor (DSP), etc.), a memory, one or more communication interfaces, and / or one or more components in other circuits), and may optionally execute the received program code and / or include embedded firmware to perform functions.
[0142] According to another aspect of the present disclosure, a computer device is provided, which includes a memory, a processor, and a computer program stored in the memory. When the computer program is executed by the processor, the processor executes the computer program to implement the steps of any method embodiment described above.
[0143] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any method embodiment described above are implemented.
[0144] According to another aspect of the present disclosure, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the steps of any method embodiment described above are implemented.
[0145] In the following, combined Fig.10 Illustrative examples of such a computer device, non-transitory computer-readable storage medium, and computer program product are described.
[0146] Fig.10 An example configuration of a computer device 1000 that can be used to implement the methods described herein is shown. For example, Figure 1 The server 120 and / or the client device 110 shown in the figure may include an architecture similar to the computer device 900. The training device 800 for the neural network model for generating code or the device 900 for generating code based on the neural network model may also be implemented in whole or in part by the computer device 1000 or a similar device or system.
[0147] Computer device 1000 can be a variety of different types of devices. Examples of computer device 1000 include, but are not limited to, desktop computers, server computers, laptop or netbook computers, mobile devices (e.g., tablet computers, cellular or other wireless phones (e.g., smart phones), notepad computers, mobile stations), wearable devices (e.g., glasses, watches), entertainment devices (e.g., entertainment appliances, set-top boxes communicatively coupled to display devices, game consoles), televisions or other display devices, automotive computers, and the like.
[0148] The computer device 1000 may include at least one processor 1002, memory 1004, communication interface(s) 1006, a display device 1008, other input / output (I / O) devices 1010, and one or more mass storage devices 1012, which can communicate with each other, such as via a system bus 1014 or other appropriate connection.
[0149] The processor 1002 may be a single processing unit or multiple processing units, all of which may include a single or multiple computing units or multiple cores. The processor 1002 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operating instructions. Among other capabilities, the processor 1002 may be configured to obtain and execute computer-readable instructions stored in the memory 1004, mass storage device 1012, or other computer-readable media, such as program code of an operating system 1016, program code of an application program 1018, program code 1022 of other programs 1020, and the like.
[0150] The memory 1004 and the mass storage device 1012 are examples of computer-readable storage media for storing instructions that are executed by the processor 1002 to implement the various functions described above. For example, the memory 1004 may generally include both volatile memory and non-volatile memory (e.g., RAM, ROM, etc.). In addition, the mass storage device 1012 may generally include a hard drive, a solid-state drive, a removable medium, including external and removable drives, a memory card, a flash memory, a floppy disk, an optical disk (e.g., a CD, a DVD), a storage array, a network attached storage, a storage area network, etc. The memory 1004 and the mass storage device 1012 may all be collectively referred to herein as memory or computer-readable storage media, and may be a non-transitory medium capable of storing computer-readable, processor-executable program instructions as computer program code, which may be executed by the processor 1002 as a specific machine configured to implement the operations and functions described in the examples herein.
[0151] A number of programs may be stored on the mass storage device 1012. These programs include an operating system 1016, one or more application programs 1018, other programs 1020, and program data 1022, and they may be loaded into the memory 1004 for execution. Examples of such applications or program modules may include, for example, computer program logic (e.g., computer program code or instructions) for implementing the following components / functions: the client application 112, the method 200 and / or the method 700, and any modules or steps thereof, and / or additional embodiments described herein.
[0152] Although in Fig.101000, but modules 1016, 1018, 1020, and 1022, or portions thereof, may be implemented using any form of computer-readable media accessible by computer device 1000. As used herein, "computer-readable media" includes at least two types of computer-readable media, namely, computer-readable storage media and communication media.
[0153] Computer-readable storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information, such as computer-readable instructions, data structures, program modules or other data. Computer-readable storage media include but are not limited to RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD), or other optical storage device, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, or any other non-transmission medium that can be used to store information for access by a computer device. In contrast, communication media can embody computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transmission mechanism. Computer-readable storage media defined herein do not include communication media.
[0154] One or more communication interfaces 1006 are used to exchange data with other devices, such as through a network, a direct connection, etc. Such communication interfaces can be one or more of the following: any type of network interface (e.g., a network interface card (NIC)), a wired or wireless (such as IEEE 802.11 wireless LAN (WLAN)) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a cellular network ... TM The communication interface 1006 may facilitate communication within a variety of network and protocol types, including wired networks (e.g., LAN, cable, etc.) and wireless networks (e.g., WLAN, cellular, satellite, etc.), the Internet, etc. The communication interface 1006 may also provide for communication with external storage devices (not shown) such as in a storage array, network attached storage, storage area network, etc.
[0155] In some examples, a display device 1008 such as a monitor may be included for displaying information and images to the user. Other I / O devices 1010 may be devices that receive various inputs from the user and provide various outputs to the user, and may include a touch input device, a gesture input device, a camera, a keyboard, a remote control, a mouse, a printer, an audio input / output device, and the like.
[0156] The technology described herein can be supported by these various configurations of the computer device 1000, and is not limited to the specific examples of the technology described herein. For example, the functionality can also be implemented in whole or in part on the "cloud" by using a distributed system. The cloud includes and / or represents a platform for resources. The platform abstracts the underlying functions of the hardware (e.g., server) and software resources of the cloud. Resources can include applications and / or data that can be used when performing computing processing on a server away from the computer device 1000. Resources can also include services provided over the Internet and / or through a subscriber network such as a cellular or Wi-Fi network. The platform can abstract resources and functions to connect the computer device 1000 to other computer devices. Therefore, the implementation of the functions described herein can be distributed throughout the cloud. For example, the functions can be implemented partially on the computer device 1000 and partially through a platform that abstracts the functions of the cloud.
[0157] Although the present disclosure has been illustrated and described in detail in the drawings and the foregoing description, such illustration and description should be considered illustrative and schematic, not restrictive; the present disclosure is not limited to the disclosed embodiments. By studying the drawings, the disclosure and the appended claims, those skilled in the art will be able to understand and implement variations to the disclosed embodiments when practicing the claimed subject matter. In the claims, the word "comprising" does not exclude other elements or steps that are not listed, the indefinite article "a" or "an" does not exclude a plurality, the term "plurality" means two or more, and the term "based on" should be interpreted as "based at least in part on". The mere fact that certain measures are recorded in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
Claims
1. A method for training a neural network model for generating code, wherein: The neural network model includes a visual encoder, an adapter and a language model, and the training method includes a main training phase, which includes: Acquire a first image, wherein the first image includes a first page layout; Annotate the first image with a first tag value, wherein the first tag value is used to represent a code for generating the first page layout, and wherein the first tag value includes an index corresponding to each first word in a plurality of first words in the code in a preset dictionary; Inputting the first image into the visual encoder, and obtaining a first feature vector of the first image output by the visual encoder; Inputting the first feature vector into the adapter, obtaining a second feature vector output by the adapter, wherein a dimension of the second feature vector is adapted to the language model; Inputting the second feature vector into the language model, obtaining a first code representation output by the language model for generating the first page layout, wherein the first code representation includes an index corresponding to each second word in the preset dictionary among a plurality of second words; Calculating a first loss value based on the first code representation and the first label value; and Based on the first loss value, parameters of the visual encoder, the adapter, and the language model are adjusted.
2. The method according to claim 1, wherein: The method further comprises a pre-training stage, wherein the pre-training stage comprises: Acquire a second image, wherein the second image includes a second page layout; Annotate the second image with a second tag value, wherein the second tag value is used to describe the second page layout; Inputting the second image into the visual encoder, and obtaining a third feature vector of the second image output by the visual encoder; Inputting the third feature vector into the adapter to obtain a fourth feature vector output by the adapter, wherein a dimension of the fourth feature vector is adapted to the language model; Inputting the fourth feature vector into the language model, and acquiring a first layout representation output by the language model for describing the second page layout; Calculating a second loss value based on the first layout representation and the second label value; and The parameters of the video encoder and the parameters of the language model are kept unchanged, and the parameters of the adapter are adjusted based on the second loss value.
3. The method according to claim 2, wherein: After the pre-training stage and before the main training stage, the method further includes an image understanding training stage, which includes: Acquire a third image, wherein the third image includes a third page layout; marking a third tag value of the third image, wherein the third tag value is used to describe the third page layout; Inputting the third image into the visual encoder, and obtaining a fifth eigenvector of the third image output by the visual encoder; Inputting the fifth feature vector into the adapter to obtain a sixth feature vector output by the adapter, wherein a dimension of the sixth feature vector is adapted to the language model; Inputting the sixth feature vector into the language model, and acquiring a second layout representation output by the language model for describing the layout of the third page; Calculating a third loss value based on the second layout representation and the third label value; and Based on the third loss value, at least parameters of the visual encoder and the adapter in the neural network model are adjusted.
4. The method according to any one of claims 1 to 3, wherein: After the main training phase, the method further includes a preference alignment training phase, and the preference alignment training phase includes: Acquire a fourth image, wherein the fourth image includes a fourth page layout; marking a preference response and a non-preference response of the fourth image, wherein the preference response is used to indicate a preference code for generating the fourth page layout, and the non-preference response is used to indicate a non-preference code for generating the fourth page layout; Inputting the fourth image into the visual encoder, and obtaining a seventh eigenvector of the fourth image output by the visual encoder; Inputting the seventh feature vector into the adapter to obtain an eighth feature vector output by the adapter, wherein a dimension of the eighth feature vector is adapted to the language model; Inputting the eighth feature vector into the language model, and acquiring a second code representation output by the language model for generating the fourth page layout; calculating a fourth loss value based on the second code representation, the preferred response, and the non-preferred response; and Based on the fourth loss value, parameters of the visual encoder, the adapter, and the language model are adjusted.
5. The method according to any one of claims 1 to 4, wherein: The visual encoder is used to decompose the input image into multiple image blocks and output a feature vector corresponding to each image block in the multiple image blocks.
6. The method according to any one of claims 1 to 5, wherein: The visual encoder is a visual encoder based on the Transformer structure.
7. The method according to any one of claims 1 to 6, wherein: The adapter comprises any one of the following structures: A neural network with one or more linear layers, a perceptron-resampler and a query generator q-former.
8. The method according to any one of claims 1 to 7, wherein: The language model is a large language model based on the Transformer structure.
9. A method for generating code based on a neural network model, comprising: Acquire a reference image, wherein the reference image includes a target page layout to be generated; Inputting the reference image into a neural network model, and obtaining a code representation output by the neural network model, wherein the code representation includes an index corresponding to each word in at least one word in a preset dictionary; Based on the preset dictionary, determining a code corresponding to the code representation; and Generate a target page based on the code, Wherein, the neural network model is trained according to the method described in any one of claims 1-8.
10. A training device for a neural network model for generating code, wherein: The neural network model includes a visual encoder, an adapter and a language model, and the training device includes a main training module, which includes: A first acquisition module is configured to acquire a first image, wherein the first image includes a first page layout; A first labeling module is configured to label the first image with a first label value, wherein the first label value is used to represent a code for generating the first page layout, and wherein the first label value includes an index corresponding to each first word in a plurality of first words in a preset dictionary; A second acquisition module is configured to input the first image into the visual encoder, and acquire a first feature vector of the first image output by the visual encoder; A third acquisition module is configured to input the first feature vector into the adapter, and acquire a second feature vector output by the adapter, wherein a dimension of the second feature vector is adapted to the language model; a fourth acquisition module, configured to input the second feature vector into the language model, and acquire a first code representation output by the language model for generating the first page layout, wherein the first code representation includes an index corresponding to each second word in the preset dictionary among a plurality of second words; a calculation module configured to calculate a first loss value based on the first code representation and the first label value; and An adjustment module is configured to adjust parameters of the visual encoder, the adapter and the language model based on the first loss value.
11. A device for generating code based on a neural network model, comprising: A fifth acquisition module is configured to acquire a reference image, wherein the reference image includes a target page layout to be generated; a sixth acquisition module, configured to input the reference image into a neural network model, and acquire a code representation output by the neural network model, wherein the code representation includes an index corresponding to each word in at least one word in a preset dictionary; a determination module configured to determine a code corresponding to the code representation based on the preset dictionary; and A generating module is configured to generate a target page based on the code, Wherein, the neural network model is trained according to the method described in any one of claims 1-8.
12. A computer device comprising: at least one processor; as well as at least one memory having a computer program stored thereon, When the computer program is executed by the at least one processor, the at least one processor executes the method according to any one of claims 1 to 9.
13. A computer-readable storage medium having instructions stored thereon, wherein when the instructions are executed by one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 9.
14. A computer program product comprising instructions, which, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 9.
Citation Information
Cited By
Front-end code generation method and device, equipment, storage medium and program product
CN120848883A