Low-code visual programming system based on bidirectional AST dynamic analysis
By analyzing user code modification features and training code enhancement models, optimizing code generation of low-code visual programming system, solving the problems of large forward parsing overhead and frequent user modifications, and achieving efficient code generation and consistency.
Patent Information
- Application Number
- CN202510781255.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-12
AI Technical Summary
In the process of low-code development, low-code visual programming systems based on bidirectional AST dynamic analysis have problems such as excessive overhead during forward resolution and frequent user code modifications, resulting in increased programming cost.
Through the user code optimization module, we obtain the code mark sequence and AST structure before and after the user's modification, analyze the changes in similarity and complexity, determine the code modification characteristics and weights, train the code enhancement model, and optimize the code automatically generated by the system.
It effectively reduces the cost of visual programming and forward parsing overhead, reduces multiple modifications of the code by users, and improves development efficiency and code consistency.
Smart Images

Figure CN120335861A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital data processing, and particularly relates to a low-code visual programming system based on bidirectional AST dynamic parsing. Background Art
[0002] With the development of information technology, the demand for software development has been increasing day by day. Traditional programming methods require developers to have professional programming knowledge and skills, and have defects such as a large amount of code writing, a long development cycle, and high costs. To address this issue, low-code development platforms have emerged. It is a software development platform that quickly constructs application systems through a graphical interface and a small amount of code, aiming to reduce the complexity and cost of software development and improve development efficiency and flexibility.
[0003] Currently, in the process of low-code development, a low-code visual programming system based on bidirectional AST (Abstract Syntax Tree) dynamic parsing has certain advantages in improving development efficiency and reducing technical thresholds. The bidirectional AST dynamic parsing mechanism is divided into forward parsing and reverse parsing. Forward parsing is to convert visual operations (dragging, configuration) into AST nodes in real time to generate executable code; reverse parsing is to dynamically update the AST and synchronize it to the visual interface when the code is directly modified to achieve bidirectional consistency. During forward parsing, since the code automatically generated by the system may not be deeply optimized, when the automatically generated executable code does not meet the user's requirements, the user will make multiple changes by themselves, thereby increasing the cost of the entire visual programming and the overhead of forward parsing. In addition, during the bidirectional AST dynamic parsing process, each update by the user will perform a global update at the visual interface and code levels, resulting in excessive overhead during the dynamic parsing process. Summary of the Invention
[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a low-code visual programming system based on bidirectional AST dynamic parsing, and the specific technical solutions adopted are as follows: In a first aspect, the present invention provides a low-code visual programming system based on bidirectional AST dynamic parsing, and the system includes a user code optimization module, and the user code optimization module includes: A data acquisition module, configured to: acquire the token sequences of the code before and after each modification and the AST structure formed by the token sequences when the user modifies the code generated by the system in the code editor; A modification before and after index acquisition module, configured to: determine the similarity between the AST structures before and after each modification and the complexity change index of the token sequences before and after each modification; A modified feature acquisition module, configured to: determine code modification features corresponding to each modification according to the similarity and the complexity change index; A weight acquisition module, configured to: determine learning weights of code modification features corresponding to each modification according to the code modification features; A model acquisition module, configured to: train a constructed code enhancement model based on the code before and after several modifications and the learning weights of the code modification features, so as to obtain a trained code enhancement model; A code optimization processing module, configured to: optimize the code automatically generated by the visualization operating system for the user each time during the drag-and-drop interface construction process by using the trained code enhancement model.
[0005] Combined with the first aspect above, in some possible implementation manners, the before-and-after modification index acquisition module includes: A similarity acquisition unit, configured to: determine the similarity between the AST structures before and after each modification according to the difference in the number of nodes and the difference in the number of layers between the AST structures before and after each modification; A complexity change index acquisition unit, configured to: determine the number of code complexity skeletons and the code complexity of the token sequence, and determine the complexity change index of the token sequence before and after each modification according to the difference in the number of code complexity skeletons between the token sequences before and after each modification and the difference in the code complexity.
[0006] Combined with the first aspect above, in some possible implementation manners, the similarity acquisition unit includes: A first processed value acquisition subunit, configured to: determine the absolute value of the first difference in the number of nodes between the AST structures before and after each modification, and perform a negative correlation mapping process on the absolute value of the first difference to obtain a first negatively correlated mapped value; A second processed value acquisition subunit, configured to: determine the absolute value of the second difference in the number of layers between the AST structures before and after each modification, and perform a negative correlation mapping process on the absolute value of the second difference to obtain a second negatively correlated mapped value; A similarity acquisition subunit, configured to: determine the product of the first negatively correlated mapped value and the second negatively correlated mapped value as the similarity between the AST structures before and after each modification.
[0007] Combined with the first aspect above, in some possible implementation manners, the complexity change index acquisition unit includes: A code complexity skeleton number acquisition subunit, configured to: determine the total number of the first keywords and delimiters included in the token sequence as the number of code complexity skeletons of the token sequence; A code complexity acquisition subunit, configured to: match the values in the value fields corresponding to the keywords and delimiters included in the token sequence with the values in the value fields corresponding to the keywords in a preset keyword set and the delimiters in a delimiter set, and determine a second total quantity of the matching keywords and delimiters; determine a ratio of the second total quantity to the number of code complexity skeletons as the code complexity of the token sequence.
[0008] Combined with the first aspect above, in some possible implementation manners, the complexity change index acquisition unit further includes: A skeleton quantity difference acquisition subunit, configured to: determine an absolute value of a third difference between the numbers of code complexity skeletons of the token sequences before and after each modification, and obtain a code complexity skeleton quantity difference; A mapped value acquisition subunit, configured to: determine a positive correlation mapped value of a difference between the code complexity of the token sequence before each modification and the code complexity of the token sequence after the modification; A complexity change index acquisition subunit, configured to: perform a normalization process on a product of the code complexity skeleton quantity difference and the positive correlation mapped value, and obtain a complexity change index of the token sequences before and after each modification.
[0009] Combined with the first aspect above, in some possible implementation manners, the modification feature acquisition module is configured to: determine a ratio of the similarity and the complexity change index before and after each modification as a code modification feature corresponding to each modification.
[0010] Combined with the first aspect above, in some possible implementation manners, the weight acquisition module includes: An accumulated value acquisition unit, configured to: determine an accumulated value of all the code modification features; A learning weight acquisition unit, configured to: determine a ratio of each code modification feature to the accumulated value as a code modification feature learning weight corresponding to each modification.
[0011] Combined with the first aspect above, in some possible implementation manners, the model acquisition module includes: A loss function determination unit, configured to: determine a loss function in the training process of the code enhancement model; A model training unit, configured to: use the code before each modification as an input of the code enhancement model, use the code after each modification as an output of the code enhancement model, and set a weight value of the code before each modification to a product of the code modification feature learning weight and a preset overall weight, and train the code enhancement model to obtain a trained code enhancement model.
[0012] Combined with the first aspect above, in some possible implementation manners, the code enhancement model is a neural network model.
[0013] Combined with the first aspect above, in some possible implementation manners, the neural network model is a CodeT5 model.
[0014] In a second aspect, the present invention also provides a method for optimizing user code applied to a low-code visual programming system based on bidirectional AST dynamic parsing. The method includes: Obtain the token sequences of the code before and after each modification and the AST structure formed by the token sequences when the user modifies the code generated by the system in the code editor. Determine the similarity between the AST structures before and after each modification, and the complexity change index of the token sequences before and after each modification. Determine the code modification features corresponding to each modification according to the similarity and the complexity change index. Determine the code modification feature learning weights corresponding to each modification according to the code modification features. Train the constructed code enhancement model based on the code before and after several modifications and the code modification feature learning weights to obtain a trained code enhancement model. Use the trained code enhancement model to optimize the code automatically generated by the system for each user's visual operation during the drag-and-drop interface construction process.
[0015] The present invention has the following beneficial effects: The present invention analyzes the token sequences of the code before and after each user modification and the AST structure formed by the token sequences, and determines the code modification feature learning weights corresponding to each modification according to the similarity between the AST structures before and after each modification and the complexity change index of the token sequences before and after each modification. Then, based on the code before and after several modifications and the code modification feature learning weights, the constructed code enhancement model is trained to enable the code enhancement model to learn the user's code modification features, obtaining a trained code enhancement model. Finally, the trained code enhancement model is used to optimize the code automatically generated by the system for each user's visual operation during the drag-and-drop interface construction process. By deeply optimizing the code automatically generated by the system based on the user's code modification style, the present invention avoids the user from modifying the code multiple times by himself, effectively reducing the cost of the entire visual programming and the overhead of forward parsing. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 Schematic diagram of the composition of the AST structure according to an embodiment of the present invention; Figure 2 Schematic diagram of the structure of the user code optimization module according to an embodiment of the present invention; Figure 3 Schematic diagram before code optimization according to an embodiment of the present invention; Figure 4 Schematic diagram after code optimization according to an embodiment of the present invention; Figure 5 Flowchart of the steps of a user code optimization method for a low-code visual programming system applied to bidirectional AST dynamic parsing according to an embodiment of the present invention. Detailed implementation manners
[0018] To clearly illustrate the technical features of this solution, the present invention will be elaborated in detail below through specific implementation manners in combination with the drawings.
[0019] The embodiments of the present invention will be described in more detail below with reference to the drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.
[0020] It should be understood that the steps recited in the method embodiments of the present invention can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.
[0021] As used herein, the term "comprising" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0022] It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order of functions executed by these devices, modules or units or their interdependent relationships.
[0023] In the embodiments of the present invention, although operations or steps are described in a specific order in the drawings, it should not be understood that these operations or steps are required to be executed in the specific order shown or in a serial order, or that all the operations or steps shown are required to be executed to obtain the desired result. In the embodiments of the present invention, these operations or steps can be executed serially; they can also be executed in parallel; or a part of these operations or steps can be executed.
[0024] At the same time, it can be understood that the data involved in the technical solution of the present invention (including but not limited to the data itself, the acquisition or use of data) should comply with the requirements of corresponding laws, regulations and related regulations. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs, and all parameters or indicators in the formulas involved in the present invention are numerical values after normalization to eliminate the influence of dimensions.
[0025] Next, a low-code visual programming system based on bidirectional AST dynamic parsing provided by the embodiments of the present invention will be introduced in detail with reference to the accompanying drawings.
[0026] The embodiments of the present invention provide a low-code visual programming system based on bidirectional AST dynamic parsing. The low-code visual system includes a front-end layer, a parsing layer and an execution layer. Among them, the front-end layer includes a visual editor (including a component library or a logic canvas), a code editor (supporting syntax highlighting and real-time AST feedback); the parsing layer includes an AST generator, a bidirectional mapping module, and a dynamic optimization engine. The execution layer includes: a JIT compiler, a distributed task scheduler.
[0027] The process of the low-code visual system implementing bidirectional AST dynamic parsing includes forward parsing and reverse parsing: Forward parsing: Generate an AST according to visual operations. That is, when the user drags and drops a component, an AST node corresponding to the component type is generated and the attributes are filled. An AST is composed of nodes, and each node represents a syntax structure. Nodes usually include type, value and other relevant information.
[0028] Reverse parsing: When the user modifies the code, perform AST reconstruction on the code and synchronously display it on the visual editor.
[0029] The low-code visual system for low-code project development mainly includes four stages: selecting a project template, visual design, learning the user code style, and executing user visual operations. The following is a detailed introduction to these four stages: 1. Selecting a project template: Select a preset project template (such as a Web application, a mobile form, IoT data processing, etc.) according to the required business scenario, and the system automatically generates an initial AST structure and an associated code framework. Figure 1 The composition schematic diagram of the AST structure is shown.
[0030] 2. Visual design: It is divided into drag-and-drop interface construction and code editor construction. Drag-and-drop interface construction focuses on visual operations and AST generation, while code editor construction focuses on two-way conversion between text and AST. The two achieve seamless cooperation through dynamic AST parsing. The following is a detailed introduction to the main implementation processes of drag-and-drop interface construction and code editor construction: (2.1) Drag-and-drop interface construction: The user drags UI elements (such as buttons, tables, input boxes) from the component library to the canvas, and the system automatically generates corresponding AST nodes (such as <uibutton id="btnSubmit">), and at the same time, component properties can be set and logic can be arranged on the canvas (such as conditional judgment, API call, loop control, etc.).
[0031] During the process of drag-and-drop interface construction, each visualization operation (such as moving components, modifying logic branches, etc.) will trigger a global update of the AST, and then it is necessary to perform forward parsing and global AST update, resulting in an excessive overhead in the overall dynamic parsing process. Therefore, the following operations are taken: For each visualization operation, first, determine the component corresponding to the operation and its corresponding dependencies. After the user performs a certain visualization operation, first obtain the specific name of this visualization operation through the event listening mechanism (such as: modify the properties of button 1). Then determine the component designed by this visualization operation and determine its corresponding dependencies. Taking the visualization operation of modifying the properties of button 1 as an example, determine that the component designed by this visualization operation is button 1, and the corresponding operation is to modify the properties (the operation corresponding to modifying the properties is the UI rendering operation). Therefore, the dependency of button 1 is UI rendering.
[0032] Then, update the AST, that is, locate the nodes corresponding to the component and its corresponding dependencies in the AST structure and perform the update.
[0033] Finally, after updating the AST, find the position of the corresponding component in the code editor, rewrite the corresponding class name and the corresponding function, and then obtain the updated code. Taking the visualization operation of modifying the properties of button 1 as an example, after updating the AST, find the position of button 1 in the code editor. For example, the name of button 1 is "btnSubmit". Search for the class name and the corresponding function containing "btnSubmit" in the code editor, and then only rewrite the corresponding class name and the corresponding function to obtain the updated code.
[0034] When the user does not actively update the global template or switch to code editor construction, each visualization operation is locally parsed forward in the above manner.
[0035] By locating the updated nodes of the AST in the forward parsing and reverse parsing processes as described above, the AST is locally updated, thereby effectively reducing the overhead in the dynamic parsing process.
[0036] Conversely, when the user actively updates the global template, the system automatically performs a global update on the AST and the code, facilitating the user's subsequent operations. When the user switches to code editor construction, it indicates that the user is dissatisfied with the effect of drag-and-drop interface construction at this time or the drag-and-drop page does not have the operations they need. At this time, in order to facilitate the accurate and timely display of the operation effect after the user operates in the code editor. When the user switches to code editor construction, the system automatically performs a global update on the AST and the code.
[0037] The above code update process is to perform a lightweight conversion on the updated AST through the Babel tool, and then generate executable code. Then the generated code is input into a pre-trained code enhancement model for code optimization, thereby obtaining optimized code. By optimizing the code, the overall code style of the project and the user style can be kept as unified as possible, avoiding the user from modifying the code multiple times by themselves, effectively reducing the cost of the entire visual programming and the overhead of forward parsing, and facilitating the user to view the code and reducing the code maintenance cost in the later stage. Since the process of obtaining the pre-trained code enhancement model will be introduced in detail in the subsequent user code style learning, it will not be elaborated here.
[0038] (2.2) Code editor construction: When the user directly modifies the generated code in the code editor, the improved scope of the code is captured through the IDE plugin or the file listening service. Specifically, the improved class name and the corresponding function are obtained, and then the following operations are performed: First, the string of the code improved by the user is decomposed into a series of tokens by the lexical analyzer. This process is a prior art. For example, for a piece of code "net a = 1", the lexical analyzer will decompose it into the following tokens: { type: 'keyword', value: 'net'}, { type: 'identifier', value: 'a'}, {type: 'punctuator', value: '='}, { type: 'numeric', value: '1'}, { type: 'punctuator', value: ';'}.
[0039] Then, the compiler organizes the token sequence (subsequently denoted as token sequence 1) into an AST structure according to the analysis result of the lexical analyzer, denoted as AST structure 1. At the same time, the token sequence before the improvement corresponding to the code in the improved scope (subsequently denoted as token sequence 2) and the AST structure are obtained, and this AST structure is denoted as AST structure 2.
[0040] Finally, use an AST diff tool (such as GumTree) to compare the differences between AST structure 1 and AST structure 2, locate the changed nodes, and update the corresponding logical branches and UI components in the visualization canvas.
[0041] It should be understood that in special cases, when the modified code cannot be mapped to existing AST nodes (such as when the user introduces undefined variables during code editing), the user can choose to "override" or "retain the conflicting part".
[0042] 3. User code style learning: During the reverse parsing process, the code input by the user usually has an obvious personal style. Therefore, during forward parsing, when the executable code automatically generated by the system does not meet the user's requirements, the user will make changes by themselves, which increases the cost of the entire visual programming and the overhead of forward parsing. For this reason, a user code optimization module is set up in the low-code visual programming system. The user code optimization module is used to learn the user code style characteristics, obtain a pre-trained code enhancement model, and use the code enhancement model to optimize the code automatically generated by the system, so as to obtain optimized code that meets the user's requirements, avoiding the user's multiple modifications to the code, and thus helping to reduce the cost of the entire visual programming and the overhead of forward parsing.
[0043] Among them, as Figure 2 shown, the user code optimization module is specifically composed of six functional modules that can implement corresponding functions. These six functional modules are respectively a data acquisition module 100, a pre-and-post modification metric acquisition module 200, a modification feature acquisition module 300, a weight acquisition module 400, a model acquisition module 500, and a code optimization processing module 600. The following combines the implementation functions of these six functional modules to introduce in detail the specific implementation process of the user code optimization module learning the user code style characteristics, obtaining a pre-trained code enhancement model, and then using the code enhancement model to optimize the code automatically generated by the system.
[0044] The data acquisition module 100 is used to: acquire the token sequences of the code before and after each modification and the AST structure formed by the token sequences when the user modifies the code generated by the system in the code editor.
[0045] Specifically, during the construction of the code editor in the visual design stage, acquire the token sequences of the code before and after each modification when the user directly modifies the generated code in the code editor in the past, and the AST structure formed by the token sequences, that is, the token sequence 1 of the modified code and the AST structure 1 formed by the token sequence 1, and the token sequence 2 of the code before modification and the AST structure 2 formed by the token sequence 2 obtained in the above code editor construction part.
[0046] The pre- and post-modification metric acquisition module 200 is configured to: determine the similarity between the AST structures before and after each modification, and the complexity change metric of the token sequences before and after each modification.
[0047] Specifically, quantify the similarity between the AST structures before and after each modification, where this similarity is used to reflect the degree of difference between the AST structures before and after each modification. For example, the smaller the value of the similarity, the greater the difference between the AST structure corresponding to the newly input code after the user's modification and the original AST structure, and the greater the dissatisfaction of the user with the original code result or visual layout. At the same time, quantify the complexity change metric of the token sequences before and after each modification, where this complexity change metric is used to reflect the level of complexity of the newly input code after the user's modification compared to the code generated by the original system. For example, when the value of the complexity change metric is larger, it can indicate that the newly input code is more complex than the code generated by the original system.
[0048] The modification feature acquisition module 300 is configured to: determine the code modification features corresponding to each modification based on the similarity and the complexity change metric.
[0049] Specifically, determine the code modification features corresponding to each modification based on the similarity between the AST structures before and after each modification and the complexity change metric of the token sequences before and after each modification. These code modification features reflect the personal style situation of the newly input code after the user's modification. For example, when the AST structure of the newly input code after the user's modification is less similar to the AST structure of the code generated by the original system, and at the same time the newly input code after the modification is more complex than the code generated by the original system, it indicates that the newly input code after the user's modification is more likely to contain personal style, and the value of the corresponding code modification feature is larger.
[0050] The weight acquisition module 400 is configured to: determine the code modification feature learning weights corresponding to each modification based on the code modification features.
[0051] Specifically, every time the user constructs code in the code editor, that is, every time a code modification is made, the code modification features corresponding to that modification can be obtained, and then based on these code modification features, the code modification feature learning weights corresponding to each modification can be determined.
[0052] The model acquisition module 500 is configured to: train the constructed code enhancement model based on the code before and after several modifications and the code modification feature learning weights to obtain a trained code enhancement model.
[0053] Specifically, a training data set is constructed based on the code before and after multiple modifications and the learning weights of code modification features, and the constructed code enhancement model is trained to enable the code enhancement model to learn the personal style features of the user's modified code, so as to obtain a trained code enhancement model, that is, a pre-trained code enhancement model.
[0054] The code optimization processing module 600 is used to: optimize the code automatically generated by the visual operating system for the user each time during the drag-and-drop interface construction process by using the trained code enhancement model.
[0055] Specifically, the code automatically generated by the visual operating system for the user each time during the drag-and-drop interface construction process is input into the pre-trained code enhancement model, and the code enhancement model can output the optimized code that conforms to the user's personal style.
[0056] 4. Execute the user's visual operation: After the user completes the construction of the code editor, the overall AST structure is updated to adapt to the integrity and completeness of the code. Then, multiple automated tests such as logical verification and integration testing are performed on the AST by the low-code visual programming system based on bidirectional AST dynamic parsing. All the updates are deployed on the canvas and the cloud server, and thus the low-code visual programming based on bidirectional AST dynamic parsing is completed.
[0057] In a low-code visual programming system based on bidirectional AST dynamic parsing provided in this embodiment, by setting up a user code optimization module and using this user code optimization module, the features when the user directly modifies and generates code in the code editor are learned, and the learning weights of code modification features corresponding to each code modification are determined. Thus, based on the code before and after multiple modifications and the learning weights of code modification features, the code enhancement model is trained to enable the code enhancement model to learn the personal style features of the user's modified code, and the trained code enhancement model is used to optimize the code automatically generated by the system during the drag-and-drop interface construction process. Therefore, when the optimized code that conforms to the user's requirements and is adapted to the user's code style is obtained, the user's multiple modifications to the code are avoided, effectively reducing the cost of the entire visual programming and the overhead of forward parsing.
[0058] In a possible implementation manner, the above-mentioned before-and-after modification index acquisition module 200 includes: The similarity acquisition unit is used to: determine the similarity between the AST structures before and after each modification according to the differences in the number of nodes and the number of layers between the AST structures before and after each modification.
[0059] Specifically, the AST structure usually involves multiple layers and includes multiple nodes. The more layers and the more nodes there are, the more complex the AST structure is, and there may be more nested relationships. Therefore, analyze the difference in the number of nodes and the difference in the number of layers between the AST structures formed by the token sequences of the code before and after each modification to determine the similarity between the AST structures before and after each modification. When the difference in the number of nodes and the difference in the number of layers are smaller, it indicates that the similarity between the AST structures before and after the modification is higher.
[0060] In an example, the similarity acquisition unit includes: a first processing value acquisition subunit, configured to: determine the absolute value of the first difference in the number of nodes between the AST structures before and after each modification, perform a negative correlation mapping process on the absolute value of the first difference to obtain a first negatively correlated mapping processing value; a second processing value acquisition subunit, configured to: determine the absolute value of the second difference in the number of layers between the AST structures before and after each modification, perform a negative correlation mapping process on the absolute value of the second difference to obtain a second negatively correlated mapping processing value; a similarity acquisition subunit, configured to: determine the product of the first negatively correlated mapping processing value and the second negatively correlated mapping processing value as the similarity between the AST structures before and after each modification.
[0061] In this embodiment, the similarity between the AST structures before and after each modification is determined by the following formula: ; In the formula: represents the similarity between the AST structures before and after each modification, that is, the similarity between AST structure 1 and AST structure 2; represents the number of layers of the AST structure formed by the token sequence of the code after each modification, that is, the number of layers of AST structure 1; represents the number of layers of the AST structure formed by the token sequence of the code before each modification, that is, the number of layers of AST structure 2; represents the number of nodes of the AST structure formed by the token sequence of the code after each modification, that is, the number of nodes in AST structure 1; represents the number of nodes of the AST structure formed by the token sequence of the code before each modification, that is, the number of nodes in AST structure 2.
[0062] The complexity change index acquisition unit is configured to: determine the number of code complexity skeletons and the code complexity of the token sequence, and determine the complexity change index of the token sequence before and after each modification according to the difference in the number of code complexity skeletons and the difference in code complexity between the token sequences before and after each modification.
[0063] Specifically, since the AST structure cannot accurately indicate the degree of difference in the lexical types in the new code constructed during the user's code editing, and the type field in the token sequence corresponding to the user code represents the syntax type corresponding to each token. For example, keyword indicates that the token is a keyword, and identifier indicates that the token is an identifier. Among them, the keywords and delimiters related to control flow, scope, and nested logic belong to the tokens of complexity and are the skeleton of code complexity. The control flow keywords are keywords such as "if, else, while, for", etc. By analyzing the distribution of keywords and delimiters in the token sequence of the code before each modification, the number of code complexity skeletons and the code complexity of each token sequence are determined. Furthermore, based on the difference in the number of code complexity skeletons between the token sequences before and after each modification, and the difference in code complexity, the complexity change index of the token sequences before and after each modification is determined.
[0064] In an example, the above complexity change index obtaining unit includes: a code complexity skeleton number obtaining subunit, configured to: determine the first total number of keywords and delimiters included in the token sequence as the number of code complexity skeletons of the token sequence; a code complexity obtaining subunit, configured to: match the values in the value fields corresponding to the keywords and delimiters included in the token sequence with the values in the value fields corresponding to the keywords in the preset keyword set and the delimiters in the delimiter set, and determine the second total number of keywords and delimiters for which there is a match; and determine the ratio of the second total number to the number of code complexity skeletons as the code complexity of the token sequence.
[0065] Specifically, first, the preset control flow keyword set is , and the preset delimiter set is , where the value field corresponding to the token represents the specific token value.
[0066] Then, for each code modification, count the number of tokens marked as keywords , and the number of tokens marked as delimiters in the token sequence corresponding to the AST structure 1, that is, the token sequence of the modified code. Then, a total of skeletons representing code complexity are found. Denote as the number of code complexity skeletons of the token sequence corresponding to the AST structure 1, and record it as .
[0067] Next, the values in the value fields corresponding to tokens are compared with the keywords in the keyword set and the delimiter set Match the values in the value field corresponding to the delimiter. If there is a matching value, increment the value of the second total quantity by 1 to obtain the final second total quantity, which is the total number of matches. 。
[0068] Finally, determine the second total quantity. And the number of code complexity skeleton Take the ratio with the number of code complexity skeleton as the marker sequence corresponding to AST structure 1, that is, the code complexity of the marker sequence of the modified code. At this time, there is 。
[0069] Similarly, the number of keywords marked in the marker sequence corresponding to AST structure 2, that is, the marker sequence of the code before modification, And the number of delimiters marked Can be counted to obtain the number of code complexity skeletons of the marker sequence corresponding to AST structure 2. And the keywords and delimiters in the marker sequence and the keyword set And The final second total quantity And finally obtain the code complexity of the marker sequence corresponding to AST structure 2, that is, the marker sequence of the code before modification. At this time, there is 。
[0070] Since the smaller the difference between the number of code complexity skeletons and the second total quantity And The smaller the difference indicates that the user's input code has fewer non-key markers, and the complexity of the corresponding code modification is smaller. At the same time, the code complexity Compared with The smaller it is, the smaller the proportion of markers representing the skeleton of code complexity in the user's input code, and the smaller the complexity of the corresponding code modification. Therefore, according to the difference between the number of code complexity skeletons and the second total quantity And And the difference between the code complexity And The complexity change index of the corresponding marker sequence before and after modification can be determined.
[0071] In one example, the above complexity change index acquisition unit further includes: a skeleton number difference acquisition subunit, configured to: determine an absolute value of a third difference between the number of code complexity skeletons in the token sequences before and after each modification, to obtain a code complexity skeleton number difference; a mapping value acquisition subunit, configured to: determine a positive correlation mapping value of the difference between the code complexity of the token sequence before each modification and the code complexity of the token sequence after the modification; a complexity change index acquisition subunit, configured to: perform normalization processing on the product of the code complexity skeleton number difference and the positive correlation mapping value, to obtain a complexity change index of the token sequences before and after each modification.
[0072] In this embodiment, the complexity change index of the token sequences before and after each modification is determined by the following formula: ; In the formula: represents the complexity change index of the token sequences before and after each modification; represents a normalization function for performing normalization processing; represents an exponential function with the natural constant e as the base, for performing positive correlation mapping processing to obtain a positive correlation mapping value.
[0073] Further, in a possible implementation manner, the above modification feature acquisition module 300 is configured to: determine a ratio of the similarity and the complexity change index before and after each modification as a code modification feature corresponding to each modification.
[0074] Specifically, the more dissimilar the AST structure of the input code after the user's modification is from the original code, and at the same time, the higher the complexity of the input code after the user's modification is compared to the original code, the higher the possibility that the input code after the user's modification contains personal style features, and the larger the value of the corresponding code modification feature. Therefore, the ratio of the similarity and the complexity change index before and after each modification is determined as the code modification feature corresponding to each modification. At this time, there is a code modification feature corresponding to each modification .
[0075] Further, in a possible implementation manner, the above weight acquisition module 400 includes: an accumulation value acquisition unit, configured to: determine an accumulation value of all the code modification features; a learning weight acquisition unit, configured to: determine a ratio of each code modification feature to the accumulation value as a code modification feature learning weight corresponding to each modification.
[0076] Specifically, each time the user modifies the code in the code editor, the code modification feature corresponding to the user's code modification this time is calculated once, and the accumulation values of all the code modification features are accumulated to obtain an accumulation value. The ratio of each code modification feature to the accumulation value is determined, so as to obtain the code modification feature learning weight corresponding to each modification: ; In the formula: represents the code modification feature learning weight corresponding to the th modification when the user modifies the system-generated code in the code editor during the construction of the code editor, that is, the weight value corresponding to the th code (referring to the modified code) constructed by the user in the code editor in subsequent learning of the user's code style; represents the number of times the user modifies the system-generated code in the code editor, that is, the number of times the user constructs the code during this code editor construction process; and respectively represent the code modification features corresponding to the th modification and the th modification.
[0077] Furthermore, in a possible implementation manner, the above model acquisition module 500 includes: a loss function determination unit, configured to: determine the loss function of the code enhancement model during training; a model training unit, configured to: use the code before each modification as the input of the code enhancement model, use the code after each modification as the output of the code enhancement model, and set the weight value of the code before each modification to the product of the code modification feature learning weight and the preset overall weight, and train the code enhancement model to obtain a trained code enhancement model.
[0078] Specifically, construct a code enhancement model. In one example, the code enhancement model is a neural network model. More specifically, the neural network model is a CodeT5 model. Determine the loss function of the code enhancement model during training, and the loss function can be a cross-entropy loss function.
[0079] Construct a training dataset by using all the codes before and after modification and the learning weights of code modification features. Take 80% of the data in the training dataset as the training set and 20% as the test set. Use the training set and the test set to train the code enhancement model. The training type of the model is code generation, so as to obtain a trained code enhancement model. Among them, during the training process, the input of the model is the original code segment, that is, the code before modification. During the input process, set the weight value of the input code before modification to the product of the code modification feature learning weight and the preset overall weight. The value of the preset overall weight can be set to 0.7. The output of the model is the optimized code after structured enhancement processing, that is, the code after modification. By setting the weight value of the input code before modification, the model can better learn the user's own code input style. Since the pre-training process of the neural network model based on the weight value is a prior art, it will not be specifically introduced here. Use the trained code enhancement model to optimize the code automatically generated by the visual operation system for the user each time during the drag-and-drop interface construction process, so as to obtain the optimized code that conforms to the user's personal style. Figure 3 and Figure 4 respectively show the schematic diagrams before and after code optimization.
[0080] Based on the same inventive concept, as Figure 5 shown, the embodiment of the present invention also provides a user code optimization method applied to a low-code visual programming system based on bidirectional AST dynamic parsing. The method includes: Step S10: Obtain the token sequence of the code before and after each modification and the AST structure composed of the token sequences when the user modifies the code generated by the system in the code editor; Step S20: Determine the similarity between the AST structures before and after each modification and the complexity change index of the token sequences before and after each modification; Step S30: Determine the code modification features corresponding to each modification according to the similarity and the complexity change index; Step S40: Determine the code modification feature learning weights corresponding to each modification according to the code modification features; Step S50: Train the constructed code enhancement model based on the codes before and after several modifications and the code modification feature learning weights to obtain a trained code enhancement model; Step S60: Use the trained code enhancement model to optimize the code automatically generated by the visual operation system for the user each time during the drag-and-drop interface construction process.
[0081] Since the steps S10 to S60 in the user code optimization method correspond one by one to the implementation functions of the six modules included in the above user code optimization module, and the implementation functions of the six modules included in the user code optimization module have been introduced in detail above, the user code optimization method will not be elaborated here.
[0082] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention, and should all be included in the protection scope of the present invention.< / uibutton>
Claims
1. A low-code visual programming system based on bidirectional AST dynamic parsing, characterized in that The system includes a user code optimization module, and the user code optimization module includes: A data acquisition module, configured to: acquire the token sequences of the code before and after each modification and the AST structure formed by the token sequences when the user modifies the code generated by the system in a code editor; A before-and-after modification index acquisition module, configured to: determine the similarity between the AST structures before and after each modification and the complexity change index of the token sequences before and after each modification; A modification feature acquisition module, configured to: determine the code modification features corresponding to each modification according to the similarity and the complexity change index; A weight acquisition module, configured to: determine the code modification feature learning weights corresponding to each modification according to the code modification features; A model acquisition module, configured to: train the constructed code enhancement model based on the code before and after several modifications and the code modification feature learning weights, and obtain a trained code enhancement model; A code optimization processing module, configured to: optimize the code automatically generated by the system for the user's each visualization operation during the drag-and-drop interface construction process by using the trained code enhancement model.
2. The low-code visual programming system based on bidirectional AST dynamic parsing according to claim 1, characterized in that The before-and-after modification index acquisition module includes: A similarity acquisition unit, configured to: determine the similarity between the AST structures before and after each modification according to the difference in the number of nodes and the difference in the number of layers between the AST structures before and after each modification; A complexity change index acquisition unit, configured to: determine the number of code complexity skeletons and the code complexity of the token sequences, and determine the complexity change index of the token sequences before and after each modification according to the difference in the number of code complexity skeletons and the difference in code complexity between the token sequences before and after each modification.
3. The low-code visual programming system based on bidirectional AST dynamic parsing according to claim 2, characterized in that, The similarity acquisition unit includes: A first processing value acquisition subunit, configured to: determine the absolute value of the first difference in the number of nodes between the AST structures before and after each modification, and perform a negative correlation mapping process on the absolute value of the first difference to obtain a first negatively correlated mapping processing value; A second processing value acquisition subunit, configured to: determine the absolute value of the second difference in the number of layers between the AST structures before and after each modification, and perform a negative correlation mapping process on the absolute value of the second difference to obtain a second negatively correlated mapping processing value; A similarity acquisition subunit, configured to: determine the product of the first negatively correlated mapping processing value and the second negatively correlated mapping processing value as the similarity between the AST structures before and after each modification.
4. The low-code visual programming system based on bidirectional AST dynamic parsing according to claim 2, wherein The complexity change index acquisition unit includes: A code complexity skeleton number acquisition subunit, configured to: determine the first total number of keywords and delimiters included in the token sequences as the number of code complexity skeletons of the token sequences; A code complexity acquisition subunit, configured to: match the values in the value fields corresponding to the keywords and delimiters included in the token sequences with the values in the value fields corresponding to the keywords in the preset keyword set and the delimiters in the delimiter set, and determine the second total number of keywords and delimiters with matches; determine the ratio of the second total number to the number of code complexity skeletons as the code complexity of the token sequences.
5. The low-code visual programming system based on bidirectional AST dynamic parsing according to claim 4, wherein The complexity change index acquisition unit further includes: A skeleton quantity difference acquisition subunit, configured to: determine the absolute value of the third difference of the code complexity skeleton quantity between the token sequences before and after each modification, so as to obtain the code complexity skeleton quantity difference; A mapping value acquisition subunit, configured to: determine the positive correlation mapping value of the difference between the code complexity of the token sequence before each modification and the code complexity of the token sequence after the modification; A complexity change index acquisition subunit, configured to: perform normalization processing on the product of the code complexity skeleton quantity difference and the positive correlation mapping value, so as to obtain the complexity change index of the token sequences before and after each modification.
6. The low-code visual programming system based on bidirectional AST dynamic parsing according to claim 1, characterized in that, The modification feature acquisition module is configured to: determine the ratio of the similarity and the complexity change index before and after each modification as the code modification feature corresponding to each modification.
7. The low-code visual programming system based on bidirectional AST dynamic parsing according to claim 1, wherein The weight acquisition module includes: An accumulation value acquisition unit, configured to: determine the accumulation value of all the code modification features; A learning weight acquisition unit, configured to: determine the ratio of each code modification feature to the accumulation value as the code modification feature learning weight corresponding to each modification.
8. The low-code visual programming system based on bidirectional AST dynamic parsing according to claim 1, characterized in that The model acquisition module includes: A loss function determination unit, configured to: determine the loss function of the code enhancement model during the training process; A model training unit, configured to: use the code before each modification as the input of the code enhancement model, use the code after each modification as the output of the code enhancement model, and set the weight value of the code before each modification to the product of the code modification feature learning weight and the preset overall weight, and train the code enhancement model to obtain a trained code enhancement model.
9. The low-code visual programming system based on bidirectional AST dynamic parsing according to claim 1, characterized in that The code enhancement model is a neural network model.
10. The low-code visual programming system based on bidirectional AST dynamic parsing according to claim 9, characterized in that, The neural network model is a CodeT5 model.
Citation Information
Patent Citations
Meta-learning-based code self-adaptive generation method
CN112114791A
Code file difference analysis method and device based on abstract syntax tree iterative mapping
CN116560662A
Code generation method and system, electronic equipment and storage medium
CN119621543A
An intelligent low-code development method and system based on source code
CN119739374A
Source code changed portion extracting method, source code changed portion extracting program, and source code changed portion extracting device
JP2022114223A