Problem adaptive code template generation technology for online programming platform
By collecting AC codes on the online programming platform and generating high-quality code templates using genetic evolution algorithms, the problem of insufficient universality of code frameworks in existing technologies is solved, and efficient code generation and learning assistance are achieved.
Patent Information
- Application Number
- CN202510771504.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-26
AI Technical Summary
The code framework of existing online programming platforms cannot address the structural differences of different problems, resulting in users having to manually complete a large amount of basic code. The templates generated by general large models have structural instability and misunderstandings, making it difficult to meet actual practical needs.
By collecting AC codes from online programming platforms, a genetic evolution algorithm is used to generate code templates with clear structure and reasonable semantics. The linear gene sequence representation method of code features and semantic abstraction technology are combined to automatically generate high-quality code templates.
Significantly lower the user's entry threshold, reduce the burden of repetitive coding, improve programming efficiency and learning effects, and provide a universal code framework with guiding value.
Smart Images

Figure CN120704659A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer program generation technology, in particular to a problem-adaptive code template generation method for an online programming platform, belonging to the technical field of programming education and intelligent code generation. Background Art
[0002] Solving programming problems is widely considered an important means to improve programming skills. Currently, online programming platforms such as LeetCode, Codeforces and Niuke.com provide a large number of practice tasks covering data structures and algorithms to help users gradually build up their algorithm application capabilities. However, in the process of solving programming problems, users often need to repeatedly write a large amount of templated code, such as variable definitions, input and output processing, and the construction of basic algorithm frameworks (such as breadth-first search, depth-first search, etc.). Although these contents have a certain teaching role in the initial learning stage, for developers who have already mastered the relevant algorithms and data structures, their learning focus should gradually shift to the application of algorithms in complex scenarios, rather than the repetitive writing of basic logic.
[0003] Although some platforms (such as LeetCode) provide initial code frameworks, these frameworks are usually too general and cannot cover the structural differences of different problems. Users still need to manually complete a large amount of basic code. In addition, although general large models (such as ChatGPT) have certain capabilities in code generation tasks, in the absence of specialized training data, the templates they generate often have problems such as structural instability and comprehension bias, which are difficult to meet actual practical needs. Therefore, there is an urgent need for a code template generation technology with problem-adaptive capabilities that can automatically generate code templates with reasonable structure, complete content, and guiding value for specific programming problems, thereby reducing the burden of repetitive coding for users and improving learning efficiency and problem-solving focus. Summary of the Invention
[0004] The purpose of the present invention is to provide a problem-adaptive code template generation method for an online programming platform, which can automatically generate code templates with clear structure, strong versatility, and reasonable semantics for different problems, so as to reduce the burden of users writing repetitive code and improve programming efficiency and learning effects. Therefore, a problem-adaptive code template generation technology for an online programming platform is proposed. The technology aims to systematically collect and analyze AC codes in the platform, explore the structural commonalities between problems and solutions, and automatically generate high-quality code templates with the help of genetic evolution algorithms, thereby providing a universal framework code skeleton for different problem types, significantly reducing the user's entry threshold and the burden of template construction. The technology focuses on the versatility, structural rationality, readability, non-redundancy of the templates, and the appropriate hiding of the core problem-solving parts, so as to achieve multiple goals such as assisting programming, guiding thinking, and improving learning efficiency.
[0005] The present invention provides a method for generating problem-adaptive code templates for an online programming platform, comprising the following steps:
[0006] Step 1: Collect programming problems from online programming platforms at home and abroad. For each programming problem, collect the AC (Accepted) code shared by developers on the platform, which passes all test cases set for this problem on the platform, and build a <programming problem, AC code> dataset;
[0007] Step 2: Preprocess all collected AC codes and classify them based on their structural similarity. Ensure that the structural similarity between any two AC codes in each category is greater than s. Filter out low-frequency categories and construct the <programming problem / category, AC code> dataset.
[0008] Step 3: Encode each AC code segment using a linear gene sequence representation method that combines code features. For each category, add the encoded AC code to the code statement library in sequence to construct <programming problem / category, code statement library>;
[0009] Step 4: For each category of programming problems, a genetic algorithm is used based on the constructed code base to generate high-quality code templates that meet the requirements, and evolve to generate <programming problem / category, optimal individual>;
[0010] Step 5: After obtaining the optimal individual evolved by the genetic algorithm, it needs to be further decoded. Through semantic abstraction and core code hiding, the final code template is obtained to obtain <programming problem / category, template>.
[0011] In addition, the defect localization technology provided by the present invention may also have the following additional technical features:
[0012] Furthermore, the online programming platform described in step 1 includes but is not limited to the LeetCode platform and the Niuke.com platform, and the language for crawling the AC code is Java.
[0013] Furthermore, in step 2, preprocessing includes removing comments and blank lines and unifying the code style; structural similarity is calculated using the control statement similarity indicator; filtering low-frequency categories specifically involves quantitative filtering of each category, eliminating categories with less than 5 AC code segments, and ensuring that the data volume of each category is sufficient to support the evolution and evaluation of the algorithm.
[0014] Furthermore, in step 3, the linear gene sequence representation method combined with code features is specifically as follows: taking code lines as units, extracting text features, abstract semantic features, and statement types of each line of code, and establishing hash mappings for different features. The hash sequences of these features are used to jointly represent a line of code. The specific steps include:
[0015] Source code text hash value generation: For each line of code, encode it using the character ASCII code;
[0016] Source code abstract expression extraction: By analyzing the AST node attributes of the code, nodes of the same type are unified into consistent abstract expressions or identifiers. This allows two code segments with the same abstract semantic features but possibly different source code text to be converted into the same abstract expression.
[0017] Abstract expression hash value generation: For each line of code, encode its corresponding abstract expression using character ASCII code;
[0018] Source code statement type extraction: determine the statement type based on the AST node type of each line of code;
[0019] Hash map storage and management: Use multiple mapping relationships to effectively manage code information, semantic information, and statement types;
[0020] Code library construction: For each AC code under each category, code it line by line and add it to the code statement library, and build a code statement library for each category.
[0021] Furthermore, in step 4, a genetic algorithm is used based on the constructed code base to generate high-quality code templates that meet the requirements. The key steps include:
[0022] First, template individuals are constructed by randomly selecting encoded code statements from a code statement library to complete the generation of an initial population. Then, through the evolutionary process of a genetic algorithm, the population is selected and mutated using selection and crossover operators that introduce a deduplication mechanism, random crossover, and hierarchical selection strategies. A multi-objective fitness function based on heuristic rules is designed to guide the generation and evolution of code templates.
[0023] Individual fitness represents the degree to which the corresponding template meets the requirements. The fitness function is a multi-objective fitness function based on heuristic rules to guide the generation direction of code templates. Key objectives and evaluation indicators include:
[0024] The template should be composed of frequently occurring code blocks. The indicator design is completed through the frequent code block mining algorithm. For each generated template individual, the proportion of frequent code blocks it contains is calculated;
[0025] The template should maintain common procedural characteristics and capture procedural features common to AC codes, which is measured by comparing the degree of match between the procedural features of each generated individual and all AC codes under that category;
[0026] The template should maintain the common control logic characteristics and capture the common control statement features of AC codes, which can be measured by comparing the degree of control statement feature matching between each generated individual and all AC codes under the category;
[0027] The order of statements within the code blocks in the template should be correct. By calculating the average edit distance between a single code segment and all AC codes, the correct order of basic statements in the appropriate code blocks can be evaluated. A better template with the correct order and correct position will show a smaller average edit distance.
[0028] Templates should use consistent identifiers. Since individual code is constructed on the collection of all AC code through evolutionary operations such as mutation and crossover, it may happen that identifiers such as variable or method names (which should be the same) are used differently in related code statements. Code coverage metrics can help solve this problem, because a template that uses consistent naming will show higher code coverage by matching more lines related to AC code.
[0029] Templates should avoid erroneous repetitions. The same code statements may appear multiple times in the template, but not in the original AC. Edit distance acts as a constraint because the edit distance between individuals of redundant code and the original AC code is usually large. In addition, the code repetition rate evaluates the degree of code repetition by comparing the frequency of identical statements in the template and AC code. The higher the frequency of repeated statements, the lower the code repetition rate. By calculating the average repetition rate of an individual and all AC codes, the redundancy in the generated code can be quantified. These two indicators work together to minimize the occurrence of errors.
[0030] Templates should maintain a well-defined code structure, ensuring that statements are organized in a well-defined code structure with proper indentation and bracket usage, and using a stack-based algorithm to evaluate whether a single element has a well-defined code structure.
[0031] The above seven aspects generally determine the quality of an individual. They are weighted and summed according to different weights, and the final score is used as the fitness of the individual in the evolutionary process.
[0032] Furthermore, in step 5, the final code template is constructed through semantic abstraction and core code hiding. The key steps include:
[0033] First, the hash sequence corresponding to the optimal solution is converted into corresponding source code statements and source code abstract expressions based on the hash mapping relationship stored in the encoding stage; then, the core and non-core parts of each line of code are judged according to their type, the core parts are hidden, and the non-core parts are converted into components of the template through semantic abstraction, thereby generating the final code template. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flow chart of the present invention.
[0035] Figure 2 It is the design goal of the fitness function in the genetic algorithm. DETAILED DESCRIPTION
[0036] To make the objects, features, and advantages of the present invention more readily apparent, the following detailed description of specific embodiments of the present invention is provided in conjunction with the accompanying drawings. The accompanying drawings illustrate several embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the present invention.
[0037] The present invention provides a method for generating problem-adaptive code templates for an online programming platform. The method includes the following steps: S01 collects programming problems on domestic and foreign online programming platforms. For each programming problem, collects the AC (Accepted) codes shared by developers on the platform to construct a <programming problem, AC code> dataset; S02 preprocesses all collected AC codes and classifies them based on the structural similarity of the AC codes to ensure that the structural similarity of any two AC codes in each category is greater than s, and then filters out low-frequency categories to construct a <programming problem / category, AC code> dataset; S03 uses a linear gene sequence representation method combined with code features to encode each AC code. For each category, the encoded AC code is added to the code statement library in turn to construct a <programming problem / category, code statement library>; S04 For each category in the programming problem, a genetic algorithm is used based on the constructed code library to generate a high-quality code template that meets the requirements, and evolve to generate a <programming problem / category, optimal individual>; S05 After obtaining the optimal individual evolved by the genetic algorithm, it needs to be further decoded, and the final code template is obtained through semantic abstraction and core code hiding, thereby obtaining the <programming problem / category, template>.
[0038] In step S01, programming problems on domestic and foreign online programming platforms are collected. For each programming problem, the AC (Accepted) codes shared by developers on the platform are collected to construct a <programming problem, AC code> data set. The specific implementation of the present invention adopts collection from LeetCode (international platform) and Niuke.com (domestic platform). Both LeetCode and Niuke.com support programming exercises in multiple programming languages. The present invention mainly focuses on the template generation of Java programming code. All AC codes (written in Java language) of 1081 and 368 questions were collected from LeetCode and Niuke.com respectively using Python crawler. These programming questions cover a variety of different algorithms and data structures. Among them, the 1081 programming questions on LeetCode cover 58 types including depth-first search, greedy, hash table, etc.; the 368 programming questions on the Niuke platform cover 34 types including binary search, dynamic programming, tree, stack, etc. The rich and diverse programming problems of data structures and algorithms provide good data support for the practicality of the present invention.
[0039] In step S02, all collected AC codes are preprocessed and classified based on their structural similarity, ensuring that the structural similarity between any two AC codes in each category is greater than s. Low-frequency categories are then filtered out. Constructing the <programming problem / category, AC code> dataset includes the following steps:
[0040] The solution code was preprocessed, including removing comments, blank lines, and standardizing the coding style. First, the solution code was converted into an abstract syntax tree (AST) using the Java JDT library. Preprocessing operations were performed based on the AST node type. Specifically, comment removal was achieved by returning an empty string for comment statements with the AST node type LINE_COMMENT or BLOCK_COMMENT. Blank line removal was achieved by strictly controlling the use of line breaks within each formatted line of code, ensuring that the combined code segments contained no extra blank lines. Coding style standardization primarily involved standardizing indentation, code block formatting, and expression formatting. For indentation, a single tab character was used to represent one level of indentation. Within a code block, a corresponding number of tabs were added before each line of code based on the nesting depth. Code block formatting standardization involved determining whether curly braces ended at the end of a statement or on a new line. These were then standardized to include line breaks and indents. (For example, for loops that appeared in a compact format or with line breaks and indents were standardized to include line breaks and indents between the loop condition and the curly braces.) The standardization of expression formats involves the use of uniform spaces. This applies to different types of code statements, such as variable declarations (int ans = 0 and int ans = 0), function calls (func(a, b) and func(a, b)), and calculation expressions (a + b and a + b). For these statement types, the form with spaces is uniformly adopted (for example, int ans = 0 and int ans = 0, the former is used).
[0041] Codes with a structural similarity greater than s are grouped together, meaning that any two code segments within that category have a similarity greater than s. Setting s too high results in too few codes within the same category. Even codes using the same algorithm but slightly different details cannot be grouped together, reducing the practicality of the classification and hindering the versatility of the generated templates. Setting s too low results in codes using different algorithms being grouped together, thus negating the meaning of the classification. Based on actual classification results, s = 0.6 was chosen as a compromise value. Finally, a quantitative filter was performed on each category, eliminating categories with fewer than five AC code segments to ensure that the data volume for each category was sufficient to support algorithm evolution.
[0042] In step S03, each AC code segment is encoded using a linear gene sequence representation method that combines code features. For each category, the encoded AC code is sequentially added to the code statement library to construct <programming problem / category, code statement library>. The specific steps are:
[0043] Source code hash value generation: For a solution code snippet, encode it using a character-based ASCII code method. Specifically, for each line of code in the code snippet, the ASCII code values corresponding to all characters in that line of code are summed up, and the resulting value is the hash value of that line of code.
[0044] Source code abstract expression extraction: For two pieces of code such as int res=5 and int ans=5, due to differences in programming habits, different developers may use different variable names, even if they implement the same function. It is impossible to determine whether the two pieces of code have the same function or logic by simply comparing the text hash values. Therefore, paying attention to the abstract semantic features of the code is crucial for identifying code fragments with the same function but different text forms. To this end, by analyzing the AST node attributes of the code and calling its related API, the nodes of the same type are unified into consistent abstract expressions or identifiers, so that two pieces of code with the same abstract semantic features but possibly different source code texts are converted into the same abstract expression (such as int res=5 and int ans=5 are uniformly represented as int SIMPLE_NAME=Num).
[0045] Abstract expression hash value generation: After obtaining the abstract expression of each code line as described above, each abstract expression is converted into a corresponding abstract semantic hash value using the same hash encoding method used in source code text hash value generation. That is, the sum of the ASCII code values of all characters in an abstract expression is used as the hash value of the expression.
[0046] Source code statement type extraction: The statement type is determined based on the AST node type of each line of code. Specific statement types considered in this invention include loop statements, conditional statements, declaration statements, assignment statements, method call statements, return statements, and expression statements. These statement types can be obtained by viewing the AST node attributes. For example, if the AST node attribute of the code is VARIABLE_DECLARATION_STATEMENT, the statement type is determined to be a variable declaration. If the node attribute is METHOD_INVOCATION, the statement type is a method call.
[0047] Hash map storage and management: Multiple mappings are used to efficiently manage code information, semantic information, and statement types. codeToHash and hashToCode store the mapping between code hash values and original code, while codeToExp and expToHash store the mapping between code semantic expressions and semantic hash values. hashToType stores the statement type corresponding to the code semantic hash value.
[0048] In S04, for each category of programming problems, a genetic algorithm is used based on the constructed code base to generate high-quality code templates that meet the requirements, and evolve to generate <programming problem / category, optimal individual>. The specific steps are:
[0049] Initial population generation: Randomly select several statements from the code corpus and combine them into code templates to complete the construction of the initial population.
[0050] Crossover operation: First, a position is randomly selected in the coding sequence of two parent individuals as the crossover point, and then the chromosomes of the two parent individuals are exchanged at the selected crossover point. During the evolutionary process, it is found that the coding sequences of individuals with high fitness are similar. The present invention divides the parent individuals into two categories based on their fitness values: individuals with high fitness and individuals with low fitness. During each crossover, one parent is selected from the high fitness category and the other parent is selected from the low fitness category as the crossover object. In this way, genetic mixing between individuals with different fitness can be promoted, increasing the genetic diversity of the population.
[0051] Mutation operation: A variety of mutation operations including deletion, replacement, addition, exchange and move are designed, and each operation is selected and executed in a random manner. The specific operations are as follows: (1) Delete operation: Randomly select a gene position in the chromosome coding sequence and delete it. This operation helps to remove redundant or invalid code fragments and simplify the generated template. (2) Replace operation: Randomly select a gene position and replace it with other valid code fragments in the code corpus. This operation can introduce new code structures and increase the diversity of individuals. (3) Add operation: Randomly select a position in the chromosome coding sequence and insert a new code fragment. This operation can introduce new genes to the population and explore a wider solution space. (4) Swap operation: Randomly select two gene positions and exchange their positions. This operation can rearrange the code fragments and adjust the code logical structure. (5) Move operation: Randomly select a gene position and move it to another position in the coding sequence. This operation can change the execution order of the code fragments and optimize the code structure.
[0052] Selection operation: To ensure the diversity of the population, a deduplication mechanism is added to the selection operation, that is, to ensure that the coding sequences of individuals in the population are different. The specific operation steps are as follows: (1) Merge the parent and new generation individuals (2) Calculate the fitness (3) Deduplication operation (4) Select the individuals with the top L fitness values.
[0053] The fitness function design based on multi-objective satisfaction is specifically implemented as follows:
[0054] The template should be composed of frequently occurring code blocks. The indicator design is completed through the frequent code block mining algorithm. For each generated template individual, the proportion of frequent code blocks it contains is calculated;
[0055] The template should maintain common program characteristics and capture the common program features of AC codes, which are measured by comparing the degree of program feature matching between each generated individual and all AC codes under this category. For the representation of program features, this invention is based on the ideas of M.Sudhamani et al., and maps the number of occurrences of template codes such as loop statements, variable declarations, and function declarations into a program feature matrix. The program feature similarity is measured by calculating the feature matrix similarity of two code segments.
[0056] The template should maintain the general control logic characteristics and capture the common control statement features of AC codes. The degree of control statement feature matching between each generated individual and all AC codes under the category is measured. For the representation of control statement features, the type of control statement and the program features in the statement block are filled into the control statement feature matrix.
[0057] The order of statements within the code blocks in the template should be correct. By calculating the average edit distance between a single code segment and all AC codes, the correct arrangement order of basic statements in the appropriate code block can be evaluated. A better template with the correct order and correct position will show a smaller average edit distance. For the calculation of the edit distance, dynamic programming is used to efficiently calculate the distance dis (in code lines) required to convert the template to AC code, which is then normalized: 1-(dis / Math.max(template.lines, AC.lines)).
[0058] Templates should use consistent identifiers. Code coverage metrics can help solve this problem, as a template that uses consistent naming will show higher code coverage by matching more lines related to AC code.
[0059] Templates should avoid erroneous repetitions. The same code statements may appear multiple times in the template, but not in the original AC. Edit distance acts as a constraint because the edit distance between individuals of redundant code and the original AC code is usually large. In addition, the code repetition rate evaluates the degree of code repetition by comparing the frequency of identical statements in the template and AC code. The higher the frequency of repeated statements, the lower the code repetition rate. By calculating the average repetition rate of an individual and all AC codes, the redundancy in the generated code can be quantified. These two indicators work together to minimize the occurrence of errors.
[0060] Templates should maintain a well-defined code structure, ensuring that statements are organized in a well-defined code structure with proper indentation and bracket usage, and using a stack-based algorithm to evaluate whether a single element has a well-defined code structure.
[0061] After S05 obtains the optimal individual evolved by the genetic algorithm, it needs to be further decoded. Through semantic abstraction and core code hiding, the final code template is obtained, and the <programming problem / category, template> is obtained. The specific steps are:
[0062] First, based on the hash mapping relationships (hashToCode, hashToExp) stored during the encoding phase, the hash sequence corresponding to the optimal solution is converted into corresponding source code statements and source code abstract expressions. Subsequently, the core and non-core parts of each line of code are determined based on their type. The core parts are hidden, while the non-core parts are converted into template components through semantic abstraction, thus generating the final code template. Code statement types include loop statements, conditional statements, declaration statements, assignment statements, method call statements, return statements, and expression statements.
[0063] Loop statements. Loop control statements typically contain the logical thinking behind solving a problem. They are treated as core statements. During optimization, loop control statements within loop statements are hidden, and the abstract expression corresponding to that line of code is used to construct the template.
[0064] Conditional statements. Judgment conditions usually contain logical ideas for solving problems. They are also treated as core statements. When optimizing, abstract expressions are used to build templates to hide key information.
[0065] Variable declaration statement. For example, if int idx = 0, the int idx = part is treated as a non-core part and retained in the template during optimization to maintain user familiarity and efficiency. For the initial value assigned to the variable, the corresponding abstract expression is selected for replacement. The abstract expression is constructed using a highly readable identifier based on the AST node type of the initial value. For example: when the AST node type of the initial value is NUMBER_LITERAL, use Num to construct it; when the AST node type of the initial value is STRING_LITERAL, use Str to construct it; when the AST node type of the initial value is PREFIX_EXPRESSION (prefix expression), INFIX_EXPRESSION (infix expression) or POSTFIX_EXPRESSION (postfix expression), use Expression to construct it uniformly.
[0066] Method declaration statements. Method declaration statements usually do not involve relevant logical thinking or operations, so their source code is directly used during template optimization.
[0067] Class declaration statement. Similar to method declaration statement, select the source code construction template.
[0068] Assignment statements. Similar to variable declaration statements, retain the left side and replace the right side with an abstract expression.
[0069] Method call statement. The method name is retained, and each parameter in the parameter list is replaced with "Args." For example, "dfs(2, i+1)" is processed as "dfs(Args)."
[0070] Return statement: Keep return and replace the return value with the corresponding abstract expression. For example, replace return "abc" with return Str.
[0071] Expression statements, including prefix expressions, infix expressions, and postfix expressions, are uniformly replaced with abstract expressions.
[0072] By combining the statement type and converting each line of code into the final template code according to the above rules, the code template is constructed. This method can accurately convert the optimal solution generated by the genetic algorithm into a code template that meets the requirements, thereby achieving the purpose of hiding the core code.
Claims
1. A problem-adaptive code template generation technology for online programming platforms, characterized by: The following steps are involved: Step 1: Collect programming problems from online programming platforms at home and abroad. For each programming problem, collect the AC (Accepted) code shared by developers on the platform, which passes all test cases set for this problem on the platform, and build a <programming problem, AC code> dataset; Step 2: Preprocess all collected AC codes and classify them based on their structural similarity. Ensure that the structural similarity between any two AC codes in each category is greater than s. Filter out low-frequency categories and construct the <programming problem / category, AC code> dataset. Step 3: Encode each AC code segment using a linear gene sequence representation method that combines code features. For each category, add the encoded AC code to the code statement library in sequence to construct <programming problem / category, code statement library>; Step 4: For each category of programming problems, a genetic algorithm is used based on the constructed code base to generate high-quality code templates that meet the requirements, and evolve to generate <programming problem / category, optimal individual>; Step 5: After obtaining the optimal individual evolved by the genetic algorithm, it needs to be further decoded. Through semantic abstraction and core code hiding, the final code template is obtained to obtain <programming problem / category, template>.
2. The code template generation technology according to claim 1, characterized in that: The online programming platforms described in step 1 include but are not limited to the LeetCode platform and the Niuke.com platform, and the language for crawling the AC code is Java.
3. The code template generation technology according to claim 1, characterized in that: In step 2, preprocessing includes removing comments and blank lines and unifying the code style; structural similarity is calculated using the control statement similarity indicator; filtering low-frequency categories specifically involves quantitative filtering of each category, eliminating categories with fewer than 5 AC code segments, and ensuring that the data volume of each category is sufficient to support algorithm evolution and evaluation.
4. The code template generation technology according to claim 1, characterized in that: In step 3, the linear gene sequence representation method combined with code features is specifically as follows: taking code lines as units, extracting the text features, abstract semantic features, and statement types of each line of code, and establishing hash mappings for different features. The hash sequences of these features are used to jointly represent a line of code. The specific steps include: Source code text hash value generation: For each line of code, encode it using the character ASCII code; Source code abstract expression extraction: By analyzing the AST node attributes of the code, nodes of the same type are unified into consistent abstract expressions or identifiers. This allows two code segments with the same abstract semantic features but possibly different source code text to be converted into the same abstract expression. Abstract expression hash value generation: For each line of code, encode its corresponding abstract expression using character ASCII code; Source code statement type extraction: determine the statement type based on the AST node type of each line of code; Hash map storage and management: Use multiple mapping relationships to effectively manage code information, semantic information, and statement types; Code library construction: For each AC code under each category, code it line by line and add it to the code statement library, and build a code statement library for each category.
5. The code template generation technology according to claim 1, characterized in that: In step 4, a genetic algorithm is used to generate high-quality code templates that meet the requirements based on the constructed code base. The key steps include: First, template individuals are constructed by randomly selecting encoded code statements from a code statement library to complete the generation of an initial population. Then, through the evolutionary process of a genetic algorithm, the population is selected and mutated using selection and crossover operators that introduce a deduplication mechanism, random crossover, and hierarchical selection strategies. A multi-objective fitness function based on heuristic rules is designed to guide the generation and evolution of code templates. Individual fitness represents the degree to which the corresponding template meets the requirements. The fitness function is a multi-objective fitness function based on heuristic rules to guide the generation direction of code templates. Key objectives and evaluation indicators include: The template should be composed of frequently occurring code blocks. The indicator design is completed through the frequent code block mining algorithm. For each generated template individual, the proportion of frequent code blocks it contains is calculated; The template should maintain common procedural characteristics and capture procedural features common to AC codes, which is measured by comparing the degree of match between the procedural features of each generated individual and all AC codes under that category; The template should maintain the common control logic characteristics and capture the common control statement features of AC codes, which can be measured by comparing the degree of control statement feature matching between each generated individual and all AC codes under the category; The order of statements within the code blocks in the template should be correct. By calculating the average edit distance between a single code segment and all AC codes, the correct order of basic statements in the appropriate code blocks can be evaluated. A better template with the correct order and correct position will show a smaller average edit distance. Templates should use consistent identifiers. Since individual code is constructed on the collection of all AC code through evolutionary operations such as mutation and crossover, it may happen that identifiers such as variable or method names (which should be the same) are used differently in related code statements. Code coverage metrics can help solve this problem, because a template that uses consistent naming will show higher code coverage by matching more lines related to AC code. Templates should avoid erroneous repetitions. The same code statements may appear multiple times in the template, but not in the original AC. Edit distance acts as a constraint because the edit distance between individuals of redundant code and the original AC code is usually large. In addition, the code repetition rate evaluates the degree of code repetition by comparing the frequency of identical statements in the template and AC code. The higher the frequency of repeated statements, the lower the code repetition rate. By calculating the average repetition rate of an individual and all AC codes, the redundancy in the generated code can be quantified. These two indicators work together to minimize the occurrence of errors. Templates should maintain a well-defined code structure, ensuring that statements are organized in a well-defined code structure with proper indentation and bracket usage, and using a stack-based algorithm to evaluate whether a single element has a well-defined code structure. The above seven aspects generally determine the quality of an individual. They are weighted and summed according to different weights, and the final score is used as the fitness of the individual in the evolutionary process.
6. The code template generation technology according to claim 1, characterized in that: In step 5, the final code template is constructed through semantic abstraction and core code hiding. The key steps include: First, the hash sequence corresponding to the optimal solution is converted into corresponding source code statements and source code abstract expressions based on the hash mapping relationship stored in the encoding stage; then, the core and non-core parts of each line of code are judged according to their type, the core parts are hidden, and the non-core parts are converted into components of the template through semantic abstraction, thereby generating the final code template.