An intelligent code completion method in a programming education scenario

Through the method of expanding the data set and optimizing the recommendation list, the problem of insufficient training data sets and lack of scenario adaptation in the field of programming education is solved, and more efficient code completion suggestions are achieved, improving students' programming learning experience.

CN119556901BActive Publication Date: 2025-05-30ARTIFICIAL INTELLIGENCE RES INST OF HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ARTIFICIAL INTELLIGENCE LAB)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510117050.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-30
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

When the existing intelligent code completion method is applied in the field of programming education, the training data set is insufficient and the scenario adaptation is lacking, resulting in the generated code suggestions that do not match students' learning needs and cannot effectively guide students to complete targeted training tasks.

Method used

By presetting the universal large language model of programming learner identity, we simulate the answer codes of students of different styles, expand the data set, and optimize the prediction accuracy of the large language model through cross-entropy loss to generate a complete model. Combining students' answer codes and scene information, optimize the recommendation list and provide code completion suggestions that are more suitable for students' learning scenarios.

Benefits of technology

It effectively solves the problem of insufficient general data sets, improves the quality of code completion recommendations in programming education scenarios, makes the generated code suggestions more in line with students' learning needs, and enhances students' programming learning experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119556901B_ABST
    Figure CN119556901B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of code completion, and discloses an intelligent code completion method in a programming education scenario, including: simulating the answering codes of students with different styles through a general large language model preset with the identity of programming learners; encoding each symbol of the answering codes in the augmented dataset, inputting the obtained encoded sequence into the large language model, and the large language model predicts the next symbol of the encoded sequence based on the context of the encoded sequence, and optimizing the prediction accuracy of the large language model by minimizing the cross-entropy loss to obtain a completion model; inputting the student's answering code into the trained completion model to obtain an initial recommendation list, selecting the three symbols most relevant to the scenario information in the initial list, and promoting the positions of the three symbols in the initial recommendation list to generate a final recommendation list. The present invention optimizes the recommendation list with the help of scenario information, which can effectively improve the recommendation quality in the programming education scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of code completion, and particularly relates to an intelligent code completion method in the context of programming education. Background Art

[0002] Code completion, as one of the core functions in an integrated development environment, is crucial for students' programming learning. First of all, code completion can provide real-time suggestions for common syntax, functions, and variable names, which can help students quickly familiarize themselves with the rules and standard writing methods of programming languages, and lower the learning threshold. Secondly, it can effectively reduce students' spelling and grammar mistakes, help students focus on logical design and problem-solving, and avoid the frustration caused by basic mistakes. In addition, code completion also provides students with opportunities for exploration. By viewing the auto-completion content, students can gradually understand more functions and features of programming languages. Finally, code completion can directly enhance students' programming experience, making them feel the fluency and creativity of programming, thereby enhancing their learning interest and sense of achievement.

[0003] Compared with the traditional code completion method based on rules, the intelligent code completion method can understand the code's connotation and provide more accurate, efficient, and intelligent completion suggestions. Therefore, the intelligent code completion method is the mainstream method for implementing code completion currently. Relying on deep learning models, through learning a large amount of code data, it understands the context and semantics of the code, and then generates real-time code suggestions for users.

[0004] However, directly applying the current mainstream intelligent code completion to the field of programming education has the following problems:

[0005] (1) Insufficient training data set. The training data of the current intelligent code completion model mainly comes from professional codes on GitHub. Although it has guiding significance for efficient programming, it has deficiencies when used in programming education. First of all, the model often assumes that users have a certain programming background, and the generated suggestions may be too complex for beginners to understand and learn. Secondly, the model tends to optimize code efficiency and professional specifications rather than teaching friendliness, and may ignore the explanation of basic knowledge and the guidance of common problems for novices. Finally, the professional code style may not be completely consistent with the educational goals, which is likely to let students fall into the imitation misunderstanding and affect the cultivation of programming thinking and logical ability.

[0006] (2) Lack of scenario adaptation. Compared with programmers who focus on solving actual engineering problems and improving development efficiency, in the field of programming education, students need to focus on specific topics and knowledge points for programming training to cultivate basic programming skills and logical thinking abilities. However, the current intelligent code completion methods usually aim to improve generality and code generation efficiency, lacking in-depth adaptation to students' learning scenarios, and thus it is difficult to effectively provide completion suggestions.

[0007] In the technical solutions of the prior art, when facing the special scenario of programming education, a general intelligent code completion method is often used, resulting in the generated code suggestions not matching the learning needs of students and being unable to effectively guide students to complete targeted training tasks. How to effectively make up for the deficiencies of the general dataset and effectively combine the student programming practice scenario to provide an intelligent code completion method suitable for the field of programming education is still a challenging task. Summary of the Invention

[0008] To solve the above technical problems, the present invention provides an intelligent code completion method in the programming education scenario.

[0009] To solve the above technical problems, the present invention adopts the following technical solutions:

[0010] An intelligent code completion method in the programming education scenario, comprising:

[0011] By using a general large language model preset with the identity of programming learners to simulate the answer codes of students with different styles, an extended dataset is obtained;

[0012] Encode each symbol of the answer code in the extended dataset, input the obtained encoded sequence into the large language model, the large language model predicts the next symbol of the encoded sequence based on the context of the encoded sequence, and evaluates the difference between the predicted symbol probability distribution and the target symbol probability distribution by calculating the cross-entropy loss; optimize the prediction accuracy of the large language model by minimizing the cross-entropy loss to obtain a completion model;

[0013] Input the student's answer code into the trained completion model to obtain an initial recommendation list, combine the current knowledge point and the question information to form scenario information, select the symbol in the initial recommendation list that is most relevant to the scenario information, and promote the ranking of the selected symbols in the initial recommendation list to generate a final recommendation list.

[0014] Further, the symbol is the smallest unit for compiling and running the answer code.

[0015] Further, the step of using a general large language model preset with the identity of programming learners to simulate the answer codes of students with different styles to obtain an extended dataset specifically includes:

[0016] Provide multiple knowledge points for a certain programming language for students to learn, provide multiple questions for each knowledge point for students to practice programming; during the programming practice process, record the student's answer code and answer result, and set a standard answer for each question;

[0017] For each question under each knowledge point of the programming language, select a piece of answer code with a passing result, together with the corresponding knowledge point, question, and standard answer, and input them into a general large language model. Set the identity of the programming learner for the large language model to let the large language model simulate the answer codes of students with different styles; randomly insert the answer codes output by the large language model into the general dataset to obtain an augmented dataset.

[0018] Further, the large language model predicts the next symbol of the coding sequence based on the previous context of the coding sequence, and evaluates the difference between the predicted symbol probability distribution and the target symbol probability distribution by calculating the cross-entropy loss; optimize the prediction accuracy of the large language model by minimizing the cross-entropy loss to obtain a completion model, specifically including:

[0019] The large language model passes through the first t symbols of the coding sequence to predict the probability distribution of the (t + 1)-th symbol ; use the cross-entropy loss to calculate the error between the symbol probability distribution predicted by the large language model and the target symbol probability distribution. By minimizing the cross-entropy loss, the symbol probability distribution predicted by the large language model is made close to the target symbol probability distribution.

[0020] Further, the cross-entropy loss is:

[0021] ;

[0022] is the t-th symbol in the target symbol sequence; represents the conditional probability;

[0023] When the cross-entropy loss converges, that is, the training of the large language model is completed, and a completion model is obtained.

[0024] Further, input the student's answer code into the trained completion model to obtain an initial recommendation list, and combine the current knowledge point and question information to form scenario information, specifically including:

[0025] When a student practices question question under knowledge point knowledge, the student's answer code will be input into the trained completion model in real time, calculate the probability of each symbol in the vocabulary appearing in the next position of the answer code, and retain the first symbols with the highest probability to form an initial recommendation list ; is the K-th symbol in the initial recommendation list;

[0026] Concatenate the knowledge point and the question to obtain the scenario information scene:

[0027] ;

[0028] represents a splicing operation.

[0029] Further, select the symbols in the initial recommendation list that are most relevant to the scenario information, and promote the selected symbols in the initial recommendation list to generate a final recommendation list, specifically including:

[0030] Use the BERT encoder to encode the scenario information into a scenario information vector ; Use the BERT encoder to encode each symbol in the initial recommendation list into a symbol vector to obtain a recommendation list vector , is the total number of symbols in the initial recommendation list; calculate the similarity between the i-th symbol vector in the recommendation list vector and the scenario information vector , ;

[0031] Promote the positions of the symbols corresponding to the highest similarities by at least one position to obtain the final recommendation list.

[0032] Further, the calculation of the similarity between the i-th symbol vector in the recommendation list vector and the scenario information vector , specifically includes:

[0033] ;

[0034] where is an operation to find the vector length.

[0035] Further, the promotion of the positions of the symbols corresponding to the highest similarities by at least one position specifically includes:

[0036] If the position of a certain symbol is the first, the position of this symbol will not be promoted.

[0037] Compared with the prior art, the beneficial technical effects of the present invention are:

[0038] With the help of the programming education platform data and the general large language model, the present invention expands the general code completion dataset, effectively solving the problem that the original dataset is difficult to adapt to the programming learning needs of students. In addition, the present invention optimizes the recommendation list with the help of scenario information, which can effectively improve the recommendation quality in the programming education scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a flowchart of the method in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] A preferred embodiment of the present invention will be described in detail below with reference to the accompanying drawings.

[0041] This embodiment provides an intelligent code completion method in a programming education scenario. First, with the help of the programming education platform and the general large language model, the general dataset is expanded. Secondly, the expanded dataset is used to train the large language model to obtain a completion model. Finally, in combination with the student programming scenario, the recommendation list of the code completion model is optimized. In a preferred embodiment, the large language model can be the GPT-2 model.

[0042] As Figure 1 shown, the intelligent code completion method in the programming education scenario in this embodiment includes the following steps:

[0043] S1, Dataset construction: With the help of the programming education platform data and the general large language model, answer code data is generated to expand the general dataset.

[0044] S2, Model training: Use the expanded dataset to train the GPT-2 model to obtain a completion model.

[0045] S3, Post-processing of the recommendation list: Input the existing answer code of the student into the trained completion model to obtain an initial recommendation list. Combine the programming knowledge points and questions of the student to perform post-processing on the recommendation list to obtain a final recommendation list and feedback it to the student.

[0046] Furthermore, step S1 specifically includes:

[0047] Taking the Python programming language as an example, the programming education platform will provide multiple knowledge points for Python for students to learn. After selecting a certain knowledge point, the programming education platform will provide rich questions for students to practice programming. During this process, the programming education platform will record the answer code and answer results of the students. In addition, the programming education platform will set a standard answer for each question. In a preferred embodiment, there are two types of answer results, namely passed and failed.

[0048] For all the questions under all the knowledge points of the Python programming language, select a piece of answer code that passes the answer. Send the knowledge points of this question, the question stem of the question, the standard answer, and a piece of answer code that passes the answer to a conversational general large language model, and let the general large language model simulate the answer codes of students with different styles. In a preferred embodiment, the general large language model can be selected from ChatGPT, Spark Model, and ERNIE Bot. During this process, in order to improve the accuracy of the answers of the general large language model, set the identity of the general large language model as a programming learner. The prompt prompt constructed based on the knowledge points, questions, standard answers, and answer codes is:

[0049] ;

[0050] Among them, smooth means using smooth statements to splice relevant information, identity means setting the identity for the general large language model, K represents the knowledge point, Q represents the question, A represents the standard answer, and S represents the answer code that passes the answer. In a preferred embodiment, the prompt prompt is as follows:

[0051] prompt = "You are a student who is learning Python. To learn the knowledge point of [" + K + "], you are doing programming exercises for the question of [" + Q + "]. The standard answer to this question is [" + A + "], and the answer code that passes the test is [" + S + "]. Now, please synthesize the above information and simulate three different styles of student answer codes for me. These three styles are the beginner style, the experimenter style, and the minimalist style. Please note that the three simulated student answer codes are only different in style, but they are all correct codes that can achieve the requirements of the question. Please return the experimental results in the form of a set, without additional descriptions other than the set."

[0052] In a preferred embodiment, the general dataset for Python code completion training is the Py150 dataset, which consists of 150k segments of Python code and can be directly obtained. For each programming question, three simulated student answer codes can be obtained with the help of a conversational general large language model. Insert these codes randomly into the general dataset to obtain the augmented dataset DataSet.

[0053] Further, step S2 specifically includes:

[0054] Use the BERT encoder to encode each symbol (token) of the answer code in the augmented dataset DataSet, and use the encoded sequence for the training of the GPT-2 model. A symbol (token) is the smallest unit for the source code to be compiled and run.

[0055] Given the encoded sequence , GPT-2 predicts the probability distribution of the (t + 1)-th symbol based on the first t symbols . For each symbol in the input encoded sequence, the GPT-2 model predicts the probability distribution of the symbol at that position. The GPT-2 model uses cross-entropy loss to calculate the error between the predicted symbol probability distribution and the target symbol probability distribution. By minimizing the cross-entropy loss, the predicted symbol probability distribution is made closer to the target symbol probability distribution, thereby improving the prediction accuracy. Suppose the target symbol sequence is , then the cross-entropy loss is :

[0056] ;

[0057] When the training loss gradually approaches a small value and no longer decreases significantly, it indicates that the large language model has fully learned the patterns in the training data. At this point, the training of the large language model can be stopped to obtain the completion model model

[0058] Furthermore, step S3 specifically includes

[0059] When the student practices the questions question under the knowledge point knowledge, the answer code already written by the student is input to the completion model model. The completion model calculates the probability of each symbol in the vocabulary and retains the top ten recommendations with the highest probabilities to form the initial recommendation list , being the 10th symbol in the initial recommendation list

[0060] By concatenating the knowledge point information and the question information, the scenario information scene can be obtained

[0061] .

[0062] Using the BERT encoder, the scenario information is encoded from a string into a vector form to obtain the scenario information vector . For each symbol in the initial recommendation list, using the BERT encoder in the same way, it is encoded into a symbol vector to obtain the recommendation list vector . For the i-th symbol vector (i = 1, 2,..., 10) in the recommendation list vector, calculate its similarity with the scenario information vector :

[0063] ;

[0064] where represents the inner product of these two vectors, and Represents the lengths of two vectors.

[0065] Select the three symbols in the recommendation list that are most similar to the scene information. Then, move the ranking of these three symbols up by one (if a symbol is ranked first, it will not be moved up), and obtain the final recommendation list for students to choose from. For example, the initial recommendation list is ,if is most relevant to the scene information, the final recommendation list is .

[0066] It is obvious to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention, and any reference numerals in the claims should not be regarded as limiting the claims involved.

[0067] In addition, it should be understood that although the present specification is described in terms of implementation modes, not every implementation mode contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. An intelligent code completion method in a programming education scenario, characterized in that: include: Through the universal large language model with the identity of programming learners preset, the answer codes of students with different styles are simulated to obtain an expanded data set; Encode each symbol of the answer code in the expanded data set, input the obtained encoding sequence into the large language model, and the large language model predicts the next symbol of the encoding sequence based on the previous encoding sequence, and evaluates the difference between the probability distribution of the predicted symbol and the probability distribution of the target symbol by calculating the cross entropy loss; optimize the prediction accuracy of the large language model by minimizing the cross entropy loss to obtain the completion model; Input the student's answer code into the completed completion model to obtain the initial recommendation list, and combine the current knowledge point and question information to form scenario information: When the student practices the question under the knowledge point knowledge, the student's answer code will be input into the completed completion model in real time to calculate the probability of each symbol in the vocabulary appearing in the next position of the answer code, and retain the top probability. symbols, forming the initial recommendation list , is the Kth symbol in the initial recommendation list; concatenate the knowledge point with the question to obtain the scene information scene, ; Represents a splicing operation; Select the most relevant scene information from the initial recommendation list symbols and promote the selected The position of the symbols is used to generate the final recommendation list: Use the BERT encoder to encode the scene information into a scene information vector ; Use the BERT encoder to encode each symbol in the initial recommendation list into a symbol vector to obtain the recommendation list vector ;calculate The i-th symbol vector in With scene information vector Similarity , ; The highest The rank of the symbol corresponding to the similarity is increased by at least one place.

2. The intelligent code completion method in the programming education scenario according to claim 1 is characterized in that: The symbol is the smallest unit for compiling and running the answer code.

3. The intelligent code completion method in the programming education scenario according to claim 1 is characterized in that: The general large language model with the identity of programming learners is preset to simulate the answer codes of students with different styles to obtain an expanded data set, which specifically includes: Provide multiple knowledge points for a certain programming language for students to learn, and provide multiple questions for each knowledge point for students to practice programming; during the programming practice, record the students' answer codes and answer results, and set standard answers for each question; For each question under each knowledge point of the programming language, a piece of answer code with a passing answer result is selected, and sent into a general large language model together with the corresponding knowledge point, question, and standard answer, and the identity of the programming learner is set for the large language model to let the large language model simulate the answer codes of students with different styles; the answer code output by the large language model is randomly inserted into the general data set to obtain an expanded data set.

4. The intelligent code completion method in a programming education scenario according to claim 1, characterized in that: The large language model predicts the next symbol of the coding sequence according to the previous context of the coding sequence, and evaluates the difference between the probability distribution of the predicted symbol and the probability distribution of the target symbol by calculating the cross entropy loss; the prediction accuracy of the large language model is optimized by minimizing the cross entropy loss to obtain a completion model, which specifically includes: Large language models by encoding sequences The first t symbols of To predict the t+1th symbol The probability distribution of ; Use cross entropy loss to calculate the error between the symbol probability distribution predicted by the large language model and the target symbol probability distribution. By minimizing the cross entropy loss, the symbol probability distribution predicted by the large language model is close to the target symbol probability distribution.

5. The intelligent code completion method in the programming education scenario according to claim 4 is characterized in that: The cross entropy loss for: ; is the tth symbol in the target symbol sequence, represents conditional probability; When the cross entropy loss converges, the training of the large language model is completed and the completion model is obtained.

6. The intelligent code completion method in the programming education scenario according to claim 1 is characterized in that: The calculation The i-th symbol vector in With scene information vector Similarity , specifically including: ; in, An operation to find the length of a vector.

7. The intelligent code completion method in a programming education scenario according to claim 1, characterized in that: The highest The rank of the symbol corresponding to the similarity is increased by at least one, specifically including: If the rank of a symbol is the first, the rank of the symbol will not be increased.

Citation Information

Patent Citations

  • Programming information recommendation method and device based on similar code recognition

    CN112115362A

  • Enhanced Transform-based code completion method under lexical element granularity

    CN116661797A

  • Self-guiding code error correction data set generation method for teenager programming scene

    CN117971704A

  • Code generation model training method, code generation method and device

    CN118227107A