Problem embedding representation method based on code space solution similarity

By constructing the data set of programming problems, knowledge points and submitted codes, using clustering algorithms and solution similarity matrix, and combining knowledge point correlation construct problem embedding representation, the problem of failure to effectively capture solution similarity and achieve level diversity in the existing technology is solved, and more accurate programming knowledge tracking and prediction is achieved.

CN119989006AActive Publication Date: 2025-05-13JIANGXI NORMAL UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510477006.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing programming knowledge tracking model has limitations in characterizing the correlation between problems and fails to effectively capture the similarity of solutions and the diversity at the implementation level.

Method used

By constructing the data set, including programming problems, knowledge points and submitting code, the clustering algorithm is used to extract multiple solutions, and the solution similarity matrix is ​​established, and problem embedding representation is constructed in combination with the knowledge point correlation.

Benefits of technology

It effectively improves the expression ability and prediction accuracy of problem embedding representation, and can more accurately evaluate students' solution design ability and programming knowledge status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989006A_ABST
    Figure CN119989006A_ABST
Patent Text Reader

Abstract

The invention discloses a problem embedding representation method based on code space solution similarity, which comprises the following steps of: constructing a data set, setting a problem text, respectively processing the data set and the problem text to obtain a problem vertex vector, a knowledge point vertex vector and a low-dimensional problem text vector, respectively processing the obtained data, and respectively obtaining a prediction association probability of a problem vertex vector, a joint prediction association probability of the problem vertex vector and a knowledge point vertex vector, and a prediction association probability of the knowledge point vertex vector; a knowledge point set associated with the problem vertex vector is defined and processed, an average knowledge point vector is obtained, the problem vertex vector, the average knowledge point vector and the low-dimensional problem text vector are processed based on the obtained prediction association probability, and a final problem embedding representation vector is obtained. According to the method, the expression ability of the problem embedding expression vector is effectively improved by jointly considering the diversity of the solution and the hierarchical association of the knowledge points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of code space solution, and is a problem embedding representation method based on code space solution similarity. Background Art

[0002] Knowledge Tracing (KT) is a method that models and evaluates the degree of knowledge and concepts mastered by students by analyzing the interaction data between students and exercises. With the introduction of deep learning technology, the performance of KT models has been significantly improved in recent years, and it has been widely used in second language learning, STEM education and other fields. In the field of programming education, relying on the massive interaction data accumulated by online programming platforms such as LeetCode and Niuke.com, the deep learning-driven KT model has also achieved rapid development. In this context, Programming Knowledge Tracing (PKT), which focuses on the knowledge tracing task of programming education, has gradually become a research hotspot.

[0003] Programming knowledge tracking provides an effective way to dynamically evaluate and predict students' programming abilities by analyzing students' behavioral data in programming tasks. Unlike traditional knowledge tracking models, PKT focuses on the mastery assessment of programming knowledge points. Its main goals include estimating students' mastery of programming knowledge points and predicting whether their code submissions in subsequent programming problems are correct. To achieve this goal, effective embedding representation of problems is crucial. High-quality problem representations can accurately capture students' knowledge status, thereby improving the accuracy of predictions. Previous studies have shown that properly designed problem embeddings can significantly improve the performance of the PKT model and become a key step in inferring students' programming knowledge level.

[0004] The current PKT model problem embedding methods can be mainly summarized into the following four categories: generating problem embedding based on problem index and student answers, generating problem embedding based on the relationship between problems and knowledge points, generating problem embedding based on problem text content, and generating problem embedding using code information. However, these methods still have limitations in characterizing the correlation between problems, which are mainly reflected in the following two aspects: First, existing research often only focuses on the correlation of knowledge points, ignoring the similarity of solutions contained in the code submitted by students. In fact, there may be multiple solutions to the same problem, and different problems may also use similar solutions; constructing problem representation based on solutions can not only highlight the differences in ideas and implementations of different solutions, but also capture the commonality of solutions in related problems, thereby more accurately evaluating students' solution design ability and programming knowledge status. Secondly, even if two problems involve the same programming knowledge points (such as recursive algorithms), their code implementation methods may be very different-for example, one problem focuses on recursive depth optimization, while the other emphasizes complex data structure processing. It is difficult to fully reflect the diversity of these implementation levels by relying solely on knowledge point associations. Therefore, in order to more accurately characterize the essential characteristics of the problem, it is necessary to comprehensively consider the problem correlation between the "code space" and the "knowledge point level" to provide more comprehensive modeling information for the problem embedding representation. Summary of the invention

[0005] In view of the deficiencies in the prior art, the present invention provides a problem embedding representation method based on code space solution similarity to solve the problems raised in the above background technology.

[0006] To achieve the above object, the present invention provides the following technical solution: a problem embedding representation method based on code space solution similarity, comprising the following steps: Step S1: construct a data set, which includes programming problems, knowledge points, and all submitted codes corresponding to the programming problems; Step S2: Process the programming problem and the knowledge point to obtain the problem vertex vector and the knowledge point vertex vector, set a problem text, and process the problem text to obtain the problem text vector; Step S3: by processing the problem vertex vector, the predicted association probability of the problem vertex vector is obtained; Step S4: by processing the problem vertex vector and the knowledge point vertex vector, a joint prediction association probability of the problem vertex vector and the knowledge point vertex vector is obtained; Step S5: obtaining the predicted association probability of the knowledge point vertex vector by processing the knowledge point vertex vector; Step S6: Use linear neural network mapping to obtain a low-dimensional question text vector for the question text vector; define a set of knowledge points associated with the question vertex vector, process the knowledge point set by arithmetic averaging to obtain an average knowledge point vector, concatenate the question vertex vector, the average knowledge point vector and the low-dimensional question text vector based on the predicted association probability of the question vertex vector, the joint predicted association probability of the question vertex vector and the knowledge point vertex vector, and the predicted association probability of the knowledge point vertex vector, construct a first-order feature matrix, process the first-order feature matrix by defining an interaction function, construct a second-order feature matrix, process the first-order feature matrix and the second-order feature matrix, obtain the scalar of the first-order feature matrix and the scalar of the second-order feature matrix, add the scalar of the first-order feature matrix and the scalar of the second-order feature matrix, and obtain the final question embedding representation vector.

[0007] Further, suppose For programming problems A collection of Indicates A programming question, is the total number of programming problems; let For knowledge points A collection of; Indicates Knowledge points; Indicates Knowledge points; is the total number of knowledge points; Submit code for all programming questions A collection of Indicates that the corresponding Programming Questions All submitted code.

[0008] Further, the problem vertex vector includes a first problem vertex vector , the second problem vertex vector , the knowledge point vertex vector includes the first knowledge point vertex vector and the second knowledge point vertex vector ; Get the question text vector The specific process is as follows: define a question text, the large-scale pre-trained language model first decomposes the question text into a word sequence, and then processes the word sequence through the word embedding layer to obtain the word embedding sequence , For the word embedding sequence; finally, the question text vector is obtained by calculating the average value of the word embedding sequence .

[0009] Further, let the corresponding Programming Questions All submitted codes represents a set of programming problems, Indicates Programming Questions No. code vectors submitted, Indicates Programming Questions The number of code vector submissions corresponding to Programming Questions All submitted codes Use the selected clustering algorithm Processing is performed to generate multiple cluster center vectors, and multiple cluster center vectors include the first Programming Questions and Programming Questions , multiple cluster center vectors are regarded as Programming Questions and Programming Questions The solution vector set and , expressed as: (1); (2); In the formula, Represents the number of cluster center vectors generated by the clustering algorithm; Indicates Programming Questions No. The center vector of each cluster; Indicates Programming Questions No. The center vector of each cluster; Indicates Programming Questions The set of cluster center vectors; Indicates Programming Questions The set of cluster center vectors; Solution vector set and Calculate the cosine value of the normal vector and take the maximum value as the first Programming Questions and Programming Questions The similarity at the solution level is expressed as: (3); In the formula, Indicates Programming Questions and Programming Questions similarities between; and Represents the solution vector set and Any solution center vector in ; Represents a transpose operation; represents the vector norm; Indicates the maximum value of similarity; Based on Programming Questions and Programming Questions Similarity at the solution level, constructing a solution similarity matrix ,in ; Indicates Programming Questions and Programming Questions The similarity of solutions; express A real matrix of ; Set a similarity threshold , and construct a similarity matrix with the solution Problems with the same dimensions and solutions to correlation matrices ; Solution similarity matrix middle Make a judgment, when Greater than the similarity threshold , the problem solution association matrix middle Set to 1, otherwise set to 0, expressed as: (4); Finally, we get the problem solution correlation matrix ; Indicates Programming Questions and Programming Questions Is there a solution association? When 1 means there is a solution association; A value of 0 indicates that there is no association; The predicted association probability of the problem vertex vector in step S3 is as follows: Calculate the first problem vertex vector And the second problem vertex vector The inner product between , obtain the predicted similarity, and map the predicted similarity to the predicted association probability of the problem vertex vector through the activation function , expressed as: (5); In the formula, is the activation function; Represents the input value; Represents a function; Indicates the first to A programming question; The cross entropy loss function is used to quantify the predicted association probability of the problem vertex vector and Problem Solution Correlation Matrix middle The difference between , i.e. the first difference, is expressed as: (6); In the formula, Represents a logarithmic function.

[0010] Going further, using programming questions Collection And knowledge points Collection Constructing a bipartite graph ,in is the vertex set of the bipartite graph, is a binary adjacency matrix representation ; Indicates Programming Questions With Knowledge Points Is there a correlation? When 1 means there is an association; when A value of 0 indicates that there is no association; The joint prediction association probability of the problem vertex vector and the knowledge point vertex vector in step S4 is as follows: The first problem vertex vector is transformed by the activation function With the first knowledge point vertex vector Between Measure the degree of association and process it to obtain the result of the processing, and map the result of the processing into the joint prediction association probability of the problem vertex vector and the knowledge point vertex vector ; The cross entropy loss function is used to quantify the joint prediction association probability of the question vertex vector and the knowledge point vertex vector. With the binary adjacency matrix middle The difference between , i.e. the second difference, is expressed as: (7).

[0011] Furthermore, the definition and Programming Questions The associated knowledge point set is , Indicates Programming Questions No. Knowledge Points The degree of correlation between And according to Programming Questions Related knowledge point collection Constructing a Programming Problem Relevance Matrix , express Programming Questions With Programming Questions Are there any common related knowledge points? 1 indicates that there are common related knowledge points. A value of 0 indicates that there are no commonly related knowledge points; The cross entropy loss function is used to quantify the predicted association probability of the problem vertex vector Programming Problem Relevance Matrix middle The difference between , which is the third difference, is expressed as: (8).

[0012] Furthermore, the definition and Knowledge Points The set of related programming problems is , Indicates Knowledge Points With Programming Questions The degree of correlation between And according to Knowledge Points A collection of related programming problems Constructing knowledge point relevance matrix , Indicates Knowledge Points With Knowledge Points Are there common related issues? 1 indicates that there is a common association problem; A value of 0 indicates that there is no common association problem; The predicted association probability of the knowledge point vertex vector in step S5 is as follows: By calculating the vertex vector of the first knowledge point and the second knowledge point vertex vector The inner product of After the activation function mapping, the predicted association probability of the knowledge point vertex vector is obtained ; The cross entropy loss function is used to quantify the predicted association probability of the knowledge point vertex vector Relevance matrix with knowledge points middle The difference between , which is the fourth difference, is expressed as: (9).

[0013] Furthermore, the final question embedding representation vector in step S6 is as follows: Question text vector Use linear neural network mapping to obtain low-dimensional representation , expressed as: (10); In the formula, Represents a low-dimensional question text vector; Indicates that the question text vector The weight matrix that is linearly mapped to a low-dimensional space; represents the transpose of the weight matrix; represents the bias term; Define the first problem vertex vector The associated knowledge point set is , the vertex vector of the first problem is calculated by arithmetic mean The associated knowledge point set is averaged to obtain the average knowledge point vector , expressed as: (11); In the formula, Represents the first problem vertex vector The number of relevant knowledge points; Predicted association probability based on question vertex vector , the joint prediction association probability of the question vertex vector and the knowledge point vertex vector , the predicted association probability of the knowledge point vertex vector The first problem vertex vector , average knowledge point vector and the low-dimensional problem text vector Concatenate in sequence to construct a first-order feature matrix ; First-order characteristic matrix From the first problem vertex vector , average knowledge point vector and the low-dimensional problem text vector Composition, that is , , ; and Represents the first-order characteristic matrix The first problem vertex vector , average knowledge point vector and the low-dimensional problem text vector Any two vectors; Defining interaction functions , by defining the interaction function Calculate the first-order characteristic matrix middle and The interaction value between ,based on Constructing the second-order characteristic matrix ; Represents the first-order characteristic matrix middle and The interaction value between represents a 3×3 real matrix; For the first-order characteristic matrix Select Weight matrix Generate the first A scalar, expressed as: (12); In the formula, Indicates scalar; Indicates A weight matrix; Represents the first-order characteristic matrix Middle Line Elements of a column; Represents element-wise multiplication and sum operation; For the second-order characteristic matrix use Weight matrix Generate the first A scalar, expressed as: (13); In the formula, Indicates scalar; Indicates A weight matrix; Represents the second-order characteristic matrix Middle Line Elements of a column; The first-order feature matrix No. scalar and second-order characteristic matrix No. scalar and bias vector After adding each element and mapping it through the activation function, we get the final question embedding representation vector , expressed as: (14); Where ReLU represents the activation function.

[0014] Further, define The actual difficulty of the question ; Use a linear layer to embed the final question into a representation vector Mapped to prediction difficulty value , expressed as: (15); In the formula, represents the weight vector; Use squared error to construct the prediction difficulty value With The actual difficulty of the problem Function , the fifth difference, is expressed as: (16); Finally, the first difference, the second difference, the third difference, the fourth difference and the fifth difference are jointly optimized and integrated into the objective function, which is expressed as: (17); In the formula, Represents minimizing the joint loss function; is the importance coefficient for balancing the constraints.

[0015] Compared with the existing technology, the present invention has the following beneficial effects: the present invention clusters multiple correct code submissions for each problem, refines multiple solutions to the problem at the code implementation level, and establishes the similarity between the solutions at the level of problems; by combining the similarity at the solution level with the problem correlation matrix of the knowledge points, a problem embedding representation based on the solution similarity is formed; compared with the traditional method that only relies on knowledge points or text features, the present invention can capture the correlation and diversity of the code implementation ideas of the problem; by jointly considering the diversity of solutions and the hierarchical correlation of knowledge points, the present invention effectively improves the expressive power and prediction accuracy of the problem embedding representation, and provides more comprehensive and accurate support for more in-depth problem analysis and student learning status diagnosis in the field of programming education. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0017] like Figure 1 As shown, the present invention provides a technical solution: a problem embedding representation method based on code space solution similarity, comprising the following steps: Step S1: construct a data set, which includes programming problems, knowledge points, and all submitted codes corresponding to the programming problems; Step S2: Process the programming problem and the knowledge point to obtain the problem vertex vector and the knowledge point vertex vector, set a problem text, and process the problem text to obtain the problem text vector; Step S3: by processing the problem vertex vector, the predicted association probability of the problem vertex vector is obtained; Step S4: by processing the problem vertex vector and the knowledge point vertex vector, a joint prediction association probability of the problem vertex vector and the knowledge point vertex vector is obtained; Step S5: obtaining the predicted association probability of the knowledge point vertex vector by processing the knowledge point vertex vector; Step S6: Use linear neural network mapping to obtain a low-dimensional question text vector for the question text vector; define a set of knowledge points associated with the question vertex vector, process the knowledge point set by arithmetic averaging to obtain an average knowledge point vector, concatenate the question vertex vector, the average knowledge point vector and the low-dimensional question text vector based on the predicted association probability of the question vertex vector, the joint predicted association probability of the question vertex vector and the knowledge point vertex vector, and the predicted association probability of the knowledge point vertex vector, construct a first-order feature matrix, process the first-order feature matrix by defining an interaction function, construct a second-order feature matrix, process the first-order feature matrix and the second-order feature matrix, obtain the scalar of the first-order feature matrix and the scalar of the second-order feature matrix, add the scalar of the first-order feature matrix and the scalar of the second-order feature matrix, and obtain the final question embedding representation vector.

[0018] Among them, For programming problems A collection of Indicates A programming question, is the total number of programming problems; let For knowledge points A collection of; Indicates Knowledge points; Indicates Knowledge points; is the total number of knowledge points; Submit code for all programming questions A collection of Indicates that the corresponding Programming Questions All submitted code.

[0019] Among them, the problem vertex vector includes the first problem vertex vector , the second problem vertex vector , the knowledge point vertex vector includes the first knowledge point vertex vector and the second knowledge point vertex vector ; Get the question text vector The specific process is as follows: define a question text, the large-scale pre-trained language model first decomposes the question text into a word sequence, and then processes the word sequence through the word embedding layer to obtain the word embedding sequence , For the word embedding sequence; finally, the question text vector is obtained by calculating the average value of the word embedding sequence .

[0020] Among them, let the corresponding Programming Questions All submitted codes represents a set of programming problems, Indicates Programming Questions No. code vectors submitted, Indicates Programming Questions The number of code vector submissions corresponding to Programming Questions All submitted codes Use the selected clustering algorithm Processing is performed to generate multiple cluster center vectors, and multiple cluster center vectors include the first Programming Questions and Programming Questions , multiple cluster center vectors are regarded as Programming Questions and Programming Questions The solution vector set and , expressed as: (1); (2); In the formula, Represents the number of cluster center vectors generated by the clustering algorithm; Indicates Programming Questions No. The center vector of each cluster; Indicates Programming Questions No. The center vector of each cluster; Indicates Programming Questions The set of cluster center vectors; Indicates Programming Questions The set of cluster center vectors; Solution vector set and Calculate the cosine value of the normal vector and take the maximum value as the first Programming Questions and Programming Questions The similarity at the solution level is expressed as: (3); In the formula, Indicates Programming Questions and Programming Questions similarities between; and Represents the solution vector set and Any solution center vector in ; Represents a transpose operation; represents the vector norm; Indicates the maximum value of similarity; Based on Programming Questions and Programming Questions Similarity at the solution level, constructing a solution similarity matrix ,in ; Indicates Programming Questions and Programming Questions The similarity of solutions; express A real matrix of ; Set a similarity threshold , and construct a similarity matrix with the solution Problems with the same dimensions and solutions to correlation matrices ; Solution similarity matrix middle Make a judgment, when Greater than the similarity threshold , the problem solution association matrix middle Set to 1, otherwise set to 0, expressed as: (4); Finally, we get the problem solution correlation matrix ; Indicates Programming Questions and Programming Questions Is there a solution association? When 1 means there is a solution association; A value of 0 indicates that there is no association; The predicted association probability of the problem vertex vector in step S3 is as follows: Calculate the first problem vertex vector And the second problem vertex vector The inner product between , obtain the predicted similarity, and map the predicted similarity to the predicted association probability of the problem vertex vector through the activation function , expressed as: (5); In the formula, is the activation function; Represents the input value; Represents a function; Indicates the first to A programming question; The cross entropy loss function is used to quantify the predicted association probability of the problem vertex vector and Problem Solution Correlation Matrix middle The difference between , i.e. the first difference, is expressed as: (6); In the formula, Represents a logarithmic function.

[0021] Among them, using programming problems Collection And knowledge points Collection Constructing a bipartite graph ,in is the vertex set of the bipartite graph, is a binary adjacency matrix representation ; Indicates Programming Questions With Knowledge Points Is there a correlation? When 1 means there is an association; when A value of 0 indicates that there is no association; The joint prediction association probability of the problem vertex vector and the knowledge point vertex vector in step S4 is as follows: The first problem vertex vector is transformed by the activation function With the first knowledge point vertex vector Between Measure the degree of association and process it to obtain the result of the processing, and map the result of the processing into the joint prediction association probability of the problem vertex vector and the knowledge point vertex vector ; The cross entropy loss function is used to quantify the joint prediction association probability of the question vertex vector and the knowledge point vertex vector. With the binary adjacency matrix middle The difference between , i.e. the second difference, is expressed as: (7).

[0022] The definition is the same as Programming Questions The associated knowledge point set is , Indicates Programming Questions No. Knowledge Points The degree of correlation between And according to Programming Questions Related knowledge point collection Constructing a Programming Problem Relevance Matrix , express Programming Questions With Programming Questions Are there any common related knowledge points? 1 indicates that there are common related knowledge points. A value of 0 indicates that there are no commonly related knowledge points; The cross entropy loss function is used to quantify the predicted association probability of the problem vertex vector Programming Problem Relevance Matrix middle The difference between , which is the third difference, is expressed as: (8).

[0023] The definition is the same as Knowledge Points The set of related programming problems is , Indicates Knowledge Points With Programming Questions The degree of correlation between And according to Knowledge Points A collection of related programming problems Constructing knowledge point relevance matrix , Indicates Knowledge Points With Knowledge Points Are there common related issues? 1 indicates that there is a common association problem; A value of 0 indicates that there is no common association problem; The predicted association probability of the knowledge point vertex vector in step S5 is as follows: By calculating the vertex vector of the first knowledge point and the second knowledge point vertex vector The inner product of After the activation function mapping, the predicted association probability of the knowledge point vertex vector is obtained ; The cross entropy loss function is used to quantify the predicted association probability of the knowledge point vertex vector Relevance matrix with knowledge points middle The difference between , which is the fourth difference, is expressed as: (9).

[0024] The final question embedding representation vector in step S6 is specifically as follows: Question text vector Use linear neural network mapping to obtain low-dimensional representation , expressed as: (10); In the formula, Represents a low-dimensional question text vector; Indicates that the question text vector The weight matrix that is linearly mapped to a low-dimensional space; represents the transpose of the weight matrix; represents the bias term; Define the first problem vertex vector The associated knowledge point set is , the vertex vector of the first problem is calculated by arithmetic mean The associated knowledge point set is averaged to obtain the average knowledge point vector , expressed as: (11); In the formula, Represents the first problem vertex vector The number of relevant knowledge points; Predicted association probability based on question vertex vector , the joint prediction association probability of the question vertex vector and the knowledge point vertex vector , the predicted association probability of the knowledge point vertex vector The first problem vertex vector , average knowledge point vector and the low-dimensional problem text vector Concatenate in sequence to construct a first-order feature matrix ; First-order characteristic matrix From the first problem vertex vector , average knowledge point vector and the low-dimensional problem text vector Composition, that is , , ; and Represents the first-order characteristic matrix The first problem vertex vector , average knowledge point vector and the low-dimensional problem text vector Any two vectors; Defining interaction functions , by defining the interaction function Calculate the first-order characteristic matrix middle and The interaction value between ,based on Constructing the second-order characteristic matrix ; Represents the first-order characteristic matrix middle and The interaction value between represents a 3×3 real matrix; For the first-order characteristic matrix Select Weight matrix Generate the first A scalar, expressed as: (12); In the formula, Indicates scalar; Indicates A weight matrix; Represents the first-order characteristic matrix Middle Line Elements of a column; Represents element-wise multiplication and sum operation; For the second-order characteristic matrix use Weight matrix Generate the first A scalar, expressed as: (13); In the formula, Indicates scalar; Indicates A weight matrix; Represents the second-order characteristic matrix Middle Line Elements of a column; The first-order feature matrix No. scalar and second-order characteristic matrix No. scalar and bias vector After adding each element and mapping it through the activation function, we get the final question embedding representation vector , expressed as: (14); Where ReLU represents the activation function.

[0025] Among them, the definition The actual difficulty of the question ; Use a linear layer to embed the final question into a representation vector Mapped to prediction difficulty value , expressed as: (15); In the formula, represents the weight vector; Use squared error to construct the prediction difficulty value With The actual difficulty of the problem Function , the fifth difference, is expressed as: (16); Finally, the first difference, the second difference, the third difference, the fourth difference and the fifth difference are jointly optimized and integrated into the objective function, which is expressed as: (17); In the formula, Represents minimizing the joint loss function; is the importance coefficient for balancing the constraints.

[0026] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A problem embedding representation method based on code space solution similarity, characterized in that: The steps include: Step S1: construct a data set, which includes programming problems, knowledge points, and all submitted codes corresponding to the programming problems; Step S2: Process the programming problem and the knowledge point to obtain the problem vertex vector and the knowledge point vertex vector, set a problem text, and process the problem text to obtain the problem text vector; Step S3: by processing the problem vertex vector, the predicted association probability of the problem vertex vector is obtained; Step S4: by processing the problem vertex vector and the knowledge point vertex vector, a joint prediction association probability of the problem vertex vector and the knowledge point vertex vector is obtained; Step S5: obtaining the predicted association probability of the knowledge point vertex vector by processing the knowledge point vertex vector; Step S6: Use a linear neural network to map the question text vector to obtain a low-dimensional question text vector; Define a set of knowledge points associated with the question vertex vector, process the knowledge point set through the arithmetic mean method to obtain the average knowledge point vector, concatenate the question vertex vector, the average knowledge point vector and the low-dimensional question text vector based on the predicted association probability of the question vertex vector, the joint predicted association probability of the question vertex vector and the knowledge point vertex vector and the predicted association probability of the knowledge point vertex vector, construct a first-order feature matrix, process the first-order feature matrix by defining an interaction function, construct a second-order feature matrix, process the first-order feature matrix and the second-order feature matrix, obtain the scalar of the first-order feature matrix and the scalar of the second-order feature matrix, add the scalar of the first-order feature matrix and the scalar of the second-order feature matrix, and obtain the final question embedding representation vector.

2. The problem embedding representation method based on code space solution similarity according to claim 1, characterized in that: set up For programming problems A collection of Indicates A programming question, is the total number of programming problems; let For knowledge points A collection of; Indicates Knowledge points; Indicates Knowledge points; is the total number of knowledge points; Submit code for all programming questions A collection of Indicates that the corresponding Programming Questions All submitted code.

3. The problem embedding representation method based on code space solution similarity according to claim 2, characterized in that: The problem vertex vectors include the first problem vertex vector , the second problem vertex vector , the knowledge point vertex vector includes the first knowledge point vertex vector and the second knowledge point vertex vector ; Get the question text vector The specific process is as follows: define a question text, the large-scale pre-trained language model first decomposes the question text into a word sequence, and then processes the word sequence through the word embedding layer to obtain the word embedding sequence , For the word embedding sequence; finally, the question text vector is obtained by calculating the average value of the word embedding sequence .

4. The problem embedding representation method based on code space solution similarity according to claim 3, characterized in that: Let the corresponding Programming Questions All submitted code represents a set of programming problems, Indicates Programming Questions No. code vectors submitted, Indicates Programming Questions The number of code vector submissions corresponding to Programming Questions All submitted code Use the selected clustering algorithm Processing is performed to generate multiple cluster center vectors, and multiple cluster center vectors include the first Programming Questions and Programming Questions , multiple cluster center vectors are regarded as Programming Questions and Programming Questions The solution vector set and , expressed as: (1); (2); In the formula, Represents the number of cluster center vectors generated by the clustering algorithm; Indicates Programming Questions No. The center vector of each cluster; Indicates Programming Questions No. The center vector of each cluster; Indicates Programming Questions The set of cluster center vectors; Indicates Programming Questions The set of cluster center vectors; Solution vector set and Calculate the cosine value of the normal vector and take the maximum value as the first Programming Questions and Programming Questions The similarity at the solution level is expressed as: (3); In the formula, Indicates Programming Questions and Programming Questions similarities between; and Represents the solution vector set and Any solution center vector in ; Represents a transpose operation; represents the vector norm; Indicates the maximum value of similarity; Based on Programming Questions and Programming Questions Similarity at the solution level, constructing a solution similarity matrix ,in ; Indicates Programming Questions and Programming Questions The similarity of solutions; express A real matrix of ; Set a similarity threshold , and construct a similarity matrix with the solution Problems with the same dimensions and solutions to correlation matrices ; Solution similarity matrix middle Make a judgment, when Greater than the similarity threshold When the problem solution association matrix middle Set to 1, otherwise set to 0, expressed as: (4); Finally, we get the problem solution correlation matrix ; Indicates Programming Questions and Programming Questions Is there a solution association? When 1 means there is a solution association; A value of 0 indicates that there is no association; The predicted association probability of the problem vertex vector in step S3 is as follows: Calculate the first problem vertex vector And the second problem vertex vector The inner product between , obtain the predicted similarity, and map the predicted similarity to the predicted association probability of the problem vertex vector through the activation function , expressed as: (5); In the formula, is the activation function; Represents the input value; Represents a function; Indicates the first to A programming question; The cross entropy loss function is used to quantify the predicted association probability of the problem vertex vector and Problem Solution Correlation Matrix middle The difference between , i.e. the first difference, is expressed as: (6); In the formula, Represents a logarithmic function.

5. The problem embedding representation method based on code space solution similarity according to claim 4, characterized in that: Exploiting Programming Problems Collection And knowledge points Collection Constructing a bipartite graph ,in is the vertex set of the bipartite graph, is a binary adjacency matrix representation ; Indicates Programming Questions With Knowledge Points Is there a correlation? 1 indicates that there is an association relationship; A value of 0 indicates that there is no association relationship; The joint prediction association probability of the problem vertex vector and the knowledge point vertex vector in step S4 is as follows: The first problem vertex vector is transformed by the activation function With the first knowledge point vertex vector Between Measure the degree of association and process it to obtain the result of the processing, and map the result of the processing into the joint predicted association probability of the problem vertex vector and the knowledge point vertex vector ; The cross entropy loss function is used to quantify the joint prediction association probability of the question vertex vector and the knowledge point vertex vector. With the binary adjacency matrix middle The difference between , i.e. the second difference, is expressed as: (7)。 6. The problem embedding representation method based on code space solution similarity according to claim 5, characterized in that: Definition and Programming Questions The associated knowledge point set is , Indicates Programming Questions No. Knowledge Points The degree of correlation between And according to Programming Questions Related knowledge point collection Constructing a Programming Problem Relevance Matrix , express Programming Questions With Programming Questions Are there any common related knowledge points? 1 indicates that there are common related knowledge points. A value of 0 indicates that there are no commonly related knowledge points; The cross entropy loss function is used to quantify the predicted association probability of the problem vertex vector Programming Problem Relevance Matrix middle The difference between , which is the third difference, is expressed as: (8)。 7. The problem embedding representation method based on code space solution similarity according to claim 6, characterized in that: Definition and Knowledge Points The set of related programming problems is , Indicates Knowledge Points With Programming Questions The degree of correlation between And according to Knowledge Points A collection of related programming problems Constructing knowledge point relevance matrix , Indicates Knowledge Points With Knowledge Points Are there common related issues? 1 indicates that there is a common association problem; A value of 0 indicates that there is no common association problem; The predicted association probability of the knowledge point vertex vector in step S5 is as follows: By calculating the vertex vector of the first knowledge point and the second knowledge point vertex vector The inner product of After the activation function mapping, the predicted association probability of the knowledge point vertex vector is obtained ; The cross entropy loss function is used to quantify the predicted association probability of the knowledge point vertex vector Relevance matrix with knowledge points middle The difference between , which is the fourth difference, is expressed as: (9)。 8. The problem embedding representation method based on code space solution similarity according to claim 7, characterized in that: The final question embedding representation vector in step S6 is as follows: Question text vector Use linear neural network mapping to obtain low-dimensional representation , expressed as: (10); In the formula, Represents a low-dimensional question text vector; Indicates that the question text vector The weight matrix that is linearly mapped to a low-dimensional space; represents the transpose of the weight matrix; represents the bias term; Define the first problem vertex vector The associated knowledge point set is , the vertex vector of the first problem is calculated by arithmetic mean The associated knowledge point set is averaged to obtain the average knowledge point vector , expressed as: (11); In the formula, Represents the first problem vertex vector The number of relevant knowledge points; Predicted association probability based on question vertex vector , the joint prediction association probability of the question vertex vector and the knowledge point vertex vector , the predicted association probability of the knowledge point vertex vector The first problem vertex vector , average knowledge point vector and the low-dimensional problem text vector Concatenate in sequence to construct a first-order feature matrix ; First-order characteristic matrix From the first problem vertex vector , average knowledge point vector and the low-dimensional problem text vector Composition, that is , , ; and Represents the first-order characteristic matrix The vertex vector in the first problem , average knowledge point vector and the low-dimensional problem text vector Any two vectors; Defining interaction functions , by defining the interaction function Calculate the first-order characteristic matrix middle and The interaction value between ,based on Construct the second-order characteristic matrix ; Represents the first-order characteristic matrix middle and The interaction value between represents a 3×3 real matrix; For the first-order characteristic matrix Select Weight Matrix Generate the first A scalar, expressed as: (12); In the formula, Indicates scalar; Indicates A weight matrix; Represents the first-order characteristic matrix Middle Line Elements of a column; Represents element-wise multiplication and sum operation; For the second-order characteristic matrix use Weight Matrix Generate the first A scalar, expressed as: (13); In the formula, Indicates scalar; Indicates A weight matrix; Represents the second-order characteristic matrix Middle Line Elements of a column; The first-order feature matrix No. scalar and second-order characteristic matrix No. scalar and bias vector After adding each element and mapping it through the activation function, we get the final question embedding representation vector , expressed as: (14); Where ReLU represents the activation function.

9. The problem embedding representation method based on code space solution similarity according to claim 8, characterized in that: Definition The actual difficulty of the question ; Use a linear layer to embed the final question into a representation vector Mapped to prediction difficulty value , expressed as: (15); In the formula, represents the weight vector; Use squared error to construct the prediction difficulty value With The actual difficulty of the problem Function , the fifth difference, is expressed as: (16); Finally, the first difference, the second difference, the third difference, the fourth difference and the fifth difference are jointly optimized and integrated into the objective function, which is expressed as: (17); In the formula, Represents minimizing the joint loss function; is the importance coefficient for balancing the constraints.

Citation Information

Patent Citations

  • Heterogeneous graph-based pre-training problem characterization method, system and device, and medium

    CN116303952A

  • Knowledge tracking method for enhancing topic similarity embedding based on weighted meta-path

    CN116502713A

  • Knowledge proficiency calculation method based on Bloom cognitive theory

    CN117556343A

  • Multi-task information enhancement test question recommendation method based on video understanding

    CN118296243A

  • Programming knowledge tracking method fusing code and score information

    CN118569447A