Knowledge point prediction method and device, electronic equipment and storage medium

By using a pre-trained knowledge point prediction model and processing test text with self-attention and knowledge point attention modules, the problems of slow recall speed and dependence on question banks in existing methods are solved, and efficient and accurate test knowledge point prediction is achieved.

CN115630696BActive Publication Date: 2026-05-15BEIJING CENTURY TAL EDUCATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING CENTURY TAL EDUCATION TECH CO LTD
Filing Date
2022-10-18
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing test question knowledge point prediction methods have slow recall speed, cannot accurately predict newly emerging test question knowledge points outside the question bank, and rely on a large question bank reserve.

Method used

A pre-trained knowledge point prediction model is used to process text data using a self-attention module and a knowledge point attention module to determine the representation vector of the questions to be tested. The output module and post-processing module are used to classify knowledge points and determine their probabilities, and to query the correspondence between preset numbers and knowledge points.

Benefits of technology

It enables accurate prediction of test knowledge points, avoids dependence on question banks, improves the accuracy and efficiency of prediction, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115630696B_ABST
    Figure CN115630696B_ABST
Patent Text Reader

Abstract

The present disclosure provides a knowledge point prediction method and device, electronic equipment and storage medium, the method comprises: obtaining text data of a test question to be predicted; using the output results of the self-attention module and the knowledge point attention module of the pre-trained knowledge point prediction model to process the text data respectively to determine the representation vector of the test question to be predicted; using the output module of the knowledge point prediction model, classifying the knowledge points according to the representation vector to obtain a classification result; using the post-processing module of the knowledge point prediction model, determining the prediction probability of each knowledge point in all preset knowledge points to which the test question to be predicted belongs according to the classification result, and determining the knowledge point number corresponding to the test question to be predicted according to the prediction probability; querying the preset correspondence between the number and the knowledge point to determine the target knowledge point corresponding to the knowledge point number. The present scheme can accurately predict the knowledge points in the test question.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of machine learning technology, and in particular to a knowledge point prediction method, apparatus, electronic device, and storage medium. Background Technology

[0002] In test question input and personalized test question recommendation systems, accurately predicting the knowledge points in test questions is crucial. Currently, the commonly used method for predicting test knowledge points is the similar question recall method based on the question stem text. However, this method is slow, relies on a large question bank, and cannot accurately predict the knowledge points of newly added test questions outside the question bank. Summary of the Invention

[0003] To solve the above-mentioned technical problems, or at least partially solve them, this disclosure provides a knowledge point prediction method, apparatus, electronic device, and storage medium.

[0004] According to one aspect of this disclosure, a knowledge point prediction method is provided, comprising:

[0005] Obtain the text data of the questions to be tested;

[0006] The representation vector of the question to be tested is determined by processing the text data using the self-attention module and knowledge point attention module of the pre-trained knowledge point prediction model.

[0007] Using the output module of the knowledge point prediction model, knowledge points are classified according to the representation vector to obtain the classification result;

[0008] Using the post-processing module of the knowledge point prediction model, the predicted probability of the question to be tested belonging to each of the preset knowledge points is determined according to the classification result, and the knowledge point number corresponding to the question to be tested is determined according to the predicted probability.

[0009] Query the preset correspondence between numbers and knowledge points to determine the target knowledge point corresponding to the knowledge point number.

[0010] According to another aspect of this disclosure, a knowledge point prediction device is provided, comprising:

[0011] The acquisition module is used to acquire the text data of the questions to be tested.

[0012] The first determining module is used to determine the representation vector of the question to be tested by processing the text data using the output results of the self-attention module and the knowledge point attention module of the pre-trained knowledge point prediction model; to classify the knowledge points according to the representation vector using the output module of the knowledge point prediction model to obtain the classification result; and to determine the prediction probability of the question to be tested belonging to each of the preset knowledge points according to the classification result using the post-processing module of the knowledge point prediction model, and to determine the knowledge point number corresponding to the question to be tested according to the prediction probability.

[0013] The second determination module is used to query the correspondence between preset numbers and knowledge points, and to determine the target knowledge point corresponding to the knowledge point number.

[0014] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0015] Processor; and

[0016] Stored program memory,

[0017] The program includes instructions that, when executed by the processor, cause the processor to perform the knowledge point prediction method according to the foregoing aspect.

[0018] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the knowledge point prediction method according to the foregoing aspect.

[0019] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the knowledge point prediction method described in the foregoing aspect.

[0020] One or more technical solutions provided in this disclosure acquire text data of a test question to be tested, utilize a pre-trained knowledge point prediction model, determine the knowledge point number corresponding to the test question based on the text data, wherein the knowledge point prediction model includes a self-attention module and a knowledge point attention module, utilizes the output results of the self-attention module and the knowledge point attention module to determine the representation vector of the test question to be tested, queries a preset correspondence between the number and the knowledge point, and determines the target knowledge point corresponding to the knowledge point number. Using the solutions of this disclosure, the knowledge points in the test questions can be accurately predicted. Attached Figure Description

[0021] Further details, features, and advantages of this disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0022] Figure 1 A flowchart of a knowledge point prediction method according to an exemplary embodiment of the present disclosure is shown;

[0023] Figure 2 A schematic diagram of the structure of a knowledge point prediction model provided according to an exemplary embodiment of the present disclosure is shown;

[0024] Figure 3 A flowchart of a knowledge point prediction method according to another exemplary embodiment of this disclosure is shown;

[0025] Figure 4 A flowchart illustrating a training method for a knowledge point prediction model according to an exemplary embodiment of the present disclosure is shown.

[0026] Figure 5 A flowchart illustrating a training method for a knowledge point prediction model according to another exemplary embodiment of the present disclosure is shown;

[0027] Figure 6 A schematic block diagram of a knowledge point prediction apparatus according to an exemplary embodiment of the present disclosure is shown;

[0028] Figure 7 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0029] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0030] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0031] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0032] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0033] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0034] Before explaining the methods provided in this disclosure, the terms that may be used in this disclosure will be explained, wherein:

[0035] BERT: Bidirectional Encoder Representation from Transformers, is a pre-trained language representation model that leverages the specificities of pre-training tasks to generate deep bidirectional language models.

[0036] HIVE: A data warehouse tool that can be used to store data.

[0037] HTML: The full name of HTML is Hyper Text Markup Language, which is a markup language.

[0038] LaTeX: A typesetting system.

[0039] Word vectors: a type of vectorized data that can be used to represent the semantic information of words.

[0040] The knowledge point prediction method, apparatus, electronic device and storage medium provided in this disclosure are described below with reference to the accompanying drawings.

[0041] Currently, the common method for personalized test question recommendation is to recommend questions based on the question stem information through processes such as similarity recall, coarse ranking, and fine ranking. This method only recommends questions based on textual similarity and does not consider the knowledge point information, or even multiple knowledge points. For example, consider the question: "Press a steel ruler firmly on a table, first extend one end slightly beyond the edge of the table, pluck the ruler, and listen to the sound it produces. Then extend one end further beyond the edge of the table and pluck the ruler again, listening to the sound it produces. Comparing the two cases, the pitch of the second vibration is higher. This shows that to study the relationship between loudness and amplitude, we must ensure that we pluck the steel ruler with different forces and listen to the loudness of the sound." In the question bank, the corresponding knowledge point tags are: 1. Relationship between loudness and amplitude; 2. Relationship between pitch and frequency. From the above example, it is clear that the textual information corresponding to the knowledge points of the question is of great significance for test question recommendation. In particular, "pitch" is more important for the label "relationship between pitch and frequency", while "loudness, amplitude" is more important for the label "relationship between loudness and amplitude".

[0042] Regarding test question input, current methods of labeling knowledge points often rely on manual annotation, where teachers label the corresponding knowledge points based on their personal experience. If a test question contains too many knowledge points, it becomes difficult to comprehensively consider all the information. Furthermore, finding the relevant knowledge points from numerous tags is time-consuming and laborious.

[0043] To address the aforementioned problems, this disclosure provides a knowledge point prediction method. It involves acquiring text data of the test questions to be predicted, processing the text data using the outputs of a pre-trained knowledge point prediction model's self-attention module and knowledge point attention module, determining the representation vector of the test questions, and then using the output module of the knowledge point prediction model to classify the knowledge points based on the representation vectors. Finally, using the post-processing module of the knowledge point prediction model, it determines the predicted probability that the test questions belong to each of a preset set of knowledge points based on the classification results, and determines the corresponding knowledge point number based on the predicted probability. Then, it queries the preset correspondence between the numbers and knowledge points to determine the target knowledge point corresponding to the knowledge point number. Using the scheme of this disclosure, the knowledge points in the test questions can be accurately predicted.

[0044] Figure 1 A flowchart of a knowledge point prediction method according to an exemplary embodiment of the present disclosure is shown. The method can be executed by a knowledge point prediction device, which can be implemented in software and / or hardware and is generally integrated into an electronic device, including devices such as mobile phones, tablets, and servers.

[0045] like Figure 1 As shown, this knowledge point prediction method may include the following steps:

[0046] Step 101: Obtain the text data of the questions to be tested.

[0047] The test questions can be either text or images. When the test questions are images, Optical Character Recognition (OCR) can be used to perform image recognition and extract the text information of the test questions from the image.

[0048] In this embodiment of the disclosure, the text data of the questions to be tested can be obtained.

[0049] For example, the stem of the questions to be tested can be cleaned, and the cleaned text can be used as the text data of the questions to be tested. Data cleaning can include, but is not limited to, cleaning some formatting and non-standard LaTeX symbols in HTML. For example, regular expressions can be defined to remove spaces, ampersands, and other symbols from the question stem.

[0050] For example, the text data of the questions to be tested can be processed structured data. Structured data describes the question stem text in the form of labels. However, some labels in the structured data do not contain semantic information and are redundant for the task of predicting knowledge points. Therefore, the obtained structured data of the questions to be tested needs to be standardized. Standardization specifically includes deleting useless labels, which can improve the accuracy and computation speed of the model to a certain extent. For example, removing labels such as "\u3000", "", "\xa0", "\n", "\t", "\\left", "\\right", "\\text", "[\{\}]", "\$\$", etc., from the structured data of the questions to be tested can be standardized. "、" "、" "、"After deleting unnecessary tags, only the tags necessary for the question stem text are retained. Then, the target tags included in the structured data of the questions to be tested are replaced. This unifies the tags representing the same characters in structured data obtained from different methods, preventing tags from incorrectly forming other tags or obscuring the current tag when connected to the question stem text. For example, "Rt" is replaced with "right angle", "||||" with "", "\\\\prime" with "\'", "\f" with "\\\\f", and "\a" with "\\\\a". "Rt" is denoted as the target tag, and "right angle" is the replaced tag. This unifies the tags and improves the accuracy of model recognition. Secondly, replacing "\f" with "\\\\f" facilitates the differentiation of similar tags, and replacing "\\\\prime" with "\'" simplifies the tags. Finally, after standardization operations such as deleting unnecessary tags and replacing tags, the text data of the questions to be tested is obtained.

[0051] Step 102: Using the output results of the self-attention module and knowledge point attention module of the pre-trained knowledge point prediction model to process the text data, the representation vector of the question to be tested is determined.

[0052] In this embodiment of the disclosure, the obtained text data of the questions to be tested can be input into a pre-trained knowledge point prediction model, and the knowledge point prediction model can input the knowledge point number corresponding to the questions to be tested.

[0053] The knowledge point prediction model includes at least a self-attention module and a knowledge point attention module. The main function of the self-attention module is to calculate the contribution of each word in the test question text to all knowledge points using the test question text data. The main function of the knowledge point attention module is to calculate the contribution of each word in the test question text to all knowledge points using the text information of the knowledge points. Therefore, the output results of processing the text data of the test questions by the self-attention module and the knowledge point attention module respectively determine the representation vector of the test question. The representation vector of the test question is used to determine the prediction probability of the test question belonging to each knowledge point, and then the knowledge point number corresponding to the test question is determined based on the prediction probability.

[0054] Step 103: Using the output module of the knowledge point prediction model, classify the knowledge points according to the representation vector to obtain the classification result.

[0055] In this embodiment of the disclosure, the knowledge point prediction model further includes an output module. The main function of the output module is to classify the knowledge point labels by passing the representation vector of the test questions output from the hybrid module through a classification layer.

[0056] Therefore, in this embodiment of the disclosure, the representation vector of the determined test question is input into the output module. In the output module, knowledge points are classified according to the representation vector of the test question to obtain the classification result.

[0057] For example, the representation vector of the determined test question is classified into knowledge points in the output module through a classification layer. The parameters of the classification layer are (h, n_l), and the classification result is output. The classification result can be determined by the following formula (1).

[0058] Logits = nn.Linear(Ques') (1)

[0059] Here, nn.Linear() is used to set up fully connected layers in the network to classify labels, Logits represents the classification result, which is the unnormalized probability of each label, and Ques' represents the representation vector of the question to be tested.

[0060] Step 104: Using the post-processing module of the knowledge point prediction model, determine the predicted probability that the question to be tested belongs to each of the preset knowledge points based on the classification result, and determine the knowledge point number corresponding to the question to be tested based on the predicted probability.

[0061] Here, "all knowledge points" refers to all knowledge points that are pre-set for the knowledge point prediction model. For example, if the knowledge point prediction model can output prediction results for 2000 knowledge points, that is, the knowledge point prediction model predicts the probability that the test question belongs to these 2000 knowledge points, then "all knowledge points" refers to these 2000 knowledge points.

[0062] In this embodiment of the disclosure, the knowledge point prediction module further includes a post-processing module. The main function of the post-processing module is to process the classification results output by the output module to obtain the prediction probability of each knowledge point, and to determine the final predicted knowledge point number based on the prediction probability.

[0063] Therefore, in this embodiment of the disclosure, the classification results output by the output module are processed by the post-processing module. First, the predicted probability of the question to be tested belonging to each of the preset knowledge points is determined based on the classification results.

[0064] For example, in the post-processing module, the classification results output by the output module are first normalized by the sigmoid layer to obtain the prediction probability corresponding to each knowledge point. The prediction probability can be obtained by the following formula (2).

[0065] Out = Sigmoid(Logits) (2)

[0066] Where Out represents the prediction probability corresponding to each knowledge point.

[0067] Next, based on the predicted probability of each knowledge point, the knowledge point number corresponding to the question to be tested can be determined.

[0068] As an optional implementation, the knowledge point number corresponding to the question to be tested can be selected from all the knowledge point numbers based on the predicted probability of each knowledge point, using a preset number of numbers with the highest predicted probability. The preset number can be set according to actual needs; for example, it can be set to two.

[0069] As an optional implementation, the maximum predicted probability can be determined from the predicted probabilities of each knowledge point. Then, based on the floor function of 1 and the maximum predicted probability, the number of knowledge points corresponding to the question to be tested can be obtained. Next, based on the predicted probability of each knowledge point, the numbers corresponding to each knowledge point in all knowledge points are sorted in descending order of predicted probability to obtain the sorting result. Then, the number of the number of knowledge points that ranks first in the sorting result is determined as the knowledge point number corresponding to the question to be tested.

[0070] For example, suppose that the maximum predicted probability for each knowledge point is 0.4, and the floor of 1 / 0.4 is 2. Then, the number of knowledge points corresponding to the question to be tested is two. Then, according to the predicted probability for each knowledge point, all knowledge point numbers are sorted in descending order of predicted probability. The first two numbers in the sorted results are selected as the knowledge point numbers corresponding to the question to be tested.

[0071] Step 105: Query the preset correspondence between the number and the knowledge point to determine the target knowledge point corresponding to the knowledge point number.

[0072] Each knowledge point has a pre-defined number. After setting the number for each knowledge point, a correspondence between the number and the knowledge point is established and saved for future reference.

[0073] In this embodiment of the disclosure, after determining the knowledge point number corresponding to the question to be tested using a pre-trained knowledge point prediction model, the correspondence between the preset number and the knowledge point can be queried to determine the target knowledge point corresponding to the knowledge point number, which is the knowledge point that matches the question to be tested.

[0074] The knowledge point prediction method of this disclosure acquires the text data of the test question to be tested, processes the text data using the output results of the self-attention module and knowledge point attention module of a pre-trained knowledge point prediction model, determines the representation vector of the test question, classifies the knowledge points based on the representation vector using the output module of the knowledge point prediction model, obtains the classification result, and uses the post-processing module of the knowledge point prediction model to determine the prediction probability of the test question belonging to each of all preset knowledge points based on the classification result, determines the knowledge point number corresponding to the test question based on the prediction probability, and then queries the preset correspondence between the number and the knowledge point to determine the target knowledge point corresponding to the knowledge point number. Using the scheme of this disclosure, a knowledge point prediction model that integrates the text information of the test question itself and the text information of the knowledge points is obtained through pre-training. Using the model to determine the knowledge points corresponding to the test question can accurately predict the knowledge points in the test question, without relying on a question bank or requiring prior question bank preparation. It offers high accuracy and low cost in predicting knowledge points.

[0075] In one optional embodiment of this disclosure, the knowledge point prediction model further includes an encoding module and a hybrid module. Figure 2 A schematic diagram of the structure of a knowledge point prediction model provided according to an exemplary embodiment of the present disclosure is shown. Figure 2 In this module, the main function is to obtain the word vector representation of each word (token) in the test text. The input to the encoding module is each word in the test text (e.g., ...). Figure 2 The output (w1, w2, w3, w4, w5 shown in the figure) is the word vector representation corresponding to each word in the test text (e.g., w1, w2, w3, w4, w5). Figure 2 As shown in h1, h2, h3, h4, and h5, the encoding module can use a further-pretained BERT as the encoder. BERT's Transformer Encoder structure allows it to capture information from both the left and right sides of the text. Furthermore, the further-pretained process makes the original model closer to the downstream task, improving the model's accuracy. Because each word (token) contributes differently to each knowledge point, to achieve this, the encoding layer obtains a vector representation of each token instead of the entire sentence. The main function of the hybrid module is to fuse the features of the self-attention module and the knowledge point attention module to construct the representation vector of the question. The main function of the output module is to classify the knowledge point labels by passing the representation vector of the question output from the hybrid module through a classification layer. The main function of the post-processing module is to process the classification results output by the output module to obtain the predicted probability of each knowledge point and determine the final predicted knowledge point number based on the predicted probability.

[0076] Therefore, in the embodiments of this disclosure, as Figure 3 As shown, based on the foregoing embodiments, the knowledge point prediction method provided in this disclosure may include the following steps:

[0077] Step 201: Obtain the text data of the questions to be tested.

[0078] It should be noted that the description of step 201 can be found in the description of step 101 in the foregoing embodiments, and will not be repeated here.

[0079] Step 202: Determine the word vector representation of each word in the text data through the encoding module of the pre-trained knowledge point prediction model.

[0080] In this embodiment of the disclosure, the text data of the questions to be tested is first encoded by the encoding module of the knowledge point prediction model to obtain the word vector representation corresponding to each word in the text data.

[0081] For example, suppose b represents the number of questions in the knowledge point prediction model for each input, s represents the number of tokens in each question, h represents the vector representation dimension of each token, Ti represents the word vector representation of each token, and H is T1, T2...T s The matrix representation of vectors is the word vector matrix formed by the word vectors corresponding to each word in the test question, and the dimension of H is s*h.

[0082] Step 203: Using the self-attention module of the knowledge point prediction model, determine the contribution matrix based on the test question according to the word vector representation of each word.

[0083] In this embodiment of the disclosure, the word vector representation of each word in the pre-test question output by the encoding module is sent to the self-attention module and the knowledge point attention module of the knowledge point prediction model for further processing.

[0084] In the self-attention module, the self-attention module determines the contribution matrix based on the word vector representation of each word in the pre-test question. The contribution matrix based on the question is used to characterize the contribution of each word in the pre-test question to all knowledge points.

[0085] In one optional embodiment of this disclosure, the self-attention module can perform a linear transformation on the word vector matrix composed of the word vector representations of each word in the pre-test question, and then add a softmax layer on the token dimension to normalize the result after the linear transformation. Then, the result of the normalization process is multiplied by the transpose of the word vector matrix to obtain the contribution of each word in the pre-test question to different knowledge points. The result is a matrix called the contribution matrix based on the test question, and the dimension of the matrix is ​​n_1*s.

[0086] In one optional embodiment of this disclosure, to make the model training process more stable, the word vector representation of each word output by the encoding module can first pass through a nonlinear layer during model training. Therefore, during the model training phase, the word vector representation of each word output by the encoding module can also first undergo a nonlinear transformation through a nonlinear layer. Thus, in this embodiment of the disclosure, when determining the contribution matrix based on the test question, a nonlinear transformation can first be performed on the word vector matrix composed of the word vector representations of each word in the test question to obtain the nonlinear matrix corresponding to the word vector matrix. Then, the nonlinear matrix obtained by the nonlinear transformation is linearly transformed to generate a transformation matrix. Next, the transformation matrix is ​​normalized in terms of word dimension to obtain a normalized matrix. Then, based on the product of the normalized matrix and the transpose of the word vector matrix, the contribution matrix based on the test question is determined, and the dimension of this matrix is ​​n_1*s.

[0087] For example, a nonlinear transformation of the word vector matrix can be achieved by the following formula (3), a linear transformation of the nonlinear matrix obtained by the nonlinear transformation can be achieved by the following formula (4), a normalization of the transformation matrix can be achieved by the following formula (5), and the contribution matrix based on the test questions can be calculated by the following formula (6).

[0088] Hs=tanh(H) (3)

[0089] Ha = Linear(Hs) (4)

[0090] Ha' = softmax(Ha, dim = 1) (5)

[0091] Self_Att=Ha'*H T (6)

[0092] Where H represents the word vector matrix composed of the word vector representations of each word in the test question, tanh represents the nonlinear transformation function, and Hs represents the result of the nonlinear transformation of the word vector matrix; Linear represents the linear transformation function, and Ha represents the transformation matrix obtained by performing a linear transformation on the result Hs of the nonlinear transformation, with corresponding parameters (h, n_l), where h represents the dimension of the vector representation of each word, i.e., the dimension of the word vector representation, and n_l represents the number of all knowledge points; dim = 1 is used to indicate normalization processing on the word (token) dimension, and Ha' represents the result of normalization, i.e., the normalization matrix; H T represents the transpose of the word vector matrix, and Self_Att represents the contribution matrix based on the test items.

[0093] It is understandable that the contribution matrix based on the test questions output by the self-attention module is based entirely on the text data of the test questions to be tested, without taking into account the text information of the knowledge points.

[0094] Step 204: Using the knowledge point attention module of the knowledge point prediction model, determine the contribution matrix based on knowledge points according to the word vector representation of each word and the feature representation of each knowledge point.

[0095] from Figure 2 As can be seen from the structure of the knowledge point prediction model shown, in this embodiment, the word vector representation of each word in the question to be predicted, output by the encoding module, is also input to the knowledge point attention module for processing. In the knowledge point attention module, a knowledge point matrix can be determined based on the feature representation of each knowledge point, wherein the dimension of the feature representation of each knowledge point is the same as the dimension of the word vector representation. Then, based on the product of the knowledge point matrix and the transpose of the word vector matrix formed by the word vector representation of each word, a contribution matrix based on the knowledge points is determined.

[0096] The feature representation of each knowledge point can be obtained through random initialization or by using other vectorization methods to vectorize each knowledge point. To ensure that the knowledge points can be directly calculated with the word vector representations of each word, the feature representation dimension of the knowledge point is the same as the word vector representation dimension.

[0097] For example, suppose we use matrix HL to represent the knowledge point matrix, and the vectors HL1, HL2...HL in the knowledge point matrix are... n_l Let n_l represent the feature representation of the corresponding knowledge point, and n_l represent the number of all knowledge points. Then the contribution matrix based on knowledge points can be calculated by the following formula (7).

[0098] Label_Att = HL * H T (7)

[0099] Among them, H T represents the transpose of the word vector matrix, and Label_Att represents the contribution matrix based on knowledge points, with the dimension of the contribution matrix based on knowledge points being n_1*s.

[0100] It is understandable that the contribution matrix output by the knowledge point attention module takes into account the textual information of the knowledge points, which helps to improve the accuracy of knowledge point prediction.

[0101] Step 205: Using the hybrid module of the knowledge point prediction model, determine the representation vector of the question to be predicted based on the contribution matrix based on the test question and the contribution matrix based on the knowledge point.

[0102] In this embodiment of the disclosure, the hybrid module can be used to fuse the contribution matrix based on test questions output by the self-attention module and the contribution matrix based on knowledge points output by the knowledge point attention module to construct the representation vector of the test question to be tested.

[0103] In one optional embodiment of this disclosure, the contribution matrix based on test questions and the contribution matrix based on knowledge points can be averaged to obtain the test question representation matrix of the test question to be tested. Each row of the test question representation matrix represents the feature representation of a knowledge point of the test question to be tested. Then, for each row of the test question representation matrix, the element value with the largest value in that row is selected. This element value is an element in the representation vector of the test question to be tested. Thus, the element values ​​selected from all rows constitute a vector, and the representation vector of the test question to be tested is obtained.

[0104] In one optional implementation of this disclosure, a first weight corresponding to the contribution matrix based on test questions and a second weight corresponding to the contribution matrix based on knowledge points can be obtained first. Then, based on the first and second weights, the contribution matrix based on test questions and the contribution matrix based on knowledge points are weighted and summed to obtain a test question representation matrix. Each row of the test question representation matrix represents the feature representation of a knowledge point for the test question to be tested. Finally, based on the test question representation matrix, the mean of each row in the test question representation matrix is ​​calculated. That is, for the feature representation of each knowledge point for the test question to be tested, the vector is converted into a value by the mean calculation. This value is an element of the representation vector of the test question to be tested. Thus, the mean values ​​calculated for all rows in the test question representation matrix constitute a vector, which is the representation vector of the test question to be tested. The number of elements in the representation vector is consistent with the number of rows in the test question representation matrix.

[0105] As an optional approach, the first and second weights can be preset according to actual needs, such as setting the first weight to 0.4 and the second weight to 0.6.

[0106] As an optional approach, the first and second weights can be determined based on the contribution matrix based on the test items and the contribution matrix based on the knowledge points. Specifically, a linear transformation function can be used to perform a dimensionality transformation on the contribution matrix based on the test items to obtain a first transformation result, and a dimensionality transformation can be performed on the contribution matrix based on the knowledge points to obtain a second transformation result. Then, an activation function is used to scale the first transformation result to determine the self-attention weights corresponding to the self-attention modules, and the activation function is used to scale the second transformation result to determine the knowledge point weights corresponding to the knowledge point attention modules. Next, the first weight is determined based on the ratio of the self-attention weights to the knowledge point weights; that is, the ratio of the self-attention weights to the knowledge point weights is used to determine the first weight. Finally, the second weight is determined based on the first weight, where the sum of the first and second weights is 1.

[0107] The linear transformation function can be a commonly used linear function, such as Linear, which can achieve dimension transformation. For example, if a matrix has dimensions [100, 20], after passing through a Linear layer with parameters [20, 5], the resulting matrix will have dimensions [100, 5]. The activation function can be a commonly used activation function, such as Sigmoid, which can scale the input values ​​to the range (0, 1).

[0108] For example, the self-attention weight can be calculated by the following formula (8), and the knowledge point weight can be calculated by the following formula (9). The self-attention weight and the knowledge point weight are simply normalized to obtain the first weight and the second weight. The first weight and the second weight can be calculated by the following formula (10).

[0109] Score_self=Sigmoid(Linear(Self_Att)) (8)

[0110] Score_lab=Sigmoid(Linear(Label_Att)) (9)

[0111] Score_self'=Score_self / Score_lab (10)

[0112] Score_lab' = 1 - Score_self

[0113] Wherein, Linear represents a linear function, which performs linear transformations on the contribution matrices based on test items and the contribution matrices based on knowledge points, respectively, with the corresponding parameters being (h,1). Sigmoid represents the activation function, Score_self represents the self-attention weight, Score_lab represents the knowledge point weight, Score_self' represents the first weight, and Score_lab' represents the second weight.

[0114] Therefore, the test item representation matrix can be calculated using the following formula (11). The average value of the test item information (i.e., the test item representation matrix) representing all knowledge points is taken along the knowledge point dimension to obtain the final representation vector of the test item to be tested. The representation vector of the test item to be tested can be calculated using the following formula (12).

[0115] Ques=Score_self'*Self_Att+Score_lab'*Label_Att (11)

[0116] Ques'=sum(Ques,dim=1) / n_l (12)

[0117] Where Ques represents the question representation matrix of the questions to be tested, Ques' represents the representation vector of the questions to be tested, dim=1 indicates that the question representation matrix is ​​calculated in the knowledge point dimension, and n_l represents the number of all knowledge points.

[0118] Step 206: Using the output module of the knowledge point prediction model, classify the knowledge points according to the representation vector to obtain the classification result.

[0119] Step 207: Using the post-processing module of the knowledge point prediction model, determine the predicted probability that the question to be tested belongs to each of the preset knowledge points based on the classification result, and determine the knowledge point number corresponding to the question to be tested based on the predicted probability.

[0120] Step 208: Query the preset correspondence between the number and the knowledge point to determine the target knowledge point corresponding to the knowledge point number.

[0121] It should be noted that the description of steps 206 to 208 can be found in the description of steps 103 to 105 in the previous embodiments, and the implementation principle is similar, so it will not be repeated here.

[0122] The knowledge point prediction method of this disclosure includes: determining the word vector representation of each word in the text data of the question to be predicted through the encoding module of the knowledge point prediction model; determining a contribution matrix based on the question using a self-attention module based on the word vector representation of each word; determining a contribution matrix based on the knowledge point using a knowledge point attention module based on the word vector representation of each word and the feature representation of each knowledge point; determining the representation vector of the question to be predicted using a hybrid module based on the contribution matrix based on the question and the contribution matrix based on the knowledge point; and classifying the knowledge points based on the representation vector using an output module to obtain the classification result. Finally, the post-processing module determines the predicted probability of the test question belonging to each of the preset knowledge points based on the classification results, and determines the corresponding knowledge point number based on the predicted probability. Then, it queries the preset correspondence between the number and the knowledge point to determine the target knowledge point corresponding to the knowledge point number. Thus, by fusing the contribution matrix based on the test question and the contribution matrix based on the knowledge point to determine the representation vector of the test question, it not only considers the text information of the test question, but also the text information of the knowledge point. Therefore, when using the obtained representation vector for knowledge point prediction, the accuracy of knowledge point prediction can be improved.

[0123] Figure 4 A flowchart illustrating a training method for a knowledge point prediction model according to an exemplary embodiment of this disclosure is shown, which enables the training of the knowledge point prediction model described in the foregoing embodiments. Figure 4 As shown, the training process of the knowledge point prediction model includes the following steps:

[0124] Step 301: Obtain a first training set and a second training set. The first training set includes a first test question sample, and the second training set includes a second test question sample and the knowledge point tags corresponding to the second test question sample.

[0125] The first training set and the second training set can be determined based on the same question set or different question sets; this disclosure does not impose any restrictions on this. When the first training set and the second training set are determined based on the same question set, the first question sample and the second question sample can be the same or different; for example, the number of first question samples can be greater than the number of second question samples. The first training set only contains question samples and does not contain the corresponding knowledge points (i.e., labels), and can be used for unsupervised training. The second training set contains both question samples and labels, and can be used for supervised training.

[0126] For example, the first training set can be determined based on a set of physics test questions, with the aim of better adapting to the downstream task of predicting physics test knowledge points and improving the accuracy of knowledge point prediction. When it is necessary to train a model for predicting knowledge points applied to the physics subject, the second training set can also be determined based on a set of physics test questions. When it is necessary to train a model for predicting knowledge points applied to other subjects, the second training set can be determined based on the test question set of the corresponding subject.

[0127] It should be noted that when obtaining the training set, the test question text can be cleaned and the cleaned text data can be used as test question samples. The cleaning method can be found in the relevant description of obtaining the text data of the test questions to be tested in the previous embodiment. The principle is similar and will not be repeated here.

[0128] Step 302: Use the first training set to perform unsupervised training on the BERT model to obtain a deep network model.

[0129] In this embodiment of the disclosure, the BERT model can be trained in an unsupervised manner (Further_Pretrain) using the acquired first training set to obtain a deep network model, so that the pretrained model can better adapt to the semantic information of test texts in educational scenarios. The data for unsupervised training can be understood as test data without knowledge points, that is, only the first test sample is used.

[0130] The BERT model is a bidirectional language model that can adaptively adjust the feature representation of words according to their context. Unsupervised pre-training of the BERT model using the first training set and its application in the training process of the knowledge point prediction model can improve the robustness of the knowledge point prediction model and thus improve the accuracy of knowledge point prediction.

[0131] Considering the characteristics of different subjects, such as physics and mathematics exam questions which may contain formulas, and for LaTeX formula symbols that were not cleaned during data cleaning, these symbols can be added to the user dictionary and then to the BERT model. These LaTeX formula symbols serve as additional text markers for the question stem, allowing the BERT model to treat them as a whole during training, preventing them from being segmented by the word segmenter and losing their original physical meaning. For example, the LaTeX identifier for a space is "\alpha". Adding all LaTeX formula symbols to the pre-trained model prevents the model from splitting "\alpha" into other identifiers during training, thus avoiding the prediction of "\alpha" as a single word. For instance, it prevents "\alpha" from being split into two words, "\al" and "pha", further improving the model's prediction accuracy.

[0132] Step 303: Train the constructed knowledge point prediction model using the second training set to obtain the knowledge point prediction model, wherein the network parameters of the word embedding layer in the encoding module of the knowledge point prediction model are initialized to the network parameters of the deep network model.

[0133] In this embodiment of the disclosure, the structure of the constructed knowledge point prediction model is as follows: Figure 2 As shown, after training the deep network model, the encoding module of the knowledge point prediction model directly loads the network parameters of the deep network model as the initial network parameters. Then, supervised training (Fine-tuning) is performed on the knowledge point prediction model using the second training set. During training, the second test sample in the second training set is used as the feature, and the knowledge point label corresponding to the second test sample is used as the ground truth label. The loss is calculated based on the output result in the output module and the ground truth label. When the loss value reaches the preset value, the model is considered to have converged, the training is complete, and the trained knowledge point prediction model is obtained, which can be used in the knowledge point prediction method of the aforementioned embodiment.

[0134] In calculating the loss value, the cross-entropy loss function (Focal Loss) can be used to calculate the loss value, and the network parameters of the multi-label model can be updated based on the loss value to solve the problem of imbalance of knowledge point categories.

[0135] In this embodiment of the disclosure, when training the knowledge point prediction model, a first training set and a second training set are first obtained. The first training set includes a first test question sample, and the second training set includes a second test question sample and the corresponding knowledge point labels. The BERT model is unsupervised trained using the first training set to obtain a deep network model. The knowledge point prediction model is then trained using the second training set to obtain a trained knowledge point prediction model. The network parameters of the token embedding layer in the encoding module of the knowledge point prediction model are initialized to the network parameters of the deep network model. Thus, by adopting a two-stage training mode of Further_Pretrain + Fine_tuning, the knowledge point prediction model is obtained, which helps to improve the robustness and accuracy of the model.

[0136] Figure 5 A flowchart illustrating a training method for a knowledge point prediction model according to another exemplary embodiment of the present disclosure is shown, such as... Figure 5 As shown, the second training set can be obtained through the following steps:

[0137] Step 401: Obtain the question texts for multiple test questions and the corresponding knowledge points for each test question.

[0138] The test question text can be text data after the test question text information has been cleaned. Cleaning can include, but is not limited to, cleaning some formatting in HTML (such as spaces, & symbols), non-standard LaTeX symbols, etc.

[0139] Generally, test questions can be divided into two main categories: simple questions and complex questions. Complex questions include a stem and sub-question information. Sub-question information may be further conditions provided in the question section of a short-answer question, and usually also includes some information related to knowledge points. Therefore, for complex questions, the stem and sub-question information can be combined to form the text information of the test question. Simple questions do not contain sub-question information; therefore, for simple questions, the stem information can be used as the text information of the test question. The text information of the test question is then cleaned to obtain the test question text.

[0140] Step 402: For each question, count the number of questions associated with each knowledge point corresponding to each question.

[0141] A single test question may contain a limited number of knowledge points, but the number of test questions associated with the same knowledge point may be numerous. For a given knowledge point, the more test questions associated with that knowledge point, the greater the likelihood that the question corresponds to that knowledge point. If each knowledge point associated with a test question has a large number of associated test questions, then the test question is more representative. Therefore, in this embodiment of the disclosure, the number of test questions associated with each knowledge point corresponding to that test question can be counted.

[0142] Step 403: Based on the number of test questions associated with each knowledge point, determine candidate test questions and the corresponding knowledge point tags from the plurality of test questions.

[0143] In this embodiment of the disclosure, after counting the number of questions associated with each knowledge point of each question, candidate questions can be selected from multiple questions based on the number of questions associated with each knowledge point, and the candidate questions are labeled with the corresponding knowledge point tags.

[0144] As an optional implementation, when screening test questions and knowledge points, for any test question, if at least one knowledge point in all knowledge points corresponding to the first test question has a number of test questions not less than a preset value, then the first test question is determined as a candidate test question, and the knowledge points with a number of associated test questions not less than the preset value are determined as the knowledge point tags corresponding to the candidate test question; if the number of test questions associated with each knowledge point in all knowledge points corresponding to the second test question is less than the preset value, then the second test question is eliminated; wherein, the first test question and the second test question are any one of the plurality of test questions.

[0145] The preset value can be set in advance according to actual needs, such as setting the preset value to 200.

[0146] For example, assuming a preset value of 150, Question 1 contains two knowledge points, A and B. Statistically, knowledge point A is associated with 133 questions, and knowledge point B is associated with 168 questions. Question 2 contains knowledge points A and C; knowledge point A is associated with 141 questions, and knowledge point C is associated with 88 questions. Since the number of questions associated with knowledge point B in Question 1 is greater than the preset value, Question 1 is identified as a candidate question. However, since the number of questions associated with knowledge point A in Question 1 is less than the preset value, knowledge point A is deleted, and only knowledge point B is identified as the corresponding knowledge point label for Question 1. Since the number of questions associated with both knowledge points A and C in Question 2 is less than the preset value, Question 2 is removed from the question set and will not be considered as a candidate question for constructing the second training set.

[0147] This step allows you to filter out all candidate questions that meet the criteria from multiple test questions, as well as the corresponding knowledge point tags for each candidate question.

[0148] Step 404: Process the text of the candidate test questions to obtain a second test question sample whose text length is no greater than the target length.

[0149] The target length can be determined in different ways.

[0150] As an optional implementation method, the target length can be preset according to actual needs, such as setting the target length to 180, 200, etc.

[0151] As an optional implementation, the target length can be determined based on the length information of the candidate questions. Specifically, the length information of each candidate question can be calculated, where the length information can be represented by the number of words. Then, the target length is determined based on the length information of the candidate questions.

[0152] For example, the length information of candidate test questions can be sorted in ascending order, and the length information corresponding to the upper quartile can be determined as the target length. It is understood that the upper quartile refers to the degree of dispersion of skewed data when describing data using quartile statistical descriptive analysis. That is, when all data are arranged from smallest to largest, the number exactly in the lower 1 / 4 position is called the lower quartile (according to the percentage, i.e., the number in the 25th percentile), also known as the first quartile. The number in the upper 1 / 4 position is called the upper quartile (according to the percentage, i.e., the number in the 75th percentile), also known as the third quartile.

[0153] For example, the average length information of the candidate questions can be calculated, and the obtained average length can be used as the target length.

[0154] In this embodiment of the disclosure, for the selected candidate questions, the length information of the question text of each candidate question can be counted, and the question text with a length greater than the target length can be segmented to obtain the question text with a length less than or equal to the target length as the second question sample.

[0155] For example, if the length information of a candidate question is not greater than the target length, the question text of the candidate question will not be processed; if the length information of a candidate question is greater than the target length, the question text of the candidate question can be segmented, retaining at most the first two sentences and the last two sentences of the question stem, so that the length of the segmented question text is not greater than the target length.

[0156] Step 405: Using the second test question sample and the knowledge point tags corresponding to the second test question sample, construct the second training set.

[0157] In this embodiment of the disclosure, the second test question sample is a candidate test question whose length is not greater than the target length or a candidate test question after the test question text has been segmented. The knowledge point tags corresponding to these candidate test questions are also the knowledge point tags corresponding to the second test question sample. Then, the second training set is constructed by using the second test question sample and the knowledge point tags corresponding to the second test question sample.

[0158] In this embodiment of the disclosure, by acquiring the question texts of multiple test questions and the knowledge points corresponding to each test question, for each test question, the number of test questions associated with each knowledge point corresponding to each test question is counted, and based on the number of test questions associated with each knowledge point, candidate test questions and the knowledge point tags corresponding to the candidate test questions are determined from the multiple test questions. Then, the question texts of the candidate test questions are processed to obtain a second test question sample with a text length less than or equal to the target length. Using the second test question sample and the knowledge point tags corresponding to the second test question sample, a second training set is constructed. Thus, a training set that meets specific conditions can be constructed, providing data support for training the knowledge point prediction model.

[0159] An exemplary embodiment of this disclosure also provides a knowledge point prediction apparatus. Figure 6 A schematic block diagram of a knowledge point prediction apparatus according to an exemplary embodiment of the present disclosure is shown, such as Figure 6 As shown, the knowledge point prediction device 60 includes: an acquisition module 610, a first determination module 620, and a second determination module 630.

[0160] Among them, the acquisition module 610 is used to acquire the text data of the questions to be tested;

[0161] The first determining module 620 is used to determine the representation vector of the question to be tested by processing the text data using the self-attention module and the knowledge point attention module of the pre-trained knowledge point prediction model; to classify the knowledge points according to the representation vector using the output module of the knowledge point prediction model to obtain the classification result; and to determine the prediction probability of the question to be tested belonging to each of the preset knowledge points according to the classification result using the post-processing module of the knowledge point prediction model, and to determine the knowledge point number corresponding to the question to be tested according to the prediction probability.

[0162] The second determining module 630 is used to query the correspondence between preset numbers and knowledge points, and to determine the target knowledge point corresponding to the knowledge point number.

[0163] Optionally, the knowledge point prediction model further includes an encoding module and a mixing module; and wherein the first determining module 620 is further configured to:

[0164] The word vector representation of each word in the text data is determined by the encoding module;

[0165] Using the self-attention module, a contribution matrix based on the test item is determined according to the word vector representation of each word;

[0166] Using the knowledge point attention module, a contribution matrix based on knowledge points is determined according to the word vector representation of each word and the feature representation of each knowledge point;

[0167] Using the hybrid module, the representation vector of the question to be tested is determined based on the contribution matrix based on the test question and the contribution matrix based on the knowledge point.

[0168] Optionally, the first determining module 620 is further configured to:

[0169] A nonlinear transformation is performed on the word vector matrix formed by the word vector representation of each word to obtain the nonlinear matrix corresponding to the word vector matrix;

[0170] Perform a linear transformation on the nonlinear matrix to generate a transformation matrix;

[0171] The transformation matrix is ​​normalized along the word dimension to obtain a normalized matrix;

[0172] The contribution matrix based on the test items is determined by multiplying the normalized matrix and the transpose of the word vector matrix.

[0173] A knowledge point matrix is ​​determined based on the feature representation of each knowledge point, wherein the dimension of the feature representation is the same as the dimension of the word vector representation;

[0174] The contribution matrix based on knowledge points is determined by multiplying the knowledge point matrix by the transpose of the word vector matrix.

[0175] Optionally, the first determining module 620 is further configured to:

[0176] Obtain the first weight corresponding to the contribution matrix based on test questions, and the second weight corresponding to the contribution matrix based on knowledge points;

[0177] Based on the first weight and the second weight, the contribution matrix based on the test question and the contribution matrix based on the knowledge point are weighted and summed to obtain the test question representation matrix, wherein each row of the test question representation matrix represents the feature representation of a knowledge point for the test question to be tested;

[0178] The mean of each row of the test item representation matrix is ​​calculated to obtain the representation vector of the test item to be tested.

[0179] Optionally, the first determining module 620 is further configured to:

[0180] Using a linear transformation function, the dimension transformation of the contribution matrix based on test questions is performed to obtain a first transformation result, and the dimension transformation of the contribution matrix based on knowledge points is performed to obtain a second transformation result;

[0181] The self-attention weights corresponding to the self-attention module are determined by scaling the first transformation result using an activation function.

[0182] The activation function is used to scale the second transformation result to determine the knowledge point weights corresponding to the knowledge point attention module;

[0183] The first weight is determined based on the ratio of the self-attention weight to the knowledge point weight;

[0184] The second weight is determined based on the first weight, wherein the sum of the first weight and the second weight is 1.

[0185] Optionally, the first determining module 620 is further configured to:

[0186] The maximum prediction probability is determined from the predicted probabilities;

[0187] Based on the floor result of 1 and the maximum predicted probability, the number of knowledge points corresponding to the question to be tested is obtained;

[0188] Based on the predicted probabilities, the numbers corresponding to each knowledge point in all the knowledge points are sorted in descending order of predicted probabilities to obtain the sorting result;

[0189] The number of the knowledge points ranked first in the sorting results is determined and used as the knowledge point number corresponding to the question to be tested.

[0190] Optionally, the knowledge point prediction model is trained in the following manner:

[0191] Obtain a first training set and a second training set. The first training set includes a first test question sample, and the second training set includes a second test question sample and the knowledge point tags corresponding to the second test question sample.

[0192] The BERT model was trained unsupervised using the first training set to obtain a deep network model.

[0193] The knowledge point prediction model is trained using the second training set to obtain the knowledge point prediction model, wherein the network parameters of the word embedding layer in the encoding module of the knowledge point prediction model are initialized to the network parameters of the deep network model.

[0194] Optionally, the second training set can be obtained in the following way:

[0195] Obtain the text of multiple test questions and the corresponding knowledge points for each question;

[0196] For each question, count the number of questions associated with each knowledge point corresponding to each question;

[0197] Based on the number of test questions associated with each knowledge point, candidate test questions and corresponding knowledge point tags are determined from the plurality of test questions;

[0198] The candidate test questions are processed to obtain a second test question sample whose length is no greater than the target length, wherein the target length is determined based on the length information of the candidate test questions;

[0199] The second training set is constructed using the second test question sample and the corresponding knowledge point tags.

[0200] Optionally, the candidate test questions and the corresponding knowledge point tags are determined in the following way:

[0201] If, among all the knowledge points corresponding to the first question, there is at least one knowledge point associated with a number of questions not less than a preset value, then the first question is determined as the candidate question, and the knowledge points associated with a number of questions not less than the preset value are determined as the knowledge point tags corresponding to the candidate question, wherein the first question is any one of the multiple questions;

[0202] If the number of questions associated with each knowledge point in the second question is less than the preset value, then the second question will be removed.

[0203] The knowledge point prediction device provided in this disclosure can execute any knowledge point prediction method applicable to electronic devices provided in this disclosure, and has the corresponding functional modules and beneficial effects for executing the method. Content not described in detail in the device embodiments of this disclosure can be referred to the description in any method embodiment of this disclosure.

[0204] Exemplary embodiments of this disclosure also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program, when executed by the at least one processor, causing the electronic device to perform a knowledge point prediction method according to embodiments of this disclosure.

[0205] Exemplary embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a knowledge point prediction method according to embodiments of this disclosure.

[0206] Exemplary embodiments of this disclosure also provide a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a knowledge point prediction method according to embodiments of this disclosure.

[0207] refer to Figure 7 The present invention describes a structural block diagram of an electronic device 1100 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0208] like Figure 7As shown, the electronic device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. The RAM 1103 may also store various programs and data required for the operation of the device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0209] Multiple components in electronic device 1100 are connected to I / O interface 1105, including: input unit 1106, output unit 1107, storage unit 1108, and communication unit 1109. Input unit 1106 can be any type of device capable of inputting information to electronic device 1100. Input unit 1106 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 1107 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1108 may include, but is not limited to, disk and optical disk. Communication unit 1109 allows electronic device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0210] The computing unit 1101 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above. For example, in some embodiments, the knowledge point prediction method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1100 via ROM 1102 and / or communication unit 1109. In some embodiments, the computing unit 1101 can be configured to perform the knowledge point prediction method by any other suitable means (e.g., by means of firmware).

[0211] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0212] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0213] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0214] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0215] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0216] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

Claims

1. A knowledge point prediction method, wherein, The method includes: Obtain the text data of the questions to be tested; The representation vector of the question to be tested is determined by processing the text data using the self-attention module and knowledge point attention module of the pre-trained knowledge point prediction model. Using the output module of the knowledge point prediction model, knowledge points are classified according to the representation vector to obtain the classification result; Using the post-processing module of the knowledge point prediction model, the predicted probability of the question to be tested belonging to each of the preset knowledge points is determined according to the classification result, and the knowledge point number corresponding to the question to be tested is determined according to the predicted probability. Query the pre-defined correspondence between numbers and knowledge points to determine the target knowledge point corresponding to the knowledge point number; The knowledge point prediction model further includes an encoding module and a hybrid module; the method further includes: The word vector representation of each word in the text data is determined by the encoding module; Furthermore, the step of using the self-attention module and knowledge point attention module of the pre-trained knowledge point prediction model to process the text data and determine the representation vector of the question to be tested includes: Using the self-attention module, a contribution matrix based on the test item is determined according to the word vector representation of each word; Using the knowledge point attention module, a contribution matrix based on knowledge points is determined according to the word vector representation of each word and the feature representation of each knowledge point; Using the hybrid module, the representation vector of the question to be tested is determined based on the contribution matrix based on the test question and the contribution matrix based on the knowledge point.

2. The knowledge point prediction method as described in claim 1, wherein, The step of determining the contribution matrix based on the question according to the word vector representation of each word includes: A nonlinear transformation is performed on the word vector matrix formed by the word vector representation of each word to obtain the nonlinear matrix corresponding to the word vector matrix; Perform a linear transformation on the nonlinear matrix to generate a transformation matrix; The transformation matrix is ​​normalized along the word dimension to obtain a normalized matrix; The contribution matrix based on the test items is determined by multiplying the normalized matrix and the transpose of the word vector matrix. Furthermore, the step of determining the contribution matrix based on knowledge points according to the word vector representation of each word and the feature representation of each knowledge point includes: A knowledge point matrix is ​​determined based on the feature representation of each knowledge point, wherein the dimension of the feature representation is the same as the dimension of the word vector representation; The contribution matrix based on knowledge points is determined by multiplying the knowledge point matrix by the transpose of the word vector matrix.

3. The knowledge point prediction method as described in claim 1, wherein, The step of determining the representation vector of the question to be tested based on the contribution matrix based on the test question and the contribution matrix based on the knowledge point includes: Obtain the first weight corresponding to the contribution matrix based on test questions, and the second weight corresponding to the contribution matrix based on knowledge points; Based on the first weight and the second weight, the contribution matrix based on the test question and the contribution matrix based on the knowledge point are weighted and summed to obtain the test question representation matrix, wherein each row of the test question representation matrix represents the feature representation of a knowledge point for the test question to be tested; The mean of each row of the test item representation matrix is ​​calculated to obtain the representation vector of the test item to be tested.

4. The knowledge point prediction method as described in claim 3, wherein, The step of obtaining the first weight corresponding to the contribution matrix based on test questions and the second weight corresponding to the contribution matrix based on knowledge points includes: Using a linear transformation function, the dimension transformation of the contribution matrix based on test questions is performed to obtain a first transformation result, and the dimension transformation of the contribution matrix based on knowledge points is performed to obtain a second transformation result; The self-attention weights corresponding to the self-attention module are determined by scaling the first transformation result using an activation function. The activation function is used to scale the second transformation result to determine the knowledge point weights corresponding to the knowledge point attention module; The first weight is determined based on the ratio of the self-attention weight to the knowledge point weight; The second weight is determined based on the first weight, wherein the sum of the first weight and the second weight is 1.

5. The knowledge point prediction method as described in claim 1, wherein, The step of determining the knowledge point number corresponding to the question to be tested based on the predicted probability includes: The maximum prediction probability is determined from the predicted probabilities; Based on the floor result of 1 and the maximum predicted probability, the number of knowledge points corresponding to the knowledge points of the question to be tested is obtained; Based on the predicted probabilities, the numbers corresponding to each knowledge point in all the knowledge points are sorted in descending order of predicted probabilities to obtain the sorting result; The number of the knowledge points ranked first in the sorting results is determined and used as the knowledge point number corresponding to the question to be tested.

6. The knowledge point prediction method as described in any one of claims 1-5, wherein, The knowledge point prediction model is trained through the following steps: Obtain a first training set and a second training set. The first training set includes a first test question sample, and the second training set includes a second test question sample and the knowledge point tags corresponding to the second test question sample. The BERT model was trained unsupervised using the first training set to obtain a deep network model. The knowledge point prediction model is trained using the second training set to obtain the knowledge point prediction model, wherein the network parameters of the word embedding layer in the encoding module of the knowledge point prediction model are initialized to the network parameters of the deep network model. Furthermore, the second training set is obtained through the following method: Obtain the text of multiple test questions and the corresponding knowledge points for each question; For each question, count the number of questions associated with each knowledge point corresponding to each question; Based on the number of test questions associated with each knowledge point, candidate test questions and corresponding knowledge point tags are determined from the plurality of test questions; The candidate test questions are processed to obtain a second test question sample whose length is no greater than the target length, wherein the target length is determined based on the length information of the candidate test questions; The second training set is constructed using the second test question sample and the corresponding knowledge point tags.

7. The knowledge point prediction method as described in claim 6, wherein, The step of determining candidate questions and their corresponding knowledge point tags from the plurality of questions based on the number of questions associated with each knowledge point includes: If, among all the knowledge points corresponding to the first question, there is at least one knowledge point associated with a number of questions not less than a preset value, then the first question is determined as the candidate question, and the knowledge points associated with a number of questions not less than the preset value are determined as the knowledge point tags corresponding to the candidate question, wherein the first question is any one of the multiple questions; If the number of questions associated with each knowledge point in the second question is less than the preset value, then the second question will be removed.

8. A knowledge point prediction device, wherein, The device includes: The acquisition module is used to acquire the text data of the questions to be tested. The first determining module is used to determine the representation vector of the question to be tested by processing the text data using the self-attention module and the knowledge point attention module of the pre-trained knowledge point prediction model; to classify the knowledge points according to the representation vector using the output module of the knowledge point prediction model to obtain the classification result; and to determine the prediction probability of the question to be tested belonging to each of the preset knowledge points according to the classification result using the post-processing module of the knowledge point prediction model, and to determine the knowledge point number corresponding to the question to be tested according to the prediction probability. The second determining module is used to query the correspondence between preset numbers and knowledge points, and to determine the target knowledge point corresponding to the knowledge point number. The knowledge point prediction model further includes an encoding module and a hybrid module; The first determining module is further configured to: The word vector representation of each word in the text data is determined by the encoding module; Using the self-attention module, a contribution matrix based on the test item is determined according to the word vector representation of each word; Using the knowledge point attention module, a contribution matrix based on knowledge points is determined according to the word vector representation of each word and the feature representation of each knowledge point; Using the hybrid module, the representation vector of the question to be tested is determined based on the contribution matrix based on the test question and the contribution matrix based on the knowledge point.

9. An electronic device, comprising: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the knowledge point prediction method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the knowledge point prediction method according to any one of claims 1-7.