A training method and a question knowledge point determination method and device
By adding knowledge point tags to the word segmentation results of the question samples and training a neural network model, the problem of excessive noise in the prediction of knowledge point tags by generative network models is solved, achieving higher prediction accuracy and computational efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-29
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, generative network models suffer from problems such as numerous noisy labels and low accuracy when predicting knowledge point labels for questions. Furthermore, they have low computational efficiency and poor generalization ability in multi-knowledge point label classification tasks.
By adding knowledge point tags as characters to the word segmentation results of the question samples, a neural network model is trained, and the objective function is optimized through a loss function to reduce noisy tags. This process is then used to train the neural network model and determine the knowledge point prediction model.
It reduces the probability of noise labels appearing, improves the accuracy of knowledge point prediction, and can directly predict the knowledge points of the question based on knowledge point labels during the reasoning stage, thus improving the accuracy of prediction.
Smart Images

Figure CN116306795B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of education, and particularly relates to a training method and a question knowledge point determination method and device. BACKGROUND
[0002] With the rapid development of Internet, big data and other technologies, various kinds of question doing learning products provide students with a convenient question doing platform, and users can browse test questions and do the questions. In the process of cataloging questions into a question bank, it is often necessary to label the questions with knowledge points for the convenience of users to understand the questions.
[0003] In related technologies, the prediction task of knowledge point labels is often regarded as a classification task, and a classification model is used to predict the knowledge point labels of questions. A generative model can also be used to extract shallow and deep data features of question texts, obtain topic distribution through a topic model, fuse features through an attention mechanism, and generate knowledge point labels based on the finally obtained feature vectors. SUMMARY
[0004] According to an aspect of the present disclosure, a training method is provided, comprising:
[0005] Adding the knowledge point label of the question sample in the form of a word to the segmentation result of the question sample to obtain a segmentation result of the training sample;
[0006] Inputting the segmentation result of the training sample into a neural network model to obtain a predicted knowledge point of the question sample;
[0007] If the neural network model does not satisfy the convergence condition, updating the neural network model based on the predicted knowledge point and the knowledge point label;
[0008] If the neural network model satisfies the convergence condition, determining the neural network model as a knowledge point prediction model.
[0009] According to another aspect of the present disclosure, a question knowledge point determination method is provided, comprising:
[0010] Inputting the question information of the question into the knowledge point prediction model to obtain a predicted knowledge point at a current time step, the knowledge point prediction model being trained by the method described in the example embodiments of the present disclosure;
[0011] Incrementally updating the question information based on the predicted knowledge point at the current time step;
[0012] Determining the knowledge point of the question based on the predicted knowledge points at multiple time steps.
[0013] According to another aspect of the present disclosure, a training device is provided, comprising:
[0014] The training module is configured to add the knowledge point label of the question sample in the form of a knowledge point label word to a segmentation result of the question sample to obtain a segmentation result of a training sample;
[0015] The training module is further configured to input the segmentation result of the training sample into the neural network model to obtain a predicted knowledge point of the question sample.
[0016] The updating module is configured to update the neural network model based on the predicted knowledge point and the knowledge point label if the neural network model does not satisfy the convergence condition.
[0017] The determining module is configured to determine the neural network model as the knowledge point prediction model if the neural network model satisfies the convergence condition.
[0018] According to another aspect of the present disclosure, a question knowledge point determination apparatus is provided, comprising:
[0019] The prediction module is configured to input question information of the question into the knowledge point prediction model to obtain a predicted knowledge point at a current time step, the knowledge point prediction model being trained by the apparatus described in the example embodiments of the present disclosure.
[0020] The updating module is configured to incrementally update the question information based on the predicted knowledge point at the current time step.
[0021] The determining module is configured to determine the knowledge point of the question based on the predicted knowledge points at multiple time steps.
[0022] According to another aspect of the present disclosure, an electronic device is provided, comprising:
[0023] a processor; and
[0024] a memory storing a program;
[0025] The program includes instructions that, when executed by the processor, cause the processor to perform the method described in the example embodiments of the present disclosure.
[0026] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to perform the method described in the example embodiments of the present disclosure.
[0027] One or more technical solutions provided in the example embodiments of the present disclosure can add the knowledge point labels of the question sample in the form of word to the segmentation result of the question sample, to obtain the segmentation result of the training sample. At this time, the knowledge point labels included in the segmentation result of the training sample exist in the form of word, and therefore, when the segmentation result of the training sample is input into the neural network model, the predicted knowledge point of the question sample obtained is a complete knowledge point, thereby reducing the probability of occurrence of noise labels. On this basis, the knowledge point prediction model trained in the inference stage can directly predict the knowledge point of the question in the form of knowledge point label, solving the noise problem existing in the generative network model for predicting the knowledge point in the prior art, and improving the prediction accuracy of the knowledge point. BRIEF DESCRIPTION OF DRAWINGS
[0028] In the following description of the example embodiments in conjunction with the drawings, more details, features and advantages of the present disclosure are disclosed, in which:
[0029] Figure 1 A flowchart of a training method of the example embodiments of the present disclosure is shown;
[0030] Figure 2 A flowchart of a question knowledge point determination method of the example embodiments of the present disclosure is shown;
[0031] Figure 3 A schematic block diagram of the modules of a training device of the example embodiments of the present disclosure is shown;
[0032] Figure 4 A schematic block diagram of the modules of a question knowledge point determination device of the example embodiments of the present disclosure is shown;
[0033] Figure 5 A schematic block diagram of a chip of the example embodiments of the present disclosure is shown;
[0034] Figure 6 A structural block diagram of an example electronic device capable of implementing the embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0035] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0036] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this regard.
[0037] The term "comprises" and variations thereof herein, do not necessarily signify the inclusion of the foregoing listed items, but rather that the list are non-exhaustive. The term "based on" means "based, at least in part, on." The term "one embodiment" is used herein to refer to at least one embodiment. The term "another embodiment" is used herein to refer to at least one additional embodiment. The term "some embodiments" is used herein to refer to at least one embodiment. Relative terms such as "first", "second" and the like can be used solely to distinguish one entity or action from another, without necessarily giving suggestions of the relative importance or constituting an implied or explicit claim to priority. The term "plurality" is used herein to refer to at least two.
[0038] It should be noted that the terms "one", "multiple", mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless the context clearly indicates otherwise, it should be understood as "one or more".
[0039] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0040] Before introducing the embodiments of the present disclosure, first, the related terms involved in the embodiments of the present disclosure are explained as follows:
[0041] Byte pair encoding (BPE) is a pruning method (combining words with sequential dependencies and high word frequency overlap) for a 1-to-n ngram tokenization method. For example, in the input sentence, abc: 50, abcd: 49, then when 49 / 50>Threshold, the word abc can be directly deleted.
[0042] Generative Pre-trained Transformer 2 (GPT-2) is a very powerful pre-trained language model proposed by Open AI, which uses more network parameters and larger data sets than GPT-1. The core idea of GPT-2 can be summarized as follows: any supervised task is a subset of the language model, when the capacity of the model is very large and the amount of data is sufficient, only by training the language model learning can complete other supervised learning tasks. That is, the learning goal of GPT-2 is to use an unsupervised pre-training model to do a supervised task.
[0043] In the related art, in the new question entry application scenario, a teaching and research teacher with certain teaching experience is required to label the knowledge points of the question, but it is a very time-consuming work to accurately and comprehensively select the correct knowledge point label from a large number of standard knowledge point label candidates. At present, the pretest question knowledge point label mainly uses a classification model. For multi-label knowledge points, the classification task of multi-knowledge point labels can be converted into multiple binary classification tasks, but for the case where the question knowledge point candidate set is large, the calculation efficiency is low. The classification task of multi-knowledge point labels can also be directly modeled, but when the number of knowledge point labels contained in the question is uncertain, the number of knowledge point labels contained in the question cannot be accurately determined, and it is difficult to achieve unified processing through the threshold setting method. Different thresholds need to be set for different knowledge point labels, so the generalization ability of the multi-knowledge point label classification task is poor, and the correlation between multiple knowledge point labels is not considered.
[0044] In addition, the prior art can also extract shallow and deep data features of the question text through a generative model, obtain a topic distribution through a topic model, fuse features through an attention mechanism, and generate a knowledge point label based on the finally obtained feature vector. However, when the generative model extracts features from the question text, the model used is a BERT pre-trained language representation model. The model structure and pre-training task of the pre-trained language representation model (Bidirectional Encoder Representation from Transformers, BERT) itself determine that the generative model trained by the pre-trained language representation model will generate a large number of noise labels in the knowledge point label generated in the word-level coding mode when predicting the knowledge point label of the question, and may not generate a standard knowledge point label. For example, for a knowledge point label such as "calculation of static friction force", the generated knowledge point sequence may be "static friction force", "static friction force calculation", "friction force calculation", and various permutations, which makes the accuracy of the knowledge point label predicted by the generative model not high, and further processing is required before use, causing unnecessary trouble.
[0045] To solve the above problems, the exemplary embodiments of the present disclosure provide a training method, which can make the obtained knowledge point prediction model directly predict the knowledge points of the question in the inference stage in the unit of knowledge point label, thereby reducing the probability of occurrence of noise labels and improving the prediction accuracy of knowledge points. It should be understood that the method of the exemplary embodiments of the present disclosure can be executed by an electronic device or a chip applied to an electronic device.
[0046] Exemplarily, the electronic device of the exemplary embodiments of the present disclosure can be executed by an electronic device having a display function, and the electronic device can be a terminal such as a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, etc.
[0047] Figure 1 A flowchart of the training method of the exemplary embodiments of the present disclosure is shown. As shown in the flowchart, the training method of the exemplary embodiments of the present disclosure can include: Figure 1
[0048] Step 101: Adding the knowledge point label of the question sample in the form of a word to the segmentation result of the question sample to obtain the segmentation result of the training sample.
[0049] In actual applications, the question sample can be cleaned to obtain a standard natural language sequence text, and then the question sample can be segmented.
[0050] The above cleaning method can include cleaning in Hyper Text Markup Language (HTML) format, cleaning of Latex strings, etc.
[0051] The above segmentation method can be a segmentation method based on string matching, a segmentation method based on understanding, or a segmentation method based on statistics. For example, when the segmentation method based on string matching is used, the question sample can be segmented by using Byte Pair Encoding (BPE), and the segmentation result can be saved in the form of a vocab.txt file. The vocab.txt file contains the "words" contained in the question sample after segmentation and the corresponding numbers. On this basis, each knowledge point label pair can be added to the vocab.txt file as a word, and the vocab.txt file can be numbered. As can be seen, the segmentation result of the training sample includes the segmentation result of the question sample and the knowledge point label of the question sample.
[0052] Step 102: Inputting the segmentation result of the training sample into the neural network model to obtain the predicted knowledge point of the question sample.
[0053] Since the segmentation result of the training sample is in the form of a knowledge point label as a word, the segmentation result of the training sample is obtained by adding the knowledge point label of the question sample in the segmentation result of the question sample. Therefore, inputting the segmentation result of the training sample into the neural network model can reduce the probability of the neural network model outputting noise labels, thereby improving the accuracy of the knowledge point label of the question sample.
[0054] In actual applications, the disclosed example embodiments can input the word segmentation result of the training sample into the neural network model in a certain format. For the word segmentation result of each training sample, the specific format is as follows:
[0055] [CLS] SENTENCE [SEP] LABEL1 [SPACE] LABEL2 [SPACE] LABEL3 [SAPCE] … [SAPCE] LABELn [SEP]
[0056] wherein [CLS] represents the start of the word segmentation result of the training sample, SENTENCE represents the word segmentation result of the question sample, [SEP] represents the separator between the word segmentation result of the question sample and the knowledge point label, and the end of the knowledge point label, [SPACE] represents the separator between the knowledge point labels, LABEL1 represents the first knowledge point label, LABEL2 represents the second knowledge point label, LABEL3 represents the third knowledge point label, and LABELn represents the nth knowledge point label. Based on this, the predicted knowledge point of the question sample can be the predicted knowledge point output by the neural network model based on the word segmentation result of the question sample and the knowledge point label of the question sample.
[0057] Step 103: If the neural network model does not meet the convergence condition, updating the neural network model based on the predicted knowledge point and the knowledge point label.
[0058] The disclosed example embodiments can use the loss function as the objective function for optimizing the neural network model. The loss function can be used to represent the gap between the predicted knowledge point of the question sample and the knowledge point label of the question sample. The training process of the neural network model is also the process of minimizing the loss function.
[0059] The convergence condition can include that the loss value of the loss function is less than or equal to a preset threshold. The convergence condition can also include that the loss value of the loss function is stable, which is determined according to actual application scenarios and is not limited herein. If the neural network model does not meet the convergence condition, it means that the generalization ability of the currently trained neural network model is poor and cannot accurately predict the knowledge point of the question. Further training is needed to meet the prediction accuracy of the question knowledge point.
[0060] Step 104: If the neural network model meets the convergence condition, determining the neural network model as a knowledge point prediction model.
[0061] When the neural network model meets the convergence condition, it means that the currently trained neural network model has strong generalization ability and can accurately predict the knowledge point of the question. At this time, the currently trained neural network model can be determined as a knowledge point prediction model.
[0062] It can be seen that the knowledge point label included in the word segmentation result of the training sample in the method of the example embodiment of the present disclosure can exist in the form of a word, the word segmentation result of the training sample is input into the neural network model, and the predicted knowledge point of the question sample obtained is a complete knowledge point, thereby reducing the probability of occurrence of noise labels. On this basis, the knowledge point prediction model trained can directly predict the knowledge point of the question in the unit of the knowledge point label in the inference stage, solving the noise problem existing in the generative network model for predicting the knowledge point in the prior art, and improving the prediction accuracy of the knowledge point.
[0063] In actual application, the knowledge point labels of the question sample are various in type, and the number of occurrences of each knowledge point label in different question samples is different. In the case of sufficient question samples, the knowledge point labels with a small number of occurrences have little contribution to the construction of the to-be-output knowledge point of the knowledge point prediction model. Therefore, the example embodiment of the present disclosure can eliminate the knowledge point labels with a small number of occurrences. Based on this, the method of the example embodiment of the present disclosure can further include:
[0064] The number of occurrences of the same initial knowledge point in different question samples is counted, and if the number of occurrences of the initial knowledge point meets the screening condition, the initial knowledge point is determined as a knowledge point label.
[0065] The knowledge point labels screened here can be considered as knowledge point labels with a large number of occurrences. In this case, the predicted knowledge point predicted by the neural network model is one of the multiple knowledge point labels with a large number of occurrences, and is not the initial knowledge point with a small number of occurrences, thereby reducing the amount of calculation while ensuring the accuracy of the knowledge point prediction and avoiding unnecessary interference of the knowledge point.
[0066] In actual application, different question samples can contain the same knowledge point or different knowledge points, and the knowledge point contained in the same question sample can be one or multiple. Based on this, the example embodiment of the present disclosure can first count the knowledge points of different question samples, which are referred to as initial knowledge points. For each initial knowledge point, the number of occurrences of the initial knowledge point in different question samples can be counted. Here, the number of occurrences of the same initial knowledge point in different question samples can be defined as the number of occurrences of the initial knowledge point, and then it is judged whether the number of occurrences of the initial knowledge point meets the screening condition.
[0067] For example, the above screening condition can be used to eliminate the initial knowledge points with a small number of occurrences, and the specific screening condition can be determined according to actual needs. At this time, the screening condition can include that the number of occurrences of the initial knowledge point is greater than or equal to a preset number of occurrences, and the preset number of occurrences can be determined according to actual needs, which is not limited here.
[0068] For example, the example embodiments of the present disclosure can be sorted in descending order of the number of occurrences of the initial knowledge points, so that the initial knowledge points are sorted in descending order of the number of occurrences. At this time, the screening condition can include the number of occurrences of the top n initial knowledge points with the largest number of occurrences, and the knowledge point label of the knowledge point prediction model can be the top n initial knowledge points with the largest number of occurrences.
[0069] For example, if the number of knowledge point labels of the question sample is n, the number of the knowledge point label with the highest frequency is 1, which is recorded as the first knowledge point; the number of the knowledge point label with the second highest frequency is 2, which is recorded as the second knowledge point; and so on. The number of the knowledge point label with the lowest frequency is n, which is recorded as the n-th knowledge point.
[0070] When the initial knowledge points are screened by the screening condition to obtain the knowledge point label, the knowledge point label can also be identified, for example, the knowledge point label can be identified in descending order of the number of occurrences of the knowledge point label, which facilitates machine learning. When the knowledge point label is added to the segmentation result of the question sample in the form of a word to obtain the segmentation result of the training sample, the segmentation result of the training sample is input into the neural network model, and the identification probability corresponding to the multiple knowledge point labels can be obtained, wherein the knowledge point label corresponding to the maximum identification probability is the predicted knowledge point.
[0071] In an optional manner, the method of the example embodiments of the present disclosure can also include determining a knowledge point co-occurrence matrix based on the knowledge point labels of the multiple question samples, and the knowledge point co-occurrence matrix is used to correct the knowledge point label predicted by the knowledge point prediction model in the reasoning stage.
[0072] After determining the knowledge point labels of the multiple question samples, the example embodiments of the present disclosure can create a knowledge point co-occurrence matrix, and the correlation information between different knowledge point labels can be embodied based on the knowledge point co-occurrence matrix. For the convenience of understanding, the example embodiments of the present disclosure show the knowledge point co-occurrence matrix in the form of a table, which actually exists in the form of a matrix, not in the form of a table. Table 1 shows the knowledge point co-occurrence matrix of the example embodiments of the present disclosure.
[0073] Table 1: Knowledge point co-occurrence matrix of the knowledge point prediction model
[0074]
[0075] As shown in Table 1, the size of the knowledge point co-occurrence matrix is (n+1) x (n+1), n represents the number of knowledge point labels, and +1 represents a special label [SPACE] when generating the knowledge point label. The exemplary embodiments of the present disclosure can use E[i][j] to represent the correlation information of the ith knowledge point and the jth knowledge point, and i and j are both integers greater than or equal to 1. E[i][j] can include the co-occurrence number of the ith knowledge point and the jth knowledge point in multiple question samples, or can include the co-occurrence probability of the ith knowledge point and the jth knowledge point in multiple question samples. The larger the value corresponding to E[i][j] is, the more likely the ith knowledge point and the jth knowledge point co-occur, and the greater the correlation between the two; on the contrary, the less likely the ith knowledge point and the jth knowledge point co-occur, and the smaller the correlation between the two. If the value corresponding to E[i][j] is 0, it means that the ith knowledge point and the jth knowledge point do not co-occur. The initialization value of the correlation information of this special label and other knowledge point labels is 0.
[0076] It can be seen that the method of the exemplary embodiments of the present disclosure can create a knowledge point co-occurrence matrix reflecting the correlation information between various knowledge point labels, which can be used to correct the predicted knowledge point label in the inference stage of the knowledge point prediction model, thereby improving the prediction accuracy of the question knowledge point.
[0077] In an optional manner, the neural network model of the exemplary embodiments of the present disclosure can include a plurality of layers of stacked transformer decoders, and each transformer decoder can include a masked self-attention module and a feedforward neural network. The masked self-attention module can mask the information of all words to the right of the current calculation position, which is used to determine the self-attention information based on the input information, and the feedforward neural network can be used to determine the predicted knowledge point based on the self-attention information.
[0078] If the transformer decoder is the first layer transformer decoder, the input information is the position encoding information of the tokenization result of the training sample. That is, the tokenization result of the training sample is essentially the word embedding vector of the training sample, which is position encoded, so that a position vector is also added to the word embedding vector, and then the position encoding information of the tokenization result of the training sample is obtained, and then the position encoding information of the tokenization result of the training sample is input into the first layer transformer decoder.
[0079] If the transformer decoder is the mth layer transformer decoder, the input information is the hidden state output by the (m-1)th layer transformer decoder, m is an integer greater than 1 and less than or equal to M, and M is the number of layers of the transformer decoder.
[0080] Exemplarily, the parameters of the configuration file of the neural network model of the exemplary embodiment of the present disclosure can be: the dictionary size (vocab_size) of the word segmentation result of the training sample is 21623, the maximum length of the position encoding is 512, the dimension of the hidden state is 768, a 6-layer stacked transform decoder is used, and the masked self-attention module is set to a multi-head attention module, the number of multi-heads of the multi-head attention module is 12, the initialization method of the model parameters adopts truncated normal distribution, and the initialization parameters such as weights can be 0.02, and the layer normalization parameter is 1e-05.
[0081] The exemplary embodiment of the present disclosure also provides a question knowledge point determination method, which can predict the knowledge points of a question in a knowledge point label unit in an inference stage based on the knowledge point prediction model trained by the exemplary embodiment of the present disclosure, thereby reducing the probability of occurrence of noise labels and improving the prediction accuracy of knowledge points. It should be understood that the method of the exemplary embodiment of the present disclosure can be executed by an electronic device or a chip applied to an electronic device. For specific content, please refer to the above text, which will not be repeated here.
[0082] Figure 2 A flowchart of a question knowledge point determination method according to an exemplary embodiment of the present disclosure is shown. As shown in Figure 2 The question knowledge point determination method according to an exemplary embodiment of the present disclosure can include:
[0083] Step 201: input the question information of a question into a knowledge point prediction model to obtain a predicted knowledge point at a current time step.
[0084] The question information of the question described above can at least include the word segmentation result of the question, and its acquisition process can refer to the acquisition process of the word segmentation result of the question sample in the above text, which will not be repeated here. The knowledge point prediction model is trained by the method described in the exemplary embodiment of the present disclosure, so that the knowledge point prediction model of the exemplary embodiment of the present disclosure can predict the knowledge points of a question in a knowledge point label unit in an inference stage, thereby reducing the probability of occurrence of noise labels and improving the prediction accuracy of knowledge points.
[0085] The predicted knowledge point at the current time step described above can be expressed in the form of conditional probability of the predicted knowledge point. Exemplarily, the conditional probability P(i) of the predicted knowledge point at the current time step can be expressed as:
[0086] P(i) = ∑ i logP(u i |u i-k ,u i-k+1 ,u i-k+2 …u i-1 ;θ);
[0087] wherein, θ is a model parameter, i represents an index of a predicted knowledge point, k represents a size of a sliding window, u i represents an i-th word contained in the topic information of the topic, u i-k represents an i-k-th word contained in the topic information of the topic, u i-k+1 represents an i-k+1-th word contained in the topic information of the topic, u i-k+2 represents an i-k+2-th word contained in the topic information of the topic, u i-1 represents an i-1-th word contained in the topic information of the topic.
[0088] Step 202: incrementally updating the topic information based on the predicted knowledge point of the current time step.
[0089] The process of incremental updating can include adding the predicted knowledge point of the current time step to the topic information of the topic, re-inputting the incrementally updated topic information of the topic to the knowledge point prediction model, and obtaining the predicted knowledge point of the next time step. If the current time step is the first time step, the incrementally updated topic information of the topic can include the word segmentation result of the topic and the predicted knowledge point of the current time step. If the current time step is a time step greater than or equal to 2, the incrementally updated topic information of the topic can include the word segmentation result of the topic, the predicted knowledge points of the historical time steps, and the predicted knowledge point of the current time step.
[0090] It can be seen that the example embodiments of the present disclosure can incrementally update the topic information based on the predicted knowledge point of the current time step, so that the knowledge point prediction model can determine the predicted knowledge point of the next time step based on the predicted knowledge point of the current time step, thereby ensuring that the predicted knowledge points have high relevance.
[0091] Step 203: determining the knowledge point of the topic based on the predicted knowledge points of multiple time steps.
[0092] Based on the iterative operations of the above steps 201 and 202, the example embodiments of the present disclosure can obtain the predicted knowledge points of multiple time steps, at this time, the knowledge point of the topic can be determined from the predicted knowledge points of multiple time steps.
[0093] It can be seen that the example embodiments of the present disclosure can predict the knowledge point of the topic in the inference stage in the unit of knowledge point label based on the knowledge point prediction model trained by the example embodiments of the present disclosure, thereby reducing the probability of occurrence of noise labels and improving the prediction accuracy of the knowledge point. At the same time, when the knowledge point of the topic is multiple, the high relevance between the multiple knowledge points can be ensured.
[0094] In an optional manner, the knowledge point prediction model of the example embodiment of the present disclosure can include a plurality of stacked transformer decoders, each of which includes a masked self-attention module and a feedforward neural network. The masked self-attention module is configured to determine self-attention information based on the input information, and the feedforward neural network is configured to determine the predicted knowledge point based on the self-attention information.
[0095] If the transformer decoder is the first transformer decoder, the input information is the position encoding information of the word segmentation result of the question information, and the detailed process can refer to the related content described above.
[0096] If the transformer decoder is the mth transformer decoder, the input information is the hidden state output by the (m-1)th transformer decoder, m is an integer greater than 1 and less than or equal to M, and M is the number of layers of the transformer decoder.
[0097] In an optional manner, the number of candidate knowledge points determined by the knowledge point prediction model at each time step can be zero, one or more. When the number of candidate knowledge points determined by the knowledge point prediction model at a certain time step is zero, it means that the conditional probability of all candidate knowledge points predicted by the knowledge point prediction model at the time step is less than the preset probability. When the number of candidate knowledge points determined by the knowledge point prediction model at a certain time step is greater than or equal to one, it means that there is a conditional probability greater than or equal to the preset probability in the conditional probability of all candidate knowledge points predicted by the knowledge point prediction model at the time step. If the number of candidate knowledge points with a conditional probability greater than or equal to the preset probability is more than one, the greater the conditional probability, the more likely the candidate knowledge point corresponding to the conditional probability is the predicted knowledge point at the current time step.
[0098] No matter how many candidate knowledge points are determined by the knowledge point prediction model at a certain time step, the predicted knowledge points determined at each time step have a mapping relationship, and at this time, the question information of the question title can include the question title and the predicted knowledge points having a mapping relationship at each historical time step, and the predicted knowledge points having a mapping relationship at each historical time step can constitute a predicted knowledge point sequence. It can be understood that the above historical time steps can include each time step before the current time step, and the above predicted knowledge point sequence can include the predicted knowledge points having a mapping relationship at each historical time step.
[0099] Based on this, inputting the question information of the question title into the knowledge point prediction model to obtain the predicted knowledge point at the current time step can include:
[0100] The title knowledge point and the predicted knowledge point sequence are input into the knowledge point prediction model to obtain a plurality of candidate knowledge points corresponding to the predicted knowledge point sequence at the current time step; a candidate knowledge point corresponding to the predicted knowledge point sequence at the current time step is selected from the plurality of candidate knowledge points, and the candidate knowledge point satisfies a candidate knowledge point screening condition; and a predicted knowledge point at the current time step is determined based on the candidate knowledge point corresponding to the predicted knowledge point sequence at the current time step, and each predicted knowledge point at the current time step and the corresponding predicted knowledge point sequence have a mapping relationship.
[0101] It should be understood that the above-mentioned candidate knowledge points are knowledge points predicted by the knowledge point prediction model based on the predicted knowledge point sequence at the current time step, at this time, a candidate knowledge point needs to be selected from the plurality of candidate knowledge points, and then combined with the predicted knowledge point sequence corresponding to the candidate knowledge point to determine whether the candidate knowledge point is the predicted knowledge point at the current time step.
[0102] For example, the above-mentioned candidate knowledge point screening condition can include that the conditional probability of the knowledge point is greater than or equal to a preset conditional probability, which can be determined according to actual needs, and is not limited here. At this time, the candidate knowledge point at the current time step can include the candidate knowledge point whose conditional probability is greater than or equal to the preset conditional probability, and the candidate knowledge point at the current time step can be one or more.
[0103] For example, the above-mentioned candidate knowledge point screening condition can include the conditional probability of the top p candidate knowledge points with the largest conditional probability after the candidate knowledge points are sorted in descending order of conditional probability, and the candidate knowledge point is the top p candidate knowledge points with the largest conditional probability, p is less than or equal to the total number of candidate knowledge points. P can be the number of candidate knowledge points corresponding to the predicted knowledge point sequence at the current time step, and the number of candidate knowledge points corresponding to the predicted knowledge point sequence at the current time step can be one or more, that is, p can be equal to 1, or an integer greater than or equal to 2, which is determined according to actual needs.
[0104] The method of the exemplary embodiment of the present disclosure can determine the predicted knowledge point at the current time step based on the candidate knowledge point corresponding to the predicted knowledge point sequence at the current time step after determining the candidate knowledge point corresponding to the predicted knowledge point sequence at the current time step. When each historical time step has a plurality of predicted knowledge points, the number of predicted knowledge point sequences, the number of predicted knowledge points at each historical time step, and the number of predicted knowledge points at the current time step are the same.
[0105] For example, the determining, by the example embodiment of the present disclosure, the predicted knowledge point of the current time step based on the candidate knowledge points corresponding to the predicted knowledge point sequence of the current time step can include: determining accumulated conditional probabilities of the candidate knowledge points corresponding to the predicted knowledge point sequence based on conditional probabilities of the candidate knowledge points and conditional probabilities of the corresponding predicted knowledge point sequence; and determining the predicted knowledge points of the current time step based on the accumulated conditional probabilities of the candidate knowledge points and the corresponding predicted knowledge point sequence.
[0106] For example, the determining, by the example embodiment of the present disclosure, the predicted knowledge point of the current time step based on the candidate knowledge points corresponding to the predicted knowledge point sequence of the current time step can include: determining accumulated conditional probabilities of the candidate knowledge points corresponding to the predicted knowledge point sequence based on conditional probabilities of the candidate knowledge points and conditional probabilities of the corresponding predicted knowledge point sequence; and determining the predicted knowledge points of the current time step based on the accumulated conditional probabilities of the candidate knowledge points and the corresponding predicted knowledge point sequence.
[0107] For example, assuming that the number of predicted knowledge points of each time step is 2, if the current time step is the second time step, the predicted knowledge points of the first time step include a first predicted knowledge point A and a second predicted knowledge point B, wherein the conditional probability of the first predicted knowledge point A is 0.4, and the conditional probability of the second predicted knowledge point B is 0.5.
[0108] Based on the first predicted knowledge point A and the target topic of the second time step, a third candidate knowledge point C and a fourth candidate knowledge point D can be determined, wherein the conditional probability of the third candidate knowledge point C is 0.9, and the conditional probability of the fourth candidate knowledge point D is 0.05. At this time, the accumulated conditional probability Q AC of the first predicted knowledge point A of the first time step to the third candidate knowledge point C of the second time step is 0.4*0.9=0.36, and the accumulated conditional probability Q AD of the second predicted knowledge point A of the first time step to the fourth candidate knowledge point D of the second time step is 0.4*0.05=0.02.
[0109] Based on the second predicted knowledge point B and the target topic of the second time step, a fourth candidate knowledge point D and a fifth candidate knowledge point E can be determined, wherein the conditional probability of the fourth candidate knowledge point D is 0.8, and the conditional probability of the fifth candidate knowledge point E is 0.15. At this time, the accumulated conditional probability Q BD of the second predicted knowledge point B of the first time step to the fourth candidate knowledge point D of the second time step is 0.5*0.8=0.4, and the accumulated conditional probability Q BE of the second predicted knowledge point B of the first time step to the fifth candidate knowledge point E of the second time step is 0.5*0.15=0.075.
[0110] On this basis, since Q BD > QAC >Q BE >Q AD , the prediction knowledge point of the second time step can be determined based on Q BD and Q AC The prediction knowledge point of the second time step is the fourth candidate knowledge point D and the third candidate knowledge point C, respectively.
[0111] Based on the above examples, it can also be seen that the number of prediction knowledge point sequences of the first time step is 2, including a first group of prediction knowledge point sequences composed of the first prediction knowledge point A and a second group of prediction knowledge point sequences composed of the second prediction knowledge point B; the number of prediction knowledge points of the first time step is 2, including the first prediction knowledge point A and the second prediction knowledge point B; the number of prediction knowledge points of the second time step is 2, including the fourth candidate knowledge point D and the third candidate knowledge point C. Moreover, the third candidate knowledge point C of the second time step has a mapping relationship with the corresponding first group of prediction knowledge point sequences, and the fourth candidate knowledge point D of the second time step has a mapping relationship with the corresponding second group of prediction knowledge point sequences.
[0112] It can be seen that the method of the exemplary embodiments of the present disclosure can determine the candidate knowledge point in the plurality of candidate knowledge points of the current time step that satisfies the candidate knowledge point screening condition as the candidate knowledge point of the current time step, and then determine the prediction knowledge point of the current time step based on the prediction knowledge point sequence of the candidate knowledge point of the current time step, eliminate the knowledge point with low possibility of the current time step, and determine the knowledge point label of the main topic with high importance to the main topic, thereby improving the prediction accuracy of the knowledge point of the main topic.
[0113] In an optional manner, the number of knowledge points contained in the question can be one or multiple. When the number of knowledge points contained in the main topic is uncertain, the method for determining the knowledge point of the main topic based on the prediction knowledge point of multiple time steps according to the exemplary embodiments of the present disclosure can include:
[0114] Based on each prediction knowledge point of the current time step and the corresponding prediction knowledge point sequence, determine a knowledge point screening parameter; if the knowledge point screening parameter satisfies a preset termination condition, generate the knowledge point of the main topic based on the prediction knowledge point sequence.
[0115] For example, when the number of predicted knowledge points of each time step is one, the predicted knowledge points of each historical time step that have a mapping relationship with the predicted knowledge point of the current time step are unique. In this case, the knowledge point screening parameter can include the number of knowledge points, which can be determined according to an actual application scenario. The preset termination condition can include that the number of knowledge points is greater than a preset number of knowledge points. If the predicted knowledge point of the current time step and the total number of predicted knowledge points included in the predicted knowledge point sequence are greater than the preset number of knowledge points, the target knowledge point is the predicted knowledge point included in the predicted knowledge point sequence. Therefore, the example embodiments of the present disclosure can limit the generation of the target knowledge point in the form of limiting the number of knowledge points.
[0116] For example, when the number of predicted knowledge points of each time step is one, the knowledge point screening parameter can include the accumulated conditional probability of the predicted knowledge point of the current time step and the predicted knowledge point sequence, which can be determined according to an actual application scenario. The accumulated conditional probability can be the multiplication result of the conditional probability of the predicted knowledge point of the current time step and the predicted knowledge point sequence. In this case, the preset termination condition can include that the accumulated conditional probability of the predicted knowledge point of the current time step and the predicted knowledge point sequence is less than a preset accumulated conditional probability, and the size of the preset accumulated conditional probability can be determined according to an actual application scenario. If the accumulated conditional probability of the predicted knowledge point of the current time step and the predicted knowledge point sequence is less than the preset accumulated conditional probability, the target knowledge point is the predicted knowledge point included in the predicted knowledge point sequence. Therefore, the example embodiments of the present disclosure can limit the generation of the target knowledge point in the form of limiting the accumulated conditional probability of the predicted knowledge point.
[0117] For example, when the number of predicted knowledge points of each time step is multiple, the example embodiments of the present disclosure can determine the knowledge point screening parameter based on each predicted knowledge point of the current time step and the corresponding predicted knowledge point sequence, which can include:
[0118] For example, when the number of predicted knowledge points of each time step is multiple, the example embodiments of the present disclosure can determine the knowledge point screening parameter based on each predicted knowledge point of the current time step and the corresponding predicted knowledge point sequence, which can include:
[0119] For example, when the number of predicted knowledge points of each time step is multiple, the example embodiments of the present disclosure can determine the knowledge point screening parameter based on each predicted knowledge point of the current time step and the corresponding predicted knowledge point sequence, which can include:
[0120] As can be seen from the foregoing examples, if the current time step is the second time step, the predicted knowledge points of the second time step include the fourth candidate knowledge point D and the third candidate knowledge point C, wherein the fourth candidate knowledge point D has a cumulative conditional probability Q BD = 0.4 with the corresponding predicted knowledge point sequence, and the third candidate knowledge point C has a cumulative conditional probability Q AC = 0.36 with the corresponding predicted knowledge point sequence, and thus the knowledge point screening parameter is Q BD . If Q BD is less than the preset cumulative conditional probability, the knowledge point of the target topic is generated based on the predicted knowledge point sequence, and the knowledge point of the target topic is the fourth candidate knowledge point D.
[0121] It can be seen that the method of the example embodiments of the present disclosure can determine the knowledge point of the target topic by setting the knowledge point screening parameter in the case where the number of knowledge points of the target topic is uncertain, and determine the knowledge point of the target topic based on each predicted knowledge point of the current time step and the corresponding predicted knowledge point sequence in the case where the number of knowledge points of the target topic is multiple, so that the obtained knowledge point can accurately describe the knowledge characteristic information of the target topic, improve the prediction accuracy of the knowledge point, and thus the knowledge point prediction model has strong generalization ability.
[0122] In an optional manner, if the current time step is the tth time step, t is an integer greater than or equal to 2, the method of the example embodiments of the present disclosure can further include: correcting the conditional probability of the predicted knowledge point of the current time step based on the co-occurrence probability of the predicted knowledge point of the (t-1)th time step and the predicted knowledge point of the tth time step.
[0123] For example, the example embodiments of the present disclosure can find the co-occurrence probability based on the predicted knowledge point of the (t-1)th time step and the predicted knowledge point of the tth time step from the knowledge point co-occurrence matrix. The knowledge point co-occurrence matrix can be a knowledge point co-occurrence matrix determined by the knowledge point prediction model based on the knowledge point labels of the plurality of topic samples in the training stage, and the specific determination method can be referred to the foregoing, which will not be described here in detail.
[0124] If the predicted knowledge point of the tth time step is the i th knowledge point in the knowledge point co-occurrence matrix, and the predicted knowledge point of the (t-1)th time step is the j th knowledge point in the knowledge point co-occurrence matrix, the co-occurrence probability (or the co-occurrence number) of the i th knowledge point and the j th knowledge point can be determined from the knowledge point co-occurrence matrix as E[j][i], and if the conditional probability of the predicted knowledge point of the current time step is P(i), the corrected conditional probability P(i)' of the predicted knowledge point of the current time step is:
[0125] P(i)' = P(i)*(E[j][i]+1)
[0126] In the method of the example embodiment of the present disclosure, the ratio of the conditional probability of the predicted knowledge points before and after the correction is negatively correlated with the co-occurrence probability. It can be understood that the predicted knowledge point sequence of the tth time step can include the predicted knowledge point of the tth time step and the predicted knowledge point sequence of the (t-1)th time step. Therefore, in the above formula, E[j][i] can also be used to represent the co-occurrence probability (or co-occurrence number) of the predicted knowledge point of the (t-1)th time step and the predicted knowledge point of the tth time step. The greater the value corresponding to E[j][i] is, the greater the co-occurrence probability (or co-occurrence number) of the predicted knowledge point sequence of the (t-1)th time step and the predicted knowledge point sequence of the tth time step is, the greater P(i)' is, and the smaller the ratio of P(i) and P(i)' is; on the contrary, the greater the ratio of P(i) and P(i)' is.
[0127] It can be seen that the method of the example embodiment of the present disclosure can correct the conditional probability of the predicted knowledge point of the current time step by using the knowledge point co-occurrence matrix determined by the knowledge point prediction model based on the knowledge point labels of the plurality of question samples in the training stage, thereby correcting the correlation between the predicted knowledge points of the two time steps before and after the correction, and further improving the prediction accuracy of the knowledge points of the question.
[0128] In the method of the example embodiment of the present disclosure, the predicted knowledge point of the (t-1)th time step and the predicted knowledge point of the tth time step have a mapping relationship. When determining E[j][i] of the i th knowledge point and the j th knowledge point from the knowledge point co-occurrence matrix, the predicted knowledge point of the (t-1)th time step which has a mapping relationship with the i th knowledge point of the tth time step can be determined as the j th knowledge point.
[0129] For example: if the number of predicted knowledge points of each time step is 1, the predicted knowledge point of the (t-1)th time step which has a mapping relationship with the predicted knowledge point of the tth time step is unique, at this time, the predicted knowledge point of the (t-1)th time step can be directly determined as the j th knowledge point.
[0130] For another example: if the number of predicted knowledge points of each time step is 2, the predicted knowledge points of the (t-1)th time step are knowledge point A and knowledge point B, the predicted knowledge points of the tth time step are knowledge point C determined based on knowledge point A and knowledge point D determined based on knowledge point B, at this time, when the i th knowledge point is knowledge point C, knowledge point A of the (t-1)th time step which has a mapping relationship with knowledge point C can be determined as the j th knowledge point; when the i th knowledge point is knowledge point D, knowledge point B of the (t-1)th time step which has a mapping relationship with knowledge point D can be determined as the j th knowledge point.
[0131] It can be seen that the method of the example embodiments of the present disclosure can quickly determine the predicted knowledge point corresponding to the predicted knowledge point of the current time step in the previous time step based on the mapping relationship between any predicted knowledge point in the current time step and the predicted knowledge point corresponding to each historical time step, thereby preparing for correcting the conditional probability of the predicted knowledge point of the current time step, and improving the determination efficiency of the target knowledge point.
[0132] One or more technical solutions provided in the example embodiments of the present disclosure can add the knowledge point label of the question sample in the form of a word to the word segmentation result of the question sample to obtain the word segmentation result of the training sample. At this time, the knowledge point label included in the word segmentation result of the training sample exists in the form of a word, so that when the word segmentation result of the training sample is input into the neural network model, the predicted knowledge point of the question sample obtained is a complete knowledge point, thereby reducing the probability of occurrence of noise labels. On this basis, the knowledge point prediction model trained in the inference stage can directly predict the knowledge point of the question in the unit of the knowledge point label, solving the noise problem existing in the generative network model for predicting the knowledge point in the prior art, and improving the prediction accuracy of the knowledge point.
[0133] The above mainly introduces the scheme provided by the embodiments of the present disclosure. It can be understood that, in order to realize the above functions, the electronic device contains the hardware structure and / or software module corresponding to each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is realized in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present disclosure.
[0134] The embodiments of the present disclosure can divide the functional units of the electronic device according to the above method examples, for example, each functional module can be divided corresponding to each function, or two or more functions can be integrated in one processing module. The above integrated module can be realized in the form of hardware or in the form of a software function module. It should be noted that the division of the module in the embodiments of the present disclosure is illustrative, and is only a logical function division. When actually implemented, there can be another division manner.
[0135] In the case of dividing each functional module corresponding to each function, the example embodiments of the present disclosure provide a training device, which can be an electronic device or a chip applied to an electronic device. Figure 3 A module schematic block diagram of the training device of the example embodiments of the present disclosure is shown. As shown in Figure 3As shown, the apparatus 300 comprises:
[0136] The training module 301 is configured to add the knowledge point label of the question sample in the form of a knowledge point label word to the segmentation result of the question sample, to obtain a segmentation result of a training sample.
[0137] The training module 301 is further configured to input the segmentation result of the training sample into the neural network model to obtain a predicted knowledge point of the question sample.
[0138] The updating module 302 is configured to update the neural network model based on the predicted knowledge point and the knowledge point label if the neural network model does not satisfy the convergence condition.
[0139] The determining module 303 is configured to determine the neural network model as a knowledge point prediction model if the neural network model satisfies the convergence condition.
[0140] As a possible implementation manner, the training module 301 is further configured to count the number of occurrences of the same initial knowledge point in different question samples.
[0141] The determining module 303 is further configured to determine the initial knowledge point as a to-be-output knowledge point of the knowledge point prediction model and select the knowledge point label of the question sample from the to-be-output knowledge point if the number of occurrences of the initial knowledge point satisfies a screening condition.
[0142] As a possible implementation manner, the screening condition comprises that the number of occurrences of the initial knowledge point is greater than or equal to a preset number of occurrences.
[0143] The to-be-output knowledge points of the knowledge point prediction model are the first n initial knowledge points with the largest number of occurrences, and n is less than the total number of initial knowledge points.
[0144] As a possible implementation manner, the neural network model comprises a plurality of layers of stacked transformer decoders, and each transformer decoder comprises a self-attention module with a mask and a feedforward neural network.
[0145] The self-attention module with a mask is configured to determine self-attention information based on input information, and the feedforward neural network is configured to determine the predicted knowledge point based on the self-attention information.
[0146] If the transformer decoder is a first layer transformer decoder, the input information is position encoding information of the segmentation result of the training sample; if the transformer decoder is an mth layer transformer decoder, the input information is a hidden state output by an (m-1)th layer transformer decoder, m is an integer greater than 1 and less than or equal to M, and M is the number of layers of the transformer decoder.
[0147] As a possible implementation manner, the determining module 303 is further configured to determine a knowledge point co-occurrence matrix based on the knowledge point labels of the plurality of question samples, and the knowledge point co-occurrence matrix is used to correct the knowledge point label predicted by the knowledge point prediction model in the inference stage.
[0148] The example embodiments of the present disclosure further provide a question knowledge point determination apparatus, which can be an electronic device or a chip applied to an electronic device. Figure 4 A schematic block diagram of the modules of the question knowledge point determination apparatus of the example embodiments of the present disclosure is shown. As shown in Figure 4 The apparatus 400 includes:
[0149] The prediction module 401 is configured to input the question information of the target question into a knowledge point prediction model to obtain a predicted knowledge point at a current time step, and the knowledge point prediction model is trained by the apparatus described in the example embodiments of the present disclosure.
[0150] The updating module 402 is configured to update the question information based on the predicted knowledge point at the current time step.
[0151] The determining module 403 is configured to determine the knowledge point of the target question based on the predicted knowledge points at a plurality of time steps.
[0152] As a possible implementation manner, the knowledge point prediction model includes a plurality of stacked transformer decoders, each of which includes a masked self-attention module and a feedforward neural network.
[0153] The masked self-attention module is configured to determine self-attention information based on input information, and the feedforward neural network is configured to determine a predicted knowledge point based on the self-attention information.
[0154] If the transformer decoder is a first layer transformer decoder, the input information is position encoding information of a word segmentation result of the question information; if the transformer decoder is an mth layer transformer decoder, the input information is a hidden state output by an (m-1)th layer transformer decoder, m is an integer greater than 1 and less than or equal to M, and M is the number of layers of the transformer decoder.
[0155] As a possible implementation manner, the topic information includes a topic and predicted knowledge points at each historical time step having a mapping relationship, and the predicted knowledge points at each historical time step constitute a predicted knowledge point sequence; the prediction module 401 is further configured to input the topic and the predicted knowledge point sequence into a knowledge point prediction model to obtain a plurality of candidate knowledge points corresponding to the predicted knowledge point sequence at a current time step; filter a candidate knowledge point corresponding to the predicted knowledge point sequence at the current time step from the plurality of candidate knowledge points, the candidate knowledge point satisfying a candidate knowledge point filtering condition; and determine a predicted knowledge point at the current time step based on the candidate knowledge point corresponding to the predicted knowledge point sequence at the current time step, each predicted knowledge point at the current time step and the corresponding predicted knowledge point sequence having a mapping relationship.
[0156] As a possible implementation manner, the candidate knowledge point filtering condition includes a condition probability of a first p candidate knowledge points in an order of the candidate knowledge points sorted according to the condition probability from large to small, the candidate knowledge point being the first p candidate knowledge points, and p being less than or equal to a total number of the candidate knowledge points.
[0157] As a possible implementation manner, each historical time step has a plurality of predicted knowledge points, a number of the predicted knowledge point sequences, a number of the predicted knowledge points at each historical time step, and a number of the predicted knowledge points at the current time step being the same; the prediction module 401 is further configured to determine a cumulative condition probability of the candidate knowledge point based on a condition probability of the candidate knowledge point and a condition probability of the corresponding predicted knowledge point sequence; and determine a plurality of predicted knowledge points at the current time step based on the cumulative condition probabilities of the plurality of candidate knowledge points and the corresponding predicted knowledge point sequences.
[0158] As a possible implementation manner, the determination module 403 is further configured to determine a knowledge point filtering parameter based on each predicted knowledge point at the current time step and the corresponding predicted knowledge point sequence; and generate a knowledge point of the topic based on the predicted knowledge point sequence if the knowledge point filtering parameter satisfies a preset termination condition.
[0159] As a possible implementation manner, the number of the predicted knowledge points at each time step is 1, the knowledge point filtering parameter includes a number of knowledge points, and the preset termination condition includes the number of knowledge points being greater than a preset number of knowledge points.
[0160] As a possible implementation manner, the determination module 403 is further configured to determine a cumulative condition probability of each predicted knowledge point at the current time step and the corresponding predicted knowledge point sequence based on each predicted knowledge point at the current time step and the corresponding predicted knowledge point sequence; and determine a knowledge point filtering parameter based on the cumulative condition probabilities of each predicted knowledge point at the current time step and the corresponding predicted knowledge point sequence, the knowledge point filtering parameter being a cumulative condition probability of a predicted knowledge point at the current time step and the corresponding predicted knowledge point sequence with the greatest cumulative condition probability.
[0161] As one possible implementation, the preset termination conditions include the predicted knowledge point at the current time step with the highest cumulative conditional probability and the cumulative conditional probability of the corresponding predicted knowledge point sequence being less than the preset cumulative conditional probability.
[0162] As one possible implementation, the device 400 further includes: a correction module 404, used to correct the conditional probability of the predicted knowledge point at the current time step based on the co-occurrence probability of the predicted knowledge point at the (t-1)th time step and the predicted knowledge point at the tth time step if the current time step is the tth time step, where t is an integer greater than or equal to 2.
[0163] As one possible implementation, the correction module 404 is also used to find the co-occurrence probability from the knowledge point co-occurrence matrix based on the predicted knowledge point at time step (t-1) and the predicted knowledge point at time step t. The knowledge point co-occurrence matrix is the knowledge point co-occurrence matrix determined by the knowledge point prediction model based on the knowledge point labels of multiple question samples during the training phase.
[0164] As one possible implementation, there is a mapping relationship between the predicted knowledge point at time step (t-1) and the predicted knowledge point at time step t; the ratio of the conditional probability of the predicted knowledge point at the current time step before and after correction is negatively correlated with the co-occurrence probability.
[0165] Figure 5 A schematic block diagram of a chip according to an exemplary embodiment of this disclosure is shown. (As follows) Figure 5 As shown, the chip 500 includes one or more (including two) processors 501 and a communication interface 502. The communication interface 502 can support the server in performing the data transmission and reception steps in the above method, and the processor 501 can support the server in performing the data processing steps in the above method.
[0166] Optional, such as Figure 5 As shown, the chip 500 also includes a memory 503, which may include read-only memory and random access memory, and provides operation instructions and data to the processor. A portion of the memory may also include non-volatile random access memory (NVRAM).
[0167] In some implementations, such as Figure 5As shown, the processor 501 executes various steps of the methods disclosed in embodiments of the present disclosure by calling stored instructions. The processor 501 controls the processing operations of any of the terminal devices, and can also be referred to as a central processing unit (CPU). The memory 503 can include a read-only memory and a random access memory, and provides instructions and data to the processor 501. A portion of the memory 503 can also include a NVRAM. The memory, the communication interface, and the bus system are coupled together via a bus system, which can include a data bus, a power supply bus, a control bus, and a state signal bus, etc. However, for the sake of clarity, all the buses are denoted as a bus system 504 in the following description. Figure 5
[0168] The method disclosed in the embodiments of the present disclosure can be applied to a processor or implemented by the processor. The processor can be an integrated circuit chip having a signal processing capability. In the implementation process, the steps of the above method can be completed by hardware integrated logic circuit or software form of instructions in the processor. The processor mentioned above can be a general processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The disclosed methods, steps and logic block diagrams in the embodiments of the present disclosure can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, or other mature storage media in the art. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps of the above method.
[0169] The exemplary embodiments of the present disclosure also provide an electronic device, including: at least one processor; and a memory connected with the at least one processor in communication. The memory stores a computer program capable of being executed by the at least one processor, and the computer program, when executed by the at least one processor, is configured to cause the electronic device to perform the method according to the embodiments of the present disclosure.
[0170] The exemplary embodiments of the present disclosure further provide a non-transitory computer readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, causes the computer to perform the method according to the embodiments of the present disclosure.
[0171] The exemplary embodiments of the present disclosure further provide a computer program product comprising a computer program, wherein the computer program, when executed by a processor of a computer, causes the computer to perform the method according to the embodiments of the present disclosure.
[0172] Reference Figure 6 will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent a wide variety of digital electronic computer devices such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computing devices. The electronic device can also represent a wide variety of mobile devices such as personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0173] As Figure 6 shown, the electronic device 600 includes a computing unit 601 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 602 or a computer program loaded into a random access memory (RAM) 603 from a storage unit 608. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0174] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, output unit 607, storage unit 608, and communication unit 609. Input unit 606 can be any type of device capable of inputting information to electronic device 600. Input unit 606 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 607 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 608 may include, but is not limited to, disks and optical discs. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0175] like Figure 6 As shown, computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 601 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 601 performs the various methods and processes described above. For example, in some embodiments, the methods of exemplary embodiments of this disclosure can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 600 via ROM 602 and / or communication unit 609. In some embodiments, computing unit 601 can be configured to perform methods by any other suitable means (e.g., by means of firmware).
[0176] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0177] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store program code for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0178] As used in this disclosure, the terms "machine-readable medium" and "computer- readable medium" refer to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal that can be used to provide machine instructions and / or data to a programmable processor.
[0179] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0180] The systems and techniques described here can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0181] The computer system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0182] In the embodiments described above, the whole or part of the embodiments can be realized by software, hardware, firmware, or any combination thereof. When realized by software, the whole or part of the embodiments can be realized in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When loaded and executed by a computer, the computer programs or instructions perform the flow or function described in the embodiments of the present disclosure. The computer can be a general purpose computer, a special purpose computer, a computer network, a terminal, user equipment, or other programmable apparatus. The computer programs or instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer programs or instructions can be transferred from one website site, computer, server, or data center to another website site, computer, server, or data center through a wired or wireless manner. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, and the like that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, a magnetic tape; an optical medium, such as a digital video disc (DVD); and a semiconductor medium, such as a solid state drive (SSD).
[0183] Although the present disclosure has been described in connection with certain specific features and embodiments thereof, it is to be understood that it is provided as an exemplification of the principles of the present disclosure and the features set forth herein are intended to be illustrative rather than limiting, and that numerous modifications and variations therein can be expected by those skilled in the art. Accordingly, it should be understood that the description and drawings are illustrative of the present disclosure and are not intended to be limiting. It should be understood that various changes can be made to the implementations described and the embodiments presented herein without departing from the spirit and scope of the present disclosure. It is intended that all such changes be considered as within the scope of the present disclosure.
Claims
1. A method for determining a topic knowledge point, characterized in that, The method comprises the following steps: inputting the question information of the target question into a knowledge point prediction model to obtain a predicted knowledge point at a current time step, wherein the knowledge point prediction model is obtained by training; incrementally updating the question information based on the predicted knowledge point at the current time step; determining the knowledge point of the target question based on predicted knowledge points at multiple time steps; the question information of the target question comprises the target question and predicted knowledge points at historical time steps having a mapping relationship, and the predicted knowledge points at the historical time steps having the mapping relationship form a predicted knowledge point sequence; the step of inputting the question information of the target question into the knowledge point prediction model to obtain the predicted knowledge point at the current time step comprises the following steps: inputting the target question and the predicted knowledge point sequence into the knowledge point prediction model to obtain multiple candidate knowledge points corresponding to the predicted knowledge point sequence at the current time step; selecting a candidate knowledge point corresponding to the predicted knowledge point sequence at the current time step from the multiple candidate knowledge points, wherein the candidate knowledge point satisfies a candidate knowledge point screening condition; determining the predicted knowledge point at the current time step based on the candidate knowledge point corresponding to the predicted knowledge point sequence at the current time step, wherein each predicted knowledge point at the current time step and the corresponding predicted knowledge point sequence have a mapping relationship.
2. The method of claim 1, wherein, the candidate knowledge point screening condition comprises a conditional probability of the first p candidate knowledge points in the order of descending conditional probability, p is less than or equal to the total number of the candidate knowledge points.
3. The method of claim 1, wherein, each historical time step has multiple predicted knowledge points, the number of the predicted knowledge point sequence, the number of the predicted knowledge points at each historical time step, and the number of the predicted knowledge points at the current time step are the same; the step of determining the predicted knowledge point at the current time step based on the candidate knowledge point corresponding to the predicted knowledge point sequence at the current time step comprises the following steps: determining a cumulative conditional probability corresponding to the candidate knowledge point based on the conditional probability of the candidate knowledge point and the conditional probability of the corresponding predicted knowledge point sequence; determining multiple predicted knowledge points at the current time step based on the cumulative conditional probabilities of the multiple candidate knowledge points and the corresponding predicted knowledge point sequence.
4. The method of claim 1, wherein, the step of determining the knowledge point of the target question based on the predicted knowledge points at multiple time steps comprises the following steps: determining a knowledge point screening parameter based on each predicted knowledge point at the current time step and the corresponding predicted knowledge point sequence; if the knowledge point screening parameter satisfies a preset termination condition, generating the knowledge point of the target question based on the predicted knowledge point sequence.
5. The method of claim 4, wherein, the number of the predicted knowledge point at each time step is 1, the knowledge point screening parameter comprises a knowledge point number, and the preset termination condition comprises that the knowledge point number is greater than a preset knowledge point number.
6. The method of claim 4, wherein, the step of determining the knowledge point screening parameter based on each predicted knowledge point at the current time step and the corresponding predicted knowledge point sequence comprises the following steps: determining, based on each of the predicted knowledge points of the current time step and the corresponding predicted knowledge point sequence, a cumulative conditional probability of each of the predicted knowledge points of the current time step and the corresponding predicted knowledge point sequence; determining, based on the cumulative conditional probabilities of each of the predicted knowledge points of the current time step and the corresponding predicted knowledge point sequence, a knowledge point screening parameter, the knowledge point screening parameter being the cumulative conditional probability of the predicted knowledge point of the current time step and the corresponding predicted knowledge point sequence with the largest cumulative conditional probability.
7. The method of claim 6, wherein, The preset termination condition comprises that the cumulative conditional probability of the predicted knowledge point of the current time step and the corresponding predicted knowledge point sequence with the largest cumulative conditional probability is less than a preset cumulative conditional probability.
8. The method according to any one of claims 1 to 7, characterized in that, If the current time step is the tth time step, t is an integer greater than or equal to 2, the method further comprises: correcting the conditional probability of the predicted knowledge point of the current time step based on the co-occurrence probability of the predicted knowledge point of the (t-1)th time step and the predicted knowledge point of the tth time step.
9. The method of claim 8, wherein, The method further comprises: finding the co-occurrence probability from the knowledge point co-occurrence matrix based on the predicted knowledge point of the (t-1)th time step and the predicted knowledge point of the tth time step, the knowledge point co-occurrence matrix being a knowledge point co-occurrence matrix determined by the knowledge point prediction model based on the knowledge point labels of the plurality of question samples in the training stage.
10. The method of claim 8, the predicted knowledge point of the (t-1)th time step and the predicted knowledge point of the tth time step having a mapping relationship; and / or, the ratio of the conditional probability of the predicted knowledge point of the current time step before and after correction is negatively correlated with the co-occurrence probability.
11. A training method for training the knowledge point prediction model according to any one of claims 1-10, characterized in that, comprising: adding the knowledge point label of the question sample in the word segmentation result of the question sample in the form of a knowledge point label as a word to obtain a word segmentation result of a training sample; inputting the word segmentation result of the training sample into a neural network model to obtain a predicted knowledge point of the question sample; if the neural network model does not satisfy the convergence condition, updating the neural network model based on the predicted knowledge point and the knowledge point label; if the neural network model satisfies the convergence condition, determining the neural network model as a knowledge point prediction model.
12. The method of claim 11, wherein, The method further comprises: counting the number of occurrences of the same initial knowledge point in different question samples; if the number of occurrences of the initial knowledge point satisfies a screening condition, determining the initial knowledge point as the knowledge point label.
13. The method of claim 12, wherein, The screening condition comprises that the number of occurrences of the initial knowledge point is greater than or equal to a preset number of occurrences; and / or, the knowledge point to be output by the knowledge point prediction model is the top n initial knowledge points with the largest number of occurrences, n being less than the total number of initial knowledge points.
14. The method of claim 11, wherein, The neural network model comprises a plurality of layers of stacked transformer decoders, each transformer decoder comprising a masked self-attention module and a feedforward neural network; the masked self-attention module is configured to determine self-attention information based on input information, and the feedforward neural network is configured to determine the predicted knowledge point based on the self-attention information. If the transformer decoder is a first layer transformer decoder, the input information is position encoding information of a word segmentation result of the training sample; if the transformer decoder is an mth layer transformer decoder, the input information is a hidden state output by an (m-1)th layer transformer decoder, m is an integer greater than 1 and less than or equal to M, and M is the number of layers of the transformer decoder.
15. The method of any one of claims 11-14, wherein, The method further comprises: determining a knowledge point co-occurrence matrix based on knowledge point labels of a plurality of question samples, the knowledge point co-occurrence matrix being used to correct the knowledge point labels predicted by the knowledge point prediction model in an inference stage.
16. A subject matter knowledge point determination apparatus characterized by comprising: Comprise: a prediction module configured to input question information of a question into a knowledge point prediction model to obtain a predicted knowledge point at a current time step, the knowledge point prediction model being obtained by training; an updating module configured to incrementally update the question information based on the predicted knowledge point at the current time step; a determination module configured to determine knowledge points of the question based on predicted knowledge points at a plurality of time steps; the question information of the question comprises the question and predicted knowledge points at historical time steps having a mapping relationship, and the predicted knowledge points at the historical time steps having the mapping relationship form a predicted knowledge point sequence; the prediction module is further configured to input the question and the predicted knowledge point sequence into the knowledge point prediction model to obtain a plurality of candidate knowledge points corresponding to the predicted knowledge point sequence at the current time step; and filter a candidate knowledge point corresponding to the predicted knowledge point sequence at the current time step from the plurality of candidate knowledge points, the candidate knowledge point satisfying a candidate knowledge point filtering condition; based on the candidate knowledge point corresponding to the predicted knowledge point sequence at the current time step, determine the predicted knowledge point at the current time step, each predicted knowledge point at the current time step and the corresponding predicted knowledge point sequence having a mapping relationship.
17. A training apparatus for training the knowledge point prediction model according to any one of claims 1-10, characterized in that, Comprise: a training module configured to add knowledge point labels of a question sample in a word segmentation result of the question sample in a form of a word of the knowledge point labels to obtain a word segmentation result of a training sample; the training module is further configured to input the word segmentation result of the training sample into a neural network model to obtain predicted knowledge points of the question sample; an updating module configured to, if the neural network model does not satisfy a convergence condition, update the neural network model based on the predicted knowledge points and the knowledge point labels; a determination module configured to, if the neural network model satisfies the convergence condition, determine the neural network model as a knowledge point prediction model.
18. An electronic device, comprising: Comprise: a processor; and a memory storing programs; wherein the programs comprise instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-15.
19. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-15.
Citation Information
Patent Citations
Knowledge point prediction method and device, electronic equipment and storage medium
CN115630696A
Character recognition method and device, equipment and storage medium
CN115690797A