A multi-label classification method for math problem text based on mathematical feature extraction

By combining the mathematical feature extraction method with prior knowledge, a feature prior tree is constructed, and the problems of insufficient feature extraction and noise in the multi-label classification of mathematical problems in the existing technology are solved, achieving higher classification accuracy and more accurate knowledge point recognition.

CN114880474BActive Publication Date: 2025-05-06JIANGSU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210485759.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-06
Publication Date
2025-05-06
Estimated Expiration
2042-05-06

AI Technical Summary

Technical Problem

The prior art has problems in the multi-label classification of mathematical problems with insufficient feature extraction, large impact on model effects due to manual selection of features, high noise and sparse data, resulting in low classification accuracy.

Method used

Using a mathematical feature extraction method combining prior knowledge, multiple knowledge point labels are extracted through the higher-order tag correlation characteristics of the sequence-to-sequence model, a feature prior tree is constructed, and mathematical feature extraction is performed to improve the classifier's recognition ability.

Benefits of technology

Effectively identify the derivation information in the text of math questions, reduce the function fitting time, improve the accuracy of multi-label text classification of math questions, significantly improve the classification effect, and is suitable for the classification of knowledge points for junior high school mathematics questions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114880474B_ABST
    Figure CN114880474B_ABST
Patent Text Reader

Abstract

The invention discloses a multi-label classification method for math problem text based on mathematical feature extraction, which takes math problem test questions as samples and knowledge points as sample labels; preprocesses and extracts features for samples and their labels, and encodes sample feature vectors to obtain hidden layer vectors; uses a self-attention mechanism to calculate the attention weights of each hidden layer vector to obtain a feature vector of text output; divides answer analysis text into leaf nodes and root nodes, and forms a feature matrix of a feature priori tree by using leaf node text information features and root node text information features; performs mathematical feature extraction on sample feature vectors and the feature matrix of the feature priori tree, inputs the feature vector of text output and the output result of the mathematical feature extraction part into a classifier, and outputs a classification result by the classifier; sets a training stop condition, and obtains a trained math text multi-label classification model when the training stops; and effectively classifies math problem text by using the math text multi-label classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, and in particular to a method for multi-label text classification of math problems by extracting mathematical features based on prior knowledge. Background Art

[0002] In recent years, with the development of computing power, artificial intelligence theory and application have made breakthrough progress, and have been widely implemented in the fields of computer vision, natural language processing, recommendation algorithms, etc., and integrated into all aspects of daily life. For example, ubiquitous biometric recognition technology; recommendation technology integrated into various consumer scenarios; intelligent customer service, machine translation, text mining, risk control, assisted driving and other technologies have greatly changed people's lifestyles. Using artificial intelligence technology to replace repetitive manual labor and improve efficiency has become an obvious trend in various industries. As an important guarantee for population quality and national reserve forces, the application and research of artificial intelligence in the field of education has become a hot topic in academia and industry. At present, my country has problems such as uneven distribution of educational resources and lack of personalized education. Using artificial intelligence technology to promote the efficiency of knowledge processing and provide students with comprehensive and personalized educational services is beneficial to the development of individuals and the country. Mathematics, as a basic subject for cultivating logical ability, deserves key research.

[0003] In the field of education, from professional books and test papers to fragmented online knowledge, the most important resource is knowledge carried by language and text. Therefore, natural language processing technology has many application scenarios in the field of education.

[0004] However, current research on natural language processing technology is mostly concentrated in scenarios such as search, news, and e-commerce, with relatively little research on the field of education, resulting in problems such as insufficient corpus resources, uneven corpus quality, and a lack of methods tailored to the characteristics of the education field. As one of the important educational infrastructures, question banks are very important for consolidating knowledge and testing learning outcomes, especially in mathematics. The development of logical thinking requires a lot of training, so elementary mathematics question banks have important application value.

[0005] Multi-label text classification of math problems solves the problem that there is relatively little research in the field of education, which leads to a lack of corpus resources and a lack of methods tailored to the characteristics of the education field. Multi-label text classification of math problems can help students consolidate their knowledge and improve their learning outcomes. It can also train students to develop logical thinking, allowing them to conduct personalized learning and improve their math scores.

[0006] Automatic classification of question types can provide abstract question features for other tasks, such as automatic construction of question banks, analysis of easy-to-make mistakes, recommendation of related questions, automatic test paper compilation, etc. It also makes it possible for question banks to organize and manage massive amounts of questions, and is one of the basic components of intelligent question banks. In many application scenarios, manual labeling is time-consuming and labor-intensive, and automatic labeling systems can save time and effort. There are few studies on the application of natural language processing based on the characteristics of mathematical texts, especially the lack of research on mathematical word segmentation and named entity recognition technology. Applying general technologies usually does not achieve good results. Therefore, this article's research on the field of education can effectively expand the scope of application of natural language technology and provide certain experience for the application of general natural language processing technology in vertical fields.

[0007] There are few existing studies on automatic classification of math problems, and they mainly focus on manually extracting features at the text level of math problems. Traditional shallow machine learning algorithms such as naive Bayes and support vector machines are used for classification. The model effect is greatly affected by manually selected features, and the text representation method based on statistical indicators such as word frequency loses a lot of information. Lv et al. adopted a first-order strategy to train multiple binary classifiers, splice the final prediction results, and ignore the correlation between labels. Moreover, using machine learning methods, the model effect is greatly affected by manually selected features. There is a lot of noise and the proportion of sparse data is reduced, which leads to the disappearance of data aggregation effect and makes feature learning difficult. Ye et al. adopted a deep learning method, combined with word vectors that express semantic information in the field of natural language processing, to train the model to generate knowledge point labels, but the math text has a strong logic, and some derived information classifiers cannot be recognized.

[0008] Based on the problems existing in the current technology, the present invention adopts the excellent high-order label correlation characteristics of the sequence-to-sequence model and proposes a mathematical feature extraction method combined with prior knowledge to generate multiple knowledge point labels, thereby solving this problem well. Summary of the invention

[0009] In order to solve the deficiencies in the prior art, the present invention proposes a multi-label classification method for math problem text based on mathematical feature extraction, which combines the mathematical feature extraction method of prior knowledge to enable the classifier to identify the deductive information present in the math problem text, reduce the function fitting time, and improve the accuracy of multi-label text classification of math problems.

[0010] The technical solution adopted by the present invention is as follows:

[0011] A multi-label classification method for math problem text based on mathematical feature extraction includes the following steps:

[0012] Step 1: Collect test questions from multiple sets of mathematics test papers as samples to form a sample set; and use the knowledge point of each question as the label of the sample;

[0013] The samples and their labels are preprocessed and feature extracted to obtain the sample feature vectors corresponding to the samples and the labels corresponding to the sample feature vectors to form a sample feature vector set w = {(x1, y1), … (x n ,y n )}, where x i is the eigenvector of the i-th sample, y i is the label corresponding to the i-th sample, i = 1, 2, ..., n, n is the number of samples;

[0014] Encode the sample feature vector to obtain the hidden layer vector h s ; Quoted from the attention mechanism to calculate each hidden layer vector h s The attention weight a s ;

[0015] Based on the obtained attention weight a s , perform weighted summation on the hidden layer vectors to obtain the feature vector of the text output

[0016] Step 2: Divide the sample set into a training set and a test set; obtain the answer parsing text corresponding to the samples in the training set. The answer parsing text is divided into leaf nodes and root nodes. The root node is the label text information of the answer parsing, and the leaf node is the text information that can directly or indirectly derive the root node label.

[0017] The answer analysis text is preprocessed and feature extracted to obtain the leaf node text information features and the root node text information features; the feature matrix of the feature prior tree is formed by the leaf node text information features and the root node text information features, which is expressed as: v = {v1, v2, ...v n};

[0018] Step 3: Perform mathematical feature extraction on the sample feature vector and the feature matrix of the feature prior tree to obtain the output result of the mathematical feature extraction part. i ;

[0019] Step 4: Output the feature vector F of the text of the training set q And the output result of the mathematical feature extraction part i Input the classifier, and the classifier outputs the classification result;

[0020] Step 5, setting the training stop condition of the training set, and obtaining the trained mathematics text multi-label classification model when the training stops; applying the trained mathematics text multi-label classification model to classify the mathematics problem text.

[0021] Furthermore, Word2vec is used in both steps 1 and 2 for word embedding to obtain feature vectors.

[0022] Furthermore, the method for encoding the sample feature vector is:

[0023] The sample feature vector is input into the encoder of the BILSTM model for encoding, and the encoding corresponding to each sample feature vector is output; it is expressed as:

[0024]

[0025]

[0026] Get the hidden layer vector in, Represents the output vector of BILSTM in two directions at time s; LSTM is a long short-term memory structure.

[0027] Furthermore, the self-attention mechanism is used to calculate the attention weights a of each hidden layer vector s :

[0028] u s =tanh(W w h s +b w )

[0029]

[0030] Among them, u s is the vector calculated at time s, s = 1, 2, ..., S, S is the total encoding time; tanh is the hyperbolic tangent function, W w ,b w are the weight parameters and bias terms of the self-attention mechanism; exp is the exponential function.

[0031] Furthermore, the mathematical feature extraction part is composed of multiple base feature extractions. In each base feature extraction, the sample feature vector and the vector v in the feature matrix of the feature prior tree are combined. i It is expressed as:

[0032] l i =f(w,v i ,W i 1 ,W i 2 )

[0033] Among them, l i is the output of base feature extraction, and the f function is the function of w,v i Vector cosine and compare operations; W i1 is used for w and v i The parameter matrix for consine calculation, W i 2 It is the calibration parameter matrix used to compare with the calculated cosine value.

[0034] Furthermore, the specific process of base feature extraction is as follows:

[0035] Step 1: Initial parameter matrix W i 1 , check parameter matrix W i 2 With the weight vector a i save;

[0036] Step 2: Calculate w,v i Similarity value: m i =consine(W i 1 ⊙w,W i 1 ⊙v i )

[0037] Step 3: Compare the similarity values ​​with the verification parameter matrix W i 2 To verify:

[0038]

[0039] l i is the output of base feature extraction, obtained according to the calculation Make a judgment, if Greater than the verification parameter W i 2 , then the text information feature vector of the root node of the subtree is output as the mathematical feature; otherwise, the text information feature vector of the root node of the subtree is discarded, and the mathematical feature vector output by the next subtree is calculated.

[0040] Further, the process of step 4 is:

[0041] The feature vector F output from the training set question text q And the output result of the mathematical feature extraction part i Perform the concat operation and convert the l after the concat operation i After being passed to the fully connected layer, it is input into the activation function ReLu: f(x) = max(0,x) to obtain the output result of the activation function;

[0042] The second fully connected layer transforms the output of the activation function into a vector with a length equal to the number of label categories: F' = (f1, f2, ... f n ), and finally the probability of each label category is obtained after softmax normalization:

[0043]

[0044] Among them, f c is the output of the cth activation function, p c is the probability value of the cth label category, c∈[1,C], C represents the total number of label categories;

[0045] Furthermore, the training stop condition in step 5 is that the loss is less than 1e-6 or the number of iterations is greater than the threshold. The method for calculating the loss is:

[0046]

[0047] Among them, L is the cross entropy loss, y c (i) is the true value of the label, p c (i) is the predicted value of the label in the previous step, n is the total number of training sets, and C is the total number of label categories.

[0048] Furthermore, the method for preprocessing samples and their labels is as follows:

[0049] Preprocessing includes removing non-text parts of the data, doing word segmentation, and removing stop words. After preprocessing, all samples and labels are in a unified format.

[0050] Beneficial effects of the present invention:

[0051] 1. The present invention performs weighted summation on the hidden layer vector based on the obtained attention weights to obtain a feature vector of the text output; the advantage of such processing is that features can be extracted based on the text context information to learn label information.

[0052] 2. Construct a feature prior tree with domain prior knowledge for the answer parsing text, which can obtain inferential features that cannot be extracted from the text information.

[0053] 3. Mathematical feature extraction based on the feature vector of the question text and the feature matrix of the feature prior tree can extract partial label information based on inferential knowledge, thereby improving the overall label classification accuracy.

[0054] 4. This application constructs a mathematical text multi-label classification model for mathematical feature extraction, and uses the text multi-label classification model to classify the problem text of junior high school mathematics problems, and outputs the classification results; the classification results can be used as a knowledge point classification outline for junior high school mathematics problems to assist teachers in assigning practice questions in a targeted manner, and can also assist students in targeted filling of weak knowledge points, effectively improving children's mathematics learning performance and providing students with a personalized learning approach.

[0055] In addition, this application solves the problem that it is difficult to mine the knowledge point information in the multi-label classification model of the math problem text, and also solves the problem of excessive noise interference in the math problem text, and can significantly improve the multi-label classification effect of the math problem text, which is of great significance to the teacher's auxiliary teaching and the students' independent personalized learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a network architecture diagram that combines feature prior knowledge.

[0057] Figure 2 It is a priori knowledge mathematical feature extraction graph.

[0058] Figure 3 This is a flow chart of a multi-label classification method for math problem text based on mathematical feature extraction in this application.

[0059] Figure 4 This is a schematic diagram of the application constructing a feature prior tree with domain knowledge based on the answer parsing text of the test questions. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0061] like Figure 1 and 3 The multi-label classification method of math problem text based on mathematical feature extraction shown in the figure includes the following steps:

[0062] Step 1: collect test questions from multiple sets of mathematics test papers as samples to form a sample set; and use the knowledge points of each question as labels of the samples; more specifically, the knowledge points corresponding to each sample (i.e., test questions) can be obtained by manual labeling.

[0063] The samples and their labels are preprocessed and feature extracted to obtain the sample feature vectors corresponding to the samples and the labels corresponding to the sample feature vectors to form a sample feature vector set.

[0064] More specifically, the method for preprocessing samples and their labels is:

[0065] Preprocessing includes removing non-text parts of the data, doing word segmentation, removing stop words, etc. After preprocessing, a unified format of all samples and labels is obtained. Take a test question as an example:

[0066] For example: In an isosceles triangle ABC, the perpendicular bisector DE of AB intersects BC at point D. The perimeter of triangle ABC is 17cm. Find the length of AB.

[0067] After preprocessing, the form obtained is:

[0068] Isosceles @Triangle0, @Line0 perpendicular bisector @Line1 intersects @Line2 at point @Point0, @Triangle0 has a perimeter of 17cm, find the length of @Line1\L 53|59

[0069] Among them, isosceles @Triangle0, @Line0 perpendicular bisector @Line1 intersects @Line2 at point @Point0, @Triangle0 has a perimeter of 17cm, find the length of @Line1 as the preprocessed sample, 53|59 is the label of the preprocessed sample;

[0070] More specifically, Word2vec is used to embed words in the preprocessed samples to obtain sample feature vectors; the sample feature vectors and the labels corresponding to the sample feature vectors form a sample feature vector set, which is expressed as:

[0071] w={(x1,y1),...(x n ,y n )},

[0072] Among them, x i is the eigenvector of the i-th sample, y i is the label corresponding to the i-th sample, i = 1, 2, ..., n, and n is the number of samples.

[0073] Encode the sample feature vector to obtain the hidden layer vector; the attention weight a of each hidden layer vector is calculated by quoting the self-attention mechanism s .

[0074] More specifically, the method for encoding the sample feature vector is:

[0075] The sample feature vector is input into the encoder of the BILSTM model for encoding, and the encoding corresponding to each sample feature vector is output; it is expressed as:

[0076]

[0077]

[0078] Get the hidden layer vector in, Represents the vector of BILSTM output in two directions at time s; LSTM is a long short-term memory structure.

[0079] Based on the attention weight a s , perform weighted summation on the hidden layer vectors to obtain the feature vector F of the text output q .

[0080] More specifically, the self-attention mechanism is used to calculate the attention weights a of each hidden layer vector s :

[0081] u s =tanh(W w h s +b w )

[0082]

[0083] Among them, u s is the vector calculated at time s, s = 1, 2, ..., S, S is the total encoding time; tanh is the hyperbolic tangent function, W w ,b w are the weight parameters and bias terms of the self-attention mechanism; exp is the exponential function.

[0084] Based on the obtained attention weight a s , perform weighted summation on the hidden layer vectors to obtain the feature vector F of the text output q , expressed as:

[0085] Step 2: Divide the sample set into a training set and a test set; obtain the answer analysis text corresponding to the samples in the training set. The answer analysis text is divided into leaf node text information and root node text information. The root node text information is the label of the answer analysis. It should be noted that for the answer analysis of each test question, the text information that can directly or indirectly deduce the root node text information in the answer analysis is used as the leaf node text information, such as Figure 4 Take part of the answer analysis of a question as an example to illustrate:

[0086] Answer analysis: Because AB is parallel to CD, and AC is parallel to BD. So ABCD is a parallelogram.

[0087] According to the answer analysis text shown above, since "AB is parallel to CD, AC is parallel to BD" can directly lead to "ABCD is a parallelogram", the answer analysis text is divided into "AB is parallel to CD, AC is parallel to BD" as the leaf node text information, and "parallelogram" as the root node text information.

[0088] The answer analysis text is preprocessed and feature extracted to obtain leaf node text information features and root node text information features; the feature matrix of the feature prior tree is formed by the leaf node text information features and the root node text information features.

[0089] More specifically, feature extraction is to embed the answer parsed text using Word2vec to obtain the feature matrix of the feature prior tree. The feature prior matrix is ​​expressed as:

[0090] v={v1,v2,...v n}

[0091] Where v is the feature matrix; v i is the feature vector formed by the i-th row of features in the feature matrix, n is the number of rows in the feature matrix; the vector is represented by v i = {v 1i ,v 2i ,...v mi},v mi is the vector v i The mth feature in the vector v i In , the tag words can be divided into feature vectors corresponding to leaf nodes and feature vectors corresponding to root nodes.

[0092] Step 3: Perform mathematical feature extraction on the sample feature vector and the feature matrix of the feature prior tree to obtain the output result of the mathematical feature extraction part. i .

[0093] More specifically, if Figure 2 The mathematical feature extraction part consists of multiple base feature extractions. In each base feature extraction, the sample feature vector and the vector v in the feature matrix of the feature prior tree are combined. i It is expressed as:

[0094] l i =f(w,v i ,W i 1 ,W i 2 )

[0095] Among them, l i is the output of base feature extraction, and the f function is the function of w,v i Vector cosine and compare operations; Wi 1 is used for w and v i The parameter matrix for consine calculation, W i 2 It is the calibration parameter matrix used to compare with the calculated cosine value.

[0096] More specifically, the specific process of base feature extraction is as follows:

[0097] Step 1: Initial parameter matrix W i 1 , check parameter matrix W i 2 With the weight vector a i save;

[0098] Step 2: Calculate w,v i Similarity value: m i =consine(W i 1 ⊙w,W i 1 ⊙v i )

[0099] Step 3: Compare the similarity values ​​with the verification parameter matrix W i 2 To verify:

[0100]

[0101] l i is the output of base feature extraction, obtained according to the calculation Make a judgment, if Greater than the verification parameter W i 2 , then the text information feature vector of the root node of the subtree is output as the mathematical feature; otherwise, the text information feature vector of the root node of the subtree is discarded, and the mathematical feature vector output by the next subtree is calculated.

[0102] Step 4: Output the feature vector F of the text of the training set q And the output result of the mathematical feature extraction part i Input the classifier and the classifier outputs the classification result.

[0103] More specifically, the feature vector F output from the training set question text q And the output result of the mathematical feature extraction part i Perform the concat operation and convert the l after the concat operation iAfter being passed to the fully connected layer, it is input into the activation function ReLu: f(x) = max(0,x) to obtain the output result of the activation function.

[0104] The second fully connected layer transforms the output of the activation function into a vector with a length equal to the number of label categories: F' = (f1, f2, ... f n ), and finally the probability of each label category is obtained after softmax normalization:

[0105]

[0106] Among them, f c is the output of the cth activation function, p c is the probability value of the cth label category, c∈[1,C], C represents the total number of label categories;

[0107] Step 5, calculate the loss function: Among them, L is the cross entropy loss, y c (i) is the true value of the label, p c (i) is the predicted value of the label in the previous step, n is the total number of training sets, and C is the total number of label categories.

[0108] Set the iteration number threshold. When the loss is less than 1e-6 or the iteration number is greater than the threshold, stop processing the training text; otherwise, continue iterative training, that is, continue training the training set to obtain a trained mathematical text multi-label classification model.

[0109] Input the test sample question text into the constructed mathematical text multi-label classification model, and output the classification results of the knowledge points designed for the question.

[0110] The above embodiments are only used to illustrate the design ideas and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, any equivalent changes or modifications made based on the principles and design ideas disclosed by the present invention are within the protection scope of the present invention.

Claims

1. A multi-label classification method for math problem text based on mathematical feature extraction, characterized in that: The steps include: Step 1: Collect test questions from multiple sets of mathematics test papers as samples to form a sample set; and use the knowledge point of each question as the label of the sample; The samples and their labels are preprocessed and feature extracted to obtain the sample feature vectors corresponding to the samples and the labels corresponding to the sample feature vectors to form a sample feature vector set w = {(x1, y1), … (x n ,y n )}, where x i is the eigenvector of the i-th sample, y i is the label corresponding to the i-th sample, i = 1, 2, ..., n, n is the number of samples; Encode the sample feature vector to obtain the hidden layer vector h s ; Quoted from the attention mechanism to calculate each hidden layer vector h s The attention weight a s ; Based on the obtained attention weight a s , perform weighted summation on the hidden layer vectors to obtain the feature vector of the text output Step 2: Divide the sample set into a training set and a test set; obtain the answer parsing text corresponding to the samples in the training set. The answer parsing text is divided into leaf nodes and root nodes. The root node is the label text information of the answer parsing, and the leaf node is the text information that can directly or indirectly derive the root node label. The answer analysis text is preprocessed and feature extracted to obtain the leaf node text information features and the root node text information features; the feature matrix of the feature prior tree is formed by the leaf node text information features and the root node text information features, which is expressed as: v = {v1, v2, ...v n }; Step 3: Perform mathematical feature extraction on the sample feature vector and the feature matrix of the feature prior tree to obtain the output result of the mathematical feature extraction part. i ; The mathematical feature extraction part consists of multiple base feature extractions. In each base feature extraction, the sample feature vector and the vector v in the feature matrix of the feature prior tree are combined. i It is expressed as: Among them, l i is the output of base feature extraction, and the f function is the function of w,v i Vector cosine and compare operations; W i 1 is used for w and v i The parameter matrix for consine calculation, It is the calibration parameter matrix used to compare with the calculated cosine value; Step 4: Output the feature vector F of the text of the training set q And the output result of the mathematical feature extraction part i Input the classifier, and the classifier outputs the classification result; Step 5, setting the training stop condition of the training set, and obtaining the trained mathematics text multi-label classification model when the training stops; applying the trained mathematics text multi-label classification model to classify the mathematics problem text.

2. According to claim 1, a multi-label classification method for math problem text based on mathematical feature extraction is characterized in that: In both steps 1 and 2, Word2vec is used for word embedding to obtain feature vectors.

3. According to claim 1, a multi-label classification method for math problem text based on mathematical feature extraction is characterized in that: The method for encoding the sample feature vector is: The sample feature vector is input into the encoder of the BILSTM model for encoding, and the encoding corresponding to each sample feature vector is output; it is expressed as: Get the hidden layer vector in, Represents the output vector of BILSTM in two directions at time s; LSTM is a long short-term memory structure.

4. According to claim 3, a multi-label classification method for math problem text based on mathematical feature extraction is characterized in that: The attention weight a of each hidden layer vector is calculated by the attention mechanism s : u s =tanh(W w h s +b w ) Among them, u s is the vector calculated at time s, s = 1, 2, ..., S, S is the total encoding time; tanh is the hyperbolic tangent function, W w ,b w are the weight parameters and bias terms of the self-attention mechanism; exp is the exponential function.

5. According to claim 1, a multi-label classification method for math problem text based on mathematical feature extraction is characterized in that: The specific process of base feature extraction is as follows: Step 1: Initial parameter matrix Calibration parameter matrix With the weight vector a i save; Step 2: Calculate w,v i Similar values ​​for: Step 3: Compare the similarity values ​​with the verification parameter matrix To verify: l i is the output of base feature extraction, obtained according to the calculation Make a judgment, if Greater than the verification parameter Then the text information feature vector of the root node of the subtree is output as a mathematical feature; Otherwise, the text information feature vector of the root node of the subtree is discarded, and the mathematical feature vector output by the next subtree is calculated.

6. The method for multi-label classification of math problem text based on mathematical feature extraction according to claim 1 is characterized in that: The process of step 4 is: The feature vector F output from the training set question text q And the output result of the mathematical feature extraction part i Perform concat operation and convert the l after concat operation i After being passed to the fully connected layer, it is input into the activation function ReLu: f(x) = max(0,x) to obtain the output result of the activation function; The second fully connected layer transforms the output of the activation function into a vector with a length equal to the number of label categories: F'=(f1,f2,...f n ), and finally the probability of each label category is obtained after softmax normalization: Among them, f c is the output of the cth activation function, p c is the probability value of the cth label category, c∈[1,C], and C represents the total number of label categories.

7. The method for multi-label classification of math problem text based on mathematical feature extraction according to claim 1 is characterized in that: The training stop condition in step 5 is that the loss is less than 1e-6 or the number of iterations is greater than the threshold. The method for calculating the loss is: Among them, L is the cross entropy loss, y c (i) is the true value of the label, p c (i) is the predicted value of the label in the previous step, n is the total number of training sets, and C is the total number of label categories.

8. The method for multi-label classification of math problem text based on mathematical feature extraction according to claim 1 is characterized in that: The method for preprocessing samples and their labels is: Preprocessing includes removing non-text parts of the data, doing word segmentation, and removing stop words. After preprocessing, all samples and labels are in a unified format.

Citation Information

Patent Citations

  • Text classification method based on generative multi-task learning model

    CN110347839A

  • Text classification method, text classification device, computer equipment and storage medium

    CN113486175A