English learning material automatic classification method and system based on RF algorithm

Through the automatic classification method based on RF algorithm, English learning materials are automatically processed and classified, solving the problems of inefficient and insufficient accuracy of manual classification, and achieving efficient and accurate learning resource management and personalized learning support.

CN120067749AInactive Publication Date: 2025-05-30XINXIANG VOCATIONAL & TECHN COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510079701.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-18
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the manual classification of English learning materials is inefficient and error-prone, and cannot meet the needs of integration and personalized learning of educational resources.

Method used

An automatic classification method based on RF algorithm is adopted to obtain and preprocess learning materials (text, audio, video) of different display types, extract data features, and build a classification model based on the voting mechanism of multiple decision trees, label and archive learning materials.

Benefits of technology

It realizes efficient automatic classification of learning materials of different display types, improves learning resource management efficiency, avoids omissions and inefficiencies in manual operations, supports personalized learning needs, and improves the accuracy and stability of classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067749A_ABST
    Figure CN120067749A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of learning material classification, and discloses an automatic English learning material classification method and system based on an RF algorithm, and the method comprises the steps: obtaining English learning materials of a plurality of display types, and carrying out the preprocessing of the English learning materials based on the display types, and the display types comprise a text type, an audio type and a video type. And extracting data features of the preprocessed English learning materials. A classification model is established based on an RF algorithm, and each English learning material is labeled based on the classification model and the data features to label attributes. And classifying and filing the English learning materials according to the label attributes. According to the method, the English learning materials of different display types are subjected to preprocessing, feature extraction, classification and archiving by using the RF algorithm, a large number of English learning materials can be efficiently, accurately and automatically classified and archived, the resource management efficiency is improved, and intelligent processing and personalized recommendation of learning contents are promoted.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A method for automatic classification of English learning materials based on RF algorithm, characterized in that: include: Acquire English learning materials of several display types, and pre-process the English learning materials based on the display types, wherein the display types include text type, audio type and video type; Extracting data features of each of the English learning materials after preprocessing; Establishing a classification model based on the RF algorithm, and labeling each of the English learning materials with label attributes based on the classification model and data features; According to the tag attributes, each of the English learning materials is classified and filed.

2. The method for automatic classification of English learning materials based on RF algorithm as claimed in claim 1, characterized in that: When the English learning material is preprocessed based on the display type, it includes: Determine the preprocessing method based on the presentation type of the English learning material: When the display type is a text type, the English learning material is not preprocessed; When the display type is an audio type, the preprocessing is to convert the English learning materials into text information based on speech recognition technology; When the display type is a video type, the preprocessing method of the English learning material is determined according to whether the English learning material contains audio, wherein: If the English learning material contains audio, determining the preprocessing method is to convert the English learning material into text information based on speech recognition technology; If the English learning material does not contain audio, the preprocessing method is determined to be to extract a number of video frame images based on the timeline, eliminate duplicate images from the video frame images, extract the same-format text from the video frame images after eliminating duplicate images, and generate text information of the English learning material based on the timestamp and the same-format text based on the video frame images.

3. The method for automatic classification of English learning materials based on RF algorithm as claimed in claim 2, characterized in that: When extracting the data features of each of the English learning materials after preprocessing, it includes: Performing stemming processing on the text information of each of the English learning materials, the stemming processing includes removing stop words and punctuation marks in the text information, extracting each stem in the text information, and restoring the word form of the stem; The text features of the text information after stemming are extracted based on the TF-IDF method: TF-IDF(t,d,D)=TF(t,d)×IDF(t,D); Wherein, TF(t, d) is the frequency of occurrence of the word stem t in the text information d, and IDF(t, D) is the inverse document frequency of the word stem t in the text information D; Vectorizing each word in the text information based on a word embedding method, extracting N-gram features, and determining context information in the text information; The text features and N-gram features are merged and a feature vector of each word is determined.

4. The method for automatic classification of English learning materials based on RF algorithm as claimed in claim 3, characterized in that: A classification model is established based on the RF algorithm, and based on the classification model and data features, each of the English learning materials is labeled with label attributes, including: Based on the extracted text features and N-gram features, a classification model is constructed using the RF algorithm. The classification model classifies the English learning materials through a voting mechanism of multiple decision trees and annotates label attributes for each type of learning material, where: The classification model includes multiple decision trees {T1, T2, ..., Tn}, each decision tree independently makes a classification decision based on input feature data, and the feature data includes text features and N-gram features; The output classification label of each decision tree is yi, and the final classification result is determined by the voting mechanism of all decision trees: Among them, yfinaI is the final classification label, I(yi=y) is the indicator function, when the output label of the i-th decision tree is y, I(yi=y)=1, otherwise it is 0.

5. The method for automatic classification of English learning materials based on RF algorithm as claimed in claim 4, characterized in that: According to the tag attributes, the English learning materials are classified and filed, including: Obtain the label attributes of the English learning materials under each category, and obtain the weight coefficient of the label attributes of each English learning material under each category: Among them, Count(Lj) is the number of times label Lj appears in all samples. is the sum of the number of times all labels appear in all samples; Acquire the label attribute whose weight coefficient is greater than that corresponding to each of the weight coefficients, and determine it as the preset label attribute of the classification; According to the relationship between the tag attribute and the preset tag attribute, the classification of the English learning material is determined and archived.

6. The automatic classification system of English learning materials based on RF algorithm adopts the automatic classification method of English learning materials based on RF algorithm as described in claims 1-5, characterized in that: include: A collection module is configured to obtain English learning materials of several display types and pre-process the English learning materials based on the display types, wherein the display types include text type, audio type and video type; A feature extraction module, electrically connected to the acquisition module, and configured to extract data features of each of the English learning materials after preprocessing; A labeling module, electrically connected to the feature extraction module, the labeling module is configured to establish a classification model based on the RF algorithm, and label each of the English learning materials with label attributes based on the classification model and data features; A classification module is electrically connected to the labeling module, and is configured to classify and archive each of the English learning materials according to the label attributes.

7. The automatic classification system for English learning materials based on RF algorithm as claimed in claim 6, characterized in that: When the acquisition module pre-processes the English learning material based on the display type, it includes: The acquisition module is further configured to determine a preprocessing method according to the presentation type of the English learning material: When the display type is a text type, the acquisition module does not pre-process the English learning materials; When the display type is an audio type, the acquisition module pre-processes the English learning materials into text information based on speech recognition technology; When the display type is a video type, the acquisition module determines the preprocessing method of the English learning material according to whether the English learning material contains audio, wherein: If the English learning material contains audio, the acquisition module determines that the pre-processing method is to convert the English learning material into text information based on speech recognition technology; If the English learning material does not contain audio, the acquisition module determines that the preprocessing method is to extract a number of video frame images based on the time axis, eliminate duplicate images from the video frame images, extract the same-format text from the video frame images after eliminating duplicate images, and generate text information of the English learning material based on the timestamp and the same-format text based on the video frame images.

8. The automatic classification system for English learning materials based on RF algorithm as claimed in claim 7, characterized in that: When the feature extraction module extracts the data features of each of the English learning materials after preprocessing, it includes: The feature extraction module is further configured to perform stemming on the text information of each of the English learning materials, wherein the stemming includes removing stop words and punctuation marks in the text information, extracting each stem in the text information, and restoring the word form of the stem; The feature extraction module is further configured to extract text features from the text information after stemming based on the TF-IDF method: TF-IDF(t,d,D)=TF(t,d)×IDF(t,D); Wherein, TF(t, d) is the frequency of occurrence of the word stem t in the text information d, and IDF(t, D) is the inverse document frequency of the word stem t in the text information D; The feature extraction module is further configured to vectorize each word in the text information based on a word embedding method, extract N-gram features, and determine context information in the text information; The feature extraction module is further configured to merge the text features and N-gram features and determine a feature vector for each of the words.

9. The automatic classification system for English learning materials based on RF algorithm as claimed in claim 8, characterized in that: The labeling module establishes a classification model based on the RF algorithm, and labels each of the English learning materials for label attributes based on the classification model and data features, including: The labeling module is further configured to construct a classification model using an RF algorithm based on the extracted text features and N-gram features. The classification model classifies the English learning materials through a voting mechanism of multiple decision trees and labels each type of learning material with label attributes, wherein: The classification model includes multiple decision trees {T1, T2, ..., Tn}, each decision tree independently makes a classification decision based on input feature data, and the feature data includes text features and N-gram features; The output classification label of each decision tree is yi, and the final classification result is determined by the voting mechanism of all decision trees: Among them, yfinal is the final classification label, I(yi=y) is the indicator function, when the output label of the i-th decision tree is y, I(yi=y)=1, otherwise it is 0.

10. The automatic classification system for English learning materials based on RF algorithm as claimed in claim 9, characterized in that: When the classification module classifies and files each of the English learning materials according to the tag attributes, it includes: The classification module is further configured to obtain the label attributes of the English learning materials under each category, and obtain the weight coefficient of the label attributes of each English learning material under each category: Among them, Count(Lj) is the number of times label Lj appears in all samples. is the sum of the number of times all labels appear in all samples; The classification module is further configured to obtain the label attribute whose weight coefficient is greater than that corresponding to each of the weight coefficients, and determine it as the preset label attribute of the classification; The classification module is further configured to determine the classification of the English learning material according to the relationship between the tag attribute and the preset tag attribute, and archive it.