Network teaching information automatic extraction system and method based on cloud computing
By collecting and analyzing the subtitles and comment data of course videos, combining the learning data of historical users, the professional words in online course videos are divided into grades, which solves the problem of inaccurate division of difficulty and importance in the existing technology, and achieves more accurate and flexible knowledge points extraction.
Patent Information
- Application Number
- CN202510541347.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing online course video knowledge point extraction technology cannot accurately divide the difficulty and importance of professional words based on the actual situation of the user, and cannot be adjusted in time based on the teaching content.
By collecting video subtitle files and comment data of course videos, extracting keywords, obtaining historical users' learning data, performing user classification and screening, and finally classifying the course-related knowledge point data to generate course-level knowledge point data.
It realizes the difficulty and importance of accurately dividing professional words according to the actual situation of users, and can be adjusted according to the changes in teaching content and user status, improving the accuracy and adaptability of knowledge point extraction.
Smart Images

Figure CN120067324A_ABST
Abstract
Claims
1. The automatic extraction system of online teaching information based on cloud computing is characterized by: It includes a text collection module, a text processing module, a learning collection module and a knowledge extraction module; The text collection module is used to collect video subtitle files of course videos and comment data of course videos to obtain course-related text data; The text processing module extracts keywords based on the course-related text data to obtain course-related knowledge point data; The learning collection module includes a data acquisition unit and a user classification unit. The data acquisition unit is used to obtain the playback data related to the historical user's learning course video, and the user classification unit is used to classify and filter the historical users to obtain relevant user data; The knowledge extraction module performs grade classification processing on the course-related knowledge point data based on the relevant user data to obtain the course graded knowledge point data.
2. The automatic extraction system of online teaching information based on cloud computing according to claim 1 is characterized in that: The text collection module is configured with a text collection strategy, which includes: For any course video in the online teaching process, record it as the first course video; obtain the video subtitle file of the first course video, convert it into an editable text file format, and mark it as the course subtitle text; Obtaining the user's comments under the first course video, including: recording any comment under the first course video as the first comment, obtaining and recording the text content of the first comment and the user ID that posted the first comment, repeatedly obtaining and recording all comments under the first course video, and converting them into a text file format that is the same as the course subtitle text, and recording them as the course comment text; The course subtitle text and the course comment text are recorded as course-related text data, and the online course to which the first course video belongs is recorded as the first course.
3. The automatic extraction system of online teaching information based on cloud computing according to claim 2 is characterized in that: The text processing module is configured with a text processing strategy, which includes: Any comment in the course review text is recorded as the second comment, and the user ID of the second comment is recorded as the second user. If there is a comment with the same text content as the second comment among all the comments posted by the second user, the second comment is marked as the first duplicate comment and removed. All comments in the course review text are processed in turn, and the initial screening review text is obtained after completion; Special characters are removed from the course subtitle text and the initial screening comment text respectively, and after completion, the first subtitle text and the first comment text are obtained in sequence; The continuous text sentences in the first subtitle text and the first comment text are divided into multiple independent words using a word segmentation tool, and the part-of-speech tagging tool is used to mark the parts of speech of the divided words in the corresponding text sentences. After completion, the second subtitle text and the second comment text are obtained in sequence.
4. The automatic extraction system of online teaching information based on cloud computing according to claim 3 is characterized in that: Text processing strategies also include: Set stop words and create a stop word list, traverse each word in the second subtitle text and the second comment text respectively, and for any word, if the same word exists in the stop word list, remove it, and after completion, obtain the third subtitle text and the third comment text; Any comment in the third comment text is recorded as the third comment, and the user ID to which the third comment belongs is recorded as the third user. If there is a comment with the same text content as the third comment among all the comments corresponding to the third user in the third comment text, the third comment will be marked as the second duplicate comment and removed. All comments in the third comment text are processed in turn, and the re-screened comment text is obtained after completion.
5. The automatic extraction system of online teaching information based on cloud computing according to claim 4 is characterized in that: Text processing strategies also include: Traversing each word in the third subtitle text and the re-screened comment text, and recording the professional words related to the first course in the third subtitle text and the re-screened comment text as course professional words; Repeatedly obtain all course professional terms in the third subtitle text and the re-screened comment text, and record them in order as subtitle word text and comment word text respectively; The total number of course professional words in the subtitle word text and the re-screened comment text are obtained respectively, and recorded as AZ and BZ in order; the number of any course professional words in the subtitle word text is obtained, and recorded as AC1, AC2, ..., ACn; the number of any course professional words in the subtitle word text is obtained, and recorded as BC1, BC2, ..., BCm; The subtitle word text and comment word text are recorded as course-related knowledge point data.
6. The automatic extraction system of online teaching information based on cloud computing according to claim 5 is characterized in that: The data acquisition unit is configured with a data acquisition strategy, which includes: For any historical user of the first course video, record it as the first historical user, obtain the cumulative time the first historical user has spent watching the first course video, record it as the first learning time; And obtain the video clips of slow play, replay and pause when the first historical user watches the first course video, record them as difficult course clips, and obtain the subtitle text corresponding to the difficult course clip according to the course subtitle text, record them as difficult subtitle text; Then obtain the percentage of the first historical user's score in the course test corresponding to the first course video to the total score, and record it as the first test score; Marking the first learning duration, difficult course segments, difficult subtitle texts, and first test scores as learning-related data of the first historical user in the first course video; The learning-related data of multiple historical users in the first course video are repeatedly obtained, and recorded as the playback-related data of the first course video according to the corresponding historical user classification.
7. The automatic extraction system of online teaching information based on cloud computing according to claim 6 is characterized in that: The user classification unit is configured with a user classification policy, which includes: Arrange the first learning durations of multiple historical users in ascending order, and remove the first learning durations that are shorter than the first course video duration, and after completion, obtain a first duration sequence; set the first threshold ratio to k1%, remove the largest k1% of the first learning durations in the first duration sequence, and obtain a second duration sequence; Obtain the first test score of any historical user belonging to the first learning duration in the second duration sequence; and sort the order of the users corresponding to the second duration sequence, recorded as the first score sequence; For any corresponding historical user in the second duration sequence, record it as the second user. If the second user meets any one of the screening conditions, the data corresponding to the second user in the second duration sequence and the first score sequence will be removed; the screening conditions include: condition one, the first learning duration of the second user is equal to the first course video duration, and the first test score of the second user is equal to 100%; condition two, the first test score of the second user is less than k2%; condition three, the first learning duration of the second user is greater than k4; k2% is the set second threshold ratio, and k4 is the set threshold duration; after completion, the third duration sequence and the second score sequence are obtained; The second score sequence is arranged in descending order and recorded as the third score sequence; the third score sequence is divided according to the ratio of p1:p2:p3:p4 to obtain a primary sequence, a secondary sequence, a tertiary sequence and a quaternary sequence in order; the historical users corresponding to the primary sequence, the secondary sequence, the tertiary sequence and the quaternary sequence are marked as primary users, secondary users, tertiary users and quaternary users in order.
8. The automatic extraction system of online teaching information based on cloud computing according to claim 7 is characterized in that: The knowledge extraction module is configured with a knowledge extraction strategy, which includes: Calculate the proportion of AC1, AC2, ..., ACn in AZ respectively, mark it as the subtitle proportion, and record it as AD1, AD2, ..., ADn in turn; and calculate the proportion of BC1, BC2, ..., BCm in BZ respectively, mark it as the comment proportion, and record it as BD1, BD2, ..., BDn in turn; Set the subtitle weight q1 and comment weight q2, q1+q2=1, 0 <q1,0<q2; For any course professional word in the subtitle word text and the comment word text, record it as the first professional word, obtain the subtitle proportion and comment proportion corresponding to the first professional word, and record them as ADi and BDi in order; and calculate the comprehensive proportion corresponding to the first professional word through the comprehensive proportion formula, which is as follows: ; Where ABi represents the comprehensive proportion of the first professional word pair; Repeatedly obtain the comprehensive proportions of all course professional words in the subtitle word text and the comment word text, and arrange them in order from large to small, which is recorded as the first proportion sequence; obtain the largest k3% comprehensive proportion in the first proportion sequence, which is recorded as the core proportion sequence, and obtain the smallest k4% comprehensive proportion in the first proportion sequence, which is recorded as the general proportion sequence, and then record the remaining comprehensive proportions in the first proportion sequence as the important proportion sequence; The course professional terms corresponding to the core proportion sequence, important proportion sequence and general proportion sequence are marked in order as core knowledge terms, important knowledge terms and general knowledge terms respectively; and the core knowledge terms, important knowledge terms, general knowledge terms and corresponding word explanations are obtained and recorded as important knowledge point information.
9. The automatic extraction system of online teaching information based on cloud computing according to claim 8 is characterized in that: Knowledge extraction strategies also include: Get the average values of the secondary sequence and the tertiary sequence respectively, record them as PU1 and PU2 in order, and calculate PU0, PU0=q3*PU1+q4*PU2, where q3 and q4 are the set weight coefficients, q3+q4=1, 0 <q3,0<q4; Obtain the course professional words in the difficult subtitle texts corresponding to the first-level user, the second-level user, the third-level user and the fourth-level user, and record them in order as the first-level word text, the second-level word text, the third-level word text and the fourth-level word text; Perform duplicate filtering on the first-level word text, second-level word text, third-level word text, and fourth-level word text in sequence, including: for any course major word, if it exists in the first-level word text, remove the same course major words in the second-level word text, third-level word text, and fourth-level word text; if it exists in the second-level word text, remove the same course major words in the third-level word text and fourth-level word text; if it exists in the third-level word text, remove the same course major words in the third and fourth-level word texts; after completion, obtain the first-level major text, second-level major text, third-level major text, and fourth-level major text in sequence; Set score thresholds PV1 and PV2, where PV2 > PV1; mark the course major words in the fourth-level major text as simple knowledge words; If PV1 < PU0 < PV2, mark the course major words in the first-level major text as difficult knowledge words, and mark the course major words in the second-level major text and third-level major text as normal knowledge words; If PV2 ≤ PU0, mark the course major words in the first-level major text as difficult knowledge words, mark the course major words in the second-level major text as normal knowledge words, and mark the course major words in the third-level major text as simple knowledge words; If PU0 ≤ PV1, mark the course major words in the first-level major text and second-level major text as difficult knowledge words, and mark the course major words in the third-level major text as normal knowledge words; Obtain difficult knowledge words, normal knowledge words, simple knowledge words, and their corresponding word explanations, denoted as difficulty knowledge point information; Merge and store the difficulty knowledge point information and important knowledge point information according to the corresponding course major words, denoted as course classification knowledge point data.
10. A method for automatically extracting network teaching information based on cloud computing, used to implement any of the automatic network teaching information extraction systems based on cloud computing according to claims 1-9, characterized in that: Include the following steps: Collect the video subtitle files of the course videos and the comment data of the course videos to obtain course-related text data; Extract keywords based on the course-related text data to obtain course-related knowledge point data; Obtain the playback-related data of historical users learning the course videos, and perform classification and filtering on the historical users to obtain relevant user data; Perform level division processing on the course-related knowledge point data based on the relevant user data to obtain course classification knowledge point data.
Citation Information
Patent Citations
Automatic knowledge extraction method for network teaching system
CN113641716A
Teaching feedback system based on knowledge point analysis
CN108520662A
Method for positioning knowledge points in course video clip
CN115797825A
Language learning data recommendation method and device, electronic equipment and readable storage medium
CN118673221A
Learning recommendation method and device for occupational development and storage medium
CN118966858A