Automatic Extraction System and Method for Network Teaching Information Based on Cloud Computing
By collecting and analyzing the subtitles and comment data of online course videos, combining user learning behaviors, the level division of professional words is solved, and the problem of inability to accurately divide the difficulty and importance of professional words in the existing technology is solved, and personalized teaching content adjustment is achieved.
Patent Information
- Application Number
- CN202510541347.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing online course video knowledge point extraction technology cannot accurately divide the difficulty and importance of professional words based on the actual situation of the user, and cannot adjust it in a timely manner based on the teaching content based on changes in user status.
By collecting video subtitle files and comment data of course videos, extracting keywords, obtaining historical users' learning data, and classifying and filtering users, classifying and classifying and classifying knowledge point data based on relevant data to obtain course grading knowledge point data.
It realizes the accuracy of the difficulty and importance of professional words based on the actual situation of the user, and can be adjusted in a timely manner according to changes in the user's status, improving the adaptability and learning efficiency of teaching content.
Smart Images

Figure CN120067324B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of online course knowledge extraction, and specifically to an automatic extraction system and method for online teaching information based on cloud computing. Background Art
[0002] The technology of extracting knowledge points from online course videos refers to the technology of using a series of algorithms and tools to analyze, process, and refine the teaching content in online course videos, and automatically identify and extract the key knowledge points therein.
[0003] In the course videos of natural science, knowledge points are often surrounded by a large number of explanations and examples; through knowledge point extraction, key professional terms can be accurately refined, helping users quickly clarify the core concepts, avoiding getting lost in the lengthy narration, and improving the quality and efficiency of learning; there are differences in the understanding difficulty and importance of these professional terms; when the existing technology of extracting knowledge points from online course videos extracts and classifies the key and difficult points of professional terms in the course videos, it often relies on traditional teaching experience for classification and processing, lacking a unified objective standard for the understanding and judgment of key and difficult points, making it difficult to scientifically quantify the difficulty and importance of professional terms, and it is difficult to take into account the individual differences of students; moreover, with the continuous development of the education field, traditional teaching experience was formed in past teaching practices and may not be able to keep up with educational changes in a timely manner; the knowledge bases, learning abilities, and learning attitudes of different users are different, and even for users of the same grade of students, there may be differences in their learning states in different semesters; relying on traditional teaching experience to classify key and difficult points may not be able to flexibly adapt to these special situations. For example, in the patent application with the publication number CN113641716A, a method for automatically extracting knowledge for an online teaching system is disclosed, and this solution lacks a unified objective standard for classifying the understanding difficulty and importance of knowledge points, making it difficult to scientifically quantify the difficulty and importance of professional terms; therefore, when the existing technology of extracting knowledge points from online course videos extracts and classifies the key and difficult points of professional terms in the course videos, it cannot accurately classify the difficulty and importance of professional terms according to the actual situation of users, and at the same time, adjust in a timely manner according to the changes in the user state based on the teaching content. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems in the prior art to a certain extent. By collecting the video subtitle files of course videos and the comment data of course videos, course-related text data is obtained; and keyword extraction is performed to obtain course-related knowledge point data; the playback-related data of historical users learning course videos is obtained, and the historical users are classified and screened to obtain relevant user data; and the course-related knowledge point data is hierarchically classified to obtain course hierarchical knowledge point data, so as to solve the problem that when the existing online course video knowledge point extraction technology extracts professional terms in course videos and classifies the key and difficult points, it is impossible to accurately classify the difficulty and importance of professional terms according to the actual situation of users, and at the same time, adjust in a timely manner according to the change of the user state based on the teaching content.
[0005] To achieve the above object, in a first aspect, the present application provides an automatic extraction system for online teaching information based on cloud computing, including a text collection module, a text processing module, a learning collection module, and a knowledge extraction module;
[0006] The text collection module is used to collect the video subtitle files of course videos and the comment data of course videos to obtain course-related text data;
[0007] The text processing module performs keyword extraction based on the course-related text data to obtain course-related knowledge point data;
[0008] The learning collection module includes a data acquisition unit and a user classification unit. The data acquisition unit is used to acquire the playback-related data of historical users learning course videos, and the user classification unit is used to classify and screen the historical users to obtain relevant user data;
[0009] The knowledge extraction module performs hierarchical classification processing on the course-related knowledge point data based on the relevant user data to obtain course hierarchical knowledge point data.
[0010] Further, the text collection module is configured with a text collection strategy, and the text collection strategy includes:
[0011] For any course video in the online teaching process, denoted as the first course video; obtain the video subtitle file of the first course video and convert it into an editable text file format, marked as the course subtitle text;
[0012] Obtain the comments of users under the first course video, including: for any comment under the first course video, denoted as the first comment, obtain and record the text content of the first comment and the user ID that posted the first comment, repeat to obtain and record all the comments under the first course video, and convert them into the same text file format as the course subtitle text, denoted as the course comment text;
[0013] Record the course subtitle text and the course review text as course-related text data, and record the online course to which the first course video belongs as the first course.
[0014] Furthermore, the text processing module is configured with a text processing strategy, and the text processing strategy includes:
[0015] Record any one comment in the course review text as the second comment, and record the user ID to which the second comment belongs as the second user. If there is a comment in all the comments posted by the second user that is exactly the same as the second comment in terms of text content, then mark the second comment as the first duplicate comment and remove it. Process all the comments in the course review text in turn. After completion, obtain the preliminarily screened review text.
[0016] Remove special characters from the course subtitle text and the preliminarily screened review text respectively. After completion, obtain the first subtitle text and the first review text in sequence.
[0017] Use a word segmentation tool to divide the continuous text sentences in the first subtitle text and the first review text into multiple independent words respectively, and use a part-of-speech tagging tool to tag the part of speech of the divided words in the corresponding text sentences. After completion, obtain the second subtitle text and the second review text in sequence.
[0018] Furthermore, the text processing strategy also includes:
[0019] Set stop words and establish a stop word list. Traverse each word in the second subtitle text and the second review text respectively. For any word, if the same word exists in the stop word list, then remove it. After completion, obtain the third subtitle text and the third review text.
[0020] Record any one comment in the third review text as the third comment, and record the user ID to which the third comment belongs as the third user. If there is a comment in all the comments corresponding to the third user in the third review text that is exactly the same as the third comment in terms of text content, then mark the third comment as the second duplicate comment and remove it. Process all the comments in the third review text in turn. After completion, obtain the re-screened review text.
[0021] Furthermore, the text processing strategy also includes:
[0022] Traverse each word in the third subtitle text and the re-screened review text, and record the professional words related to the first course in the third subtitle text and the re-screened review text as course professional words.
[0023] Repeat to obtain all the course professional words in the third subtitle text and the re-screened review text respectively, and record them as the subtitle word text and the review word text in sequence.
[0024] Obtain the total number of course major terms in the subtitle text and the re-screened review text respectively, and record them as AZ and BZ in sequence; obtain the number of occurrences of any course major term in the subtitle text, and record them as AC1, AC2, ……, ACn respectively; obtain the number of occurrences of any course major term in the subtitle text, and record them as BC1, BC2, ……, BCm respectively;
[0025] Record the subtitle text and the review text as course-related knowledge point data.
[0026] Furthermore, the data acquisition unit is configured with a data acquisition strategy, and the data acquisition strategy includes:
[0027] For any historical user of the first course video, denoted as the first historical user, obtain the cumulative duration of the first historical user watching the first course video, denoted as the first learning duration;
[0028] And obtain the video segments of slow playback, rewind, and pause when the first historical user watches the first course video, denoted as difficult course segments, and obtain the subtitle text corresponding to the difficult course segments according to the course subtitle text, denoted as difficult subtitle text;
[0029] Then obtain the percentage of the score of the first historical user in the course quiz corresponding to the first course video in the total score, denoted as the first quiz score;
[0030] Mark the first learning duration, the difficult course segments, the difficult subtitle text, and the first quiz score as the learning-related data of the first historical user in the first course video;
[0031] Repeat to obtain the learning-related data of multiple historical users in the first course video, and classify them according to the corresponding historical users, denoted as the playback-related data of the first course video.
[0032] Furthermore, the user classification unit is configured with a user classification strategy, and the user classification strategy includes:
[0033] Arrange the first learning durations of multiple historical users in ascending order, and remove the first learning durations that are less than the duration of the first course video. After completion, obtain the first duration sequence; set the first threshold ratio as k1%, and remove the largest k1% of the first learning durations in the first duration sequence to obtain the second duration sequence;
[0034] Obtain the first quiz score of the historical user to which any first learning duration in the second duration sequence belongs; and sort them in the order of the users corresponding to the second duration sequence, denoted as the first score sequence;
[0035] For any corresponding historical user in the second duration sequence, denoted as the second user, if the second user meets any one of the screening conditions, the data corresponding to the second user in the second duration sequence and the first score sequence will be removed; the screening conditions include: Condition 1, the first learning duration of the second user is equal to the first course video duration, and the first quiz score of the second user is equal to 100%; Condition 2, the first quiz score of the second user is less than k2%; Condition 3, the first learning duration of the second user is greater than k4; where k2% is the set second threshold ratio and k4 is the set threshold duration; after completion, the third duration sequence and the second score sequence are obtained.
[0036] Arrange the second score sequence in descending order, denoted as the third score sequence; divide the third score sequence according to the ratio of p1:p2:p3:p4, and sequentially obtain the first-level sequence, the second-level sequence, the third-level sequence, and the fourth-level sequence; sequentially mark the historical users corresponding to the first-level sequence, the second-level sequence, the third-level sequence, and the fourth-level sequence as first-level users, second-level users, third-level users, and fourth-level users.
[0037] Furthermore, the knowledge extraction module is configured with a knowledge extraction strategy, and the knowledge extraction strategy includes:
[0038] Calculate the proportions of AC1, AC2, ……, ACn in AZ respectively, marked as the subtitle proportion, and sequentially denoted as AD1, AD2, ……, ADn; and calculate the proportions of BC1, BC2, ……, BCm in BZ respectively, marked as the comment proportion, and sequentially denoted as BD1, BD2, ……, BDn;
[0039] Set the subtitle weight q1 and the comment weight q2, q1 + q2 = 1, 0 < q1, 0 < q2;
[0040] For any course professional word in the subtitle word text and the comment word text, denoted as the first professional word, obtain the subtitle proportion and the comment proportion corresponding to the first professional word, and sequentially denoted as ADi and BDi; and calculate the comprehensive proportion corresponding to the first professional word through the comprehensive proportion formula, and the comprehensive proportion formula is as follows: ; where ABi represents the comprehensive proportion of the first professional word pair;
[0041] Repeat to obtain the comprehensive proportions of all course professional words in the subtitle word text and the comment word text, and arrange them in descending order, denoted as the first proportion sequence; obtain the largest k3% of the comprehensive proportions in the first proportion sequence, denoted as the core proportion sequence, and obtain the smallest k4% of the comprehensive proportions in the first proportion sequence, denoted as the general proportion sequence, and then denote the remaining comprehensive proportions in the first proportion sequence as the important proportion sequence;
[0042] Sequentially mark the course professional terms corresponding to the core proportion sequence, important proportion sequence, and general proportion sequence as core knowledge terms, important knowledge terms, and general knowledge terms respectively; and obtain the core knowledge terms, important knowledge terms, general knowledge terms, and their corresponding word explanations, which are recorded as important knowledge point information.
[0043] Furthermore, the knowledge extraction strategy also includes:
[0044] Respectively obtain the averages of the secondary sequence and the tertiary sequence, which are sequentially recorded as PU1 and PU2, and calculate PU0, where PU0 = q3 * PU1 + q4 * PU2, and q3 and q4 are set weight coefficients, q3 + q4 = 1, 0 < q3, 0 < q4;
[0045] Obtain the course professional terms in the difficult subtitle texts corresponding to the first-level user, second-level user, third-level user, and fourth-level user respectively, and sequentially record them as the first-level word text, second-level word text, third-level word text, and fourth-level word text;
[0046] Sequentially perform duplicate filtering processing on the first-level word text, second-level word text, third-level word text, and fourth-level word text, including: for any course professional term, if it exists in the first-level word text, remove the same course professional terms in the second-level word text, third-level word text, and fourth-level word text; if it exists in the second-level word text, remove the same course professional terms in the third-level word text and fourth-level word text; if it exists in the third-level word text, remove the same course professional terms in the third-level and fourth-level word texts; after completion, obtain the first-level professional text, second-level professional text, third-level professional text, and fourth-level professional text in sequence;
[0047] Set score thresholds PV1 and PV2, where PV2 > PV1; mark the course professional terms in the fourth-level professional text as simple knowledge terms;
[0048] If PV1 < PU0 < PV2, mark the course professional terms in the first-level professional text as difficult knowledge terms, and mark the course professional terms in the second-level professional text and third-level professional text as normal knowledge terms;
[0049] If PV2 ≤ PU0, mark the course professional terms in the first-level professional text as difficult knowledge terms, mark the course professional terms in the second-level professional text as normal knowledge terms, and mark the course professional terms in the third-level professional text as simple knowledge terms;
[0050] If PU0 ≤ PV1, mark the course professional terms in the first-level professional text and second-level professional text as difficult knowledge terms, and mark the course professional terms in the third-level professional text as normal knowledge terms;
[0051] Obtain difficult knowledge words, normal knowledge words, simple knowledge words, and their corresponding word explanations, which are recorded as difficult and easy knowledge point information;
[0052] Merge and store the difficult and easy knowledge point information and the important knowledge point information according to the corresponding course professional words, which is recorded as course hierarchical knowledge point data.
[0053] In a second aspect, the present application provides an automatic extraction method for online teaching information based on cloud computing, including the following steps:
[0054] Collect the video subtitle files of the course video and the comment data of the course video to obtain course-related text data;
[0055] Extract keywords based on the course-related text data to obtain course-related knowledge point data;
[0056] Obtain the playback-related data of historical users learning the course video, and perform classification and screening processing on historical users to obtain relevant user data;
[0057] Based on the relevant user data, perform a hierarchical classification process on the course-related knowledge point data to obtain course hierarchical knowledge point data.
[0058] Advantages of the present invention: The present invention collects the video subtitle files of the course video and the comment data of the course video to obtain course-related text data; extracts keywords based on the course-related text data to obtain course-related knowledge point data; obtains the playback-related data of historical users learning the course video, and performs classification and screening processing on historical users to obtain relevant user data; based on the relevant user data, performs a hierarchical classification process on the course-related knowledge point data to obtain course hierarchical knowledge point data; it can accurately classify the difficulty and importance of professional words according to the actual situation of users, and at the same time, adjust in a timely manner according to the change of the user state based on the teaching content.
[0059] The present invention obtains the subtitle text and comment text of the course video, and performs two screenings on the comment text. The advantage is that it can remove similar and identical comments of the same user, avoiding the interference of duplicate data on subsequent classification processing, and making the subsequent processing based on comments more accurate; by calculating the comprehensive proportion of course professional words through weighted summation to divide the importance, it can comprehensively consider the proportion of course professional words in the subtitle text and comment text, and objectively reflect the importance of course professional words in the course content and student discussions; by dividing users into different levels and dividing the difficulty level through the learning-related data of historical users, it can well reflect the mastery situation and demand differences of different-level users for course knowledge; it can accurately find out knowledge points at different difficulty levels, which helps teaching workers adjust teaching strategies according to the actual situation of users. Brief Description of the Drawings
[0060] Figure 1 It is a schematic block diagram of the system of the present invention;
[0061] Figure 2 It is a flowchart of the steps of the method of the present invention;
[0062] Figure 3 It is a flowchart for dividing the difficulty level of the course professional terms of the present invention;
[0063] Figure 4 It is a schematic structural diagram of the electronic device of the present invention. Detailed Description of the Invention
[0064] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0065] Embodiment 1, please refer to Figure 1 As shown, the present application provides a network teaching information automatic extraction system based on cloud computing, including a text collection module, a text processing module, a learning collection module, and a knowledge extraction module;
[0066] The text collection module is used to collect the video subtitle files of the course videos and the comment data of the course videos to obtain course-related text data;
[0067] The text collection module is configured with a text collection strategy, and the text collection strategy includes: for any course video in the network teaching process, denoted as the first course video; obtaining the video subtitle file of the first course video and converting it into an editable text file format, marked as the course subtitle text; for example, it can be converted into the txt format, and the method will be processed later;
[0068] Obtaining the comments of users under the first course video, including: for any comment under the first course video, denoted as the first comment, obtaining and recording the text content of the first comment and the user ID that published the first comment, repeatedly obtaining and recording all the comments under the first course video, and converting them into the same text file format as the course subtitle text, denoted as the course comment text;
[0069] Denoting the course subtitle text and the course comment text as course-related text data, and denoting the network course to which the first course video belongs as the first course;
[0070] In the specific implementation process, the course subtitle text can accurately present the core knowledge content such as professional terms, concept explanations, and principle elaborations in the course, and can directly reflect the professional knowledge in the course content; the course review text is the feedback and discussion of students on the course content, including students' understanding, questions, and discussions of the course; obtaining the course subtitle text and the course review text can provide an accurate and comprehensive text basis for subsequent knowledge point extraction and division.
[0071] The text processing module extracts keywords based on the course-related text data to obtain course-related knowledge point data;
[0072] The text processing module is configured with a text processing strategy. The text processing strategy includes: regarding any comment in the course review text as the second comment, and regarding the user ID to which the second comment belongs as the second user. If there is a comment in all the comments posted by the second user that is exactly the same as the second comment in terms of text content, then the second comment is marked as the first repeated comment and removed. All the comments in the course review text are processed in turn. After completion, the preliminary screened review text is obtained; that is, the same comments of the same user are removed to avoid the interference caused by the repeated comments of the same user to the subsequent processing.
[0073] Special characters in the course subtitle text and the preliminary screened review text are removed respectively. After completion, the first subtitle text and the first review text are obtained in sequence; the text may contain various special characters such as @, #, and $, including emoji characters. These characters have nothing to do with the teaching content itself but will interfere with subsequent analysis.
[0074] The continuous text sentences in the first subtitle text and the first review text are respectively divided into multiple independent words by using a word segmentation tool, and the parts of speech of the divided words in the corresponding text sentences are marked by using a part-of-speech tagging tool. After completion, the second subtitle text and the second review text are obtained in sequence; for example, when performing accurate mode word segmentation on the sentence "Online teaching breaks the time and space limitations of traditional teaching", the result will be "Online teaching", "breaks the", "traditional teaching", "of", "time and space limitations". This mode can accurately reflect the semantic structure of the sentence; common parts of speech include nouns, verbs, adjectives, adverbs, etc.; through part-of-speech tagging, the grammatical function of each word in the sentence can be clarified, which is helpful for the subsequent screening of course professional words.
[0075] Stop words are set according to common and meaningless words, and a stop word list is established, such as "的", "了", "在", "和", etc. Although these words play a connecting and auxiliary role in grammar, they will increase the amount of data, distract attention, and affect the grasp of key information in the process of knowledge point extraction; therefore, removing stop words can streamline the text and highlight the core content; traverse each word in the second subtitle text and the second comment text respectively, for any word, if the same word exists in the stop word list, remove it, and after completion, obtain the third subtitle text and the third comment text;
[0076] Any comment in the third comment text is recorded as the third comment, and the user ID to which the third comment belongs is recorded as the third user. If there is a comment that is exactly the same as the third comment in the text content among all the comments corresponding to the third user in the third comment text, the third comment is marked as the second duplicate comment and removed. All comments in the third comment text are processed in turn, and a re-screened comment text is obtained after completion. The re-screening can remove the same comments of the same user after removing the stop words, that is, the original similar comments, which is helpful for the subsequent screening of course professional terms and improves the accuracy of subsequent results.
[0077] Traversing each word in the third subtitle text and the re-screened comment text, and recording the professional words related to the first course in the third subtitle text and the re-screened comment text as course professional words;
[0078] Repeatedly obtain all course professional terms in the third subtitle text and the re-screened comment text, and record them in order as subtitle word text and comment word text respectively;
[0079] Get the total number of course professional words in the subtitle word text and the re-screened comment text respectively, and record them as AZ and BZ in order; get the number of any course professional words in the subtitle word text, and record them as AC1, AC2, ..., ACn; get the number of any course professional words in the subtitle word text, and record them as BC1, BC2, ..., BCm; AC and BC have the same sequence number, and the corresponding course professional words are the same. For example, if AC1 corresponds to "gravity", then BC1 also corresponds to "gravity";
[0080] Record the subtitle word text and comment word text as course-related knowledge point data;
[0081] In the specific implementation process, the comment text is the students' feedback and discussion on the course content, which may include further explanation and expansion of the course knowledge points, as well as relevant background knowledge supplemented by the students after consulting the materials themselves; these comment texts may contain some course professional terms that do not exist in the subtitle text, that is, implicit knowledge points related to the course, so n and m may not be equal.
[0082] The learning collection module includes a data acquisition unit and a user classification unit. The data acquisition unit is used to acquire the playback-related data of historical user learning course videos, and the user classification unit is used to classify and screen historical users to obtain relevant user data;
[0083] The data acquisition unit is configured with a data acquisition strategy, which includes: for any historical user of the first course video, denoted as the first historical user, acquiring the cumulative duration of the first historical user watching the first course video, denoted as the first learning duration; the cumulative duration includes the duration of slow playback, rewind, and pause when the user watches the first course video; for example, if a course video is 40 minutes long, and in the first 10 minutes, the user watches at 0.5 speed, and in the second 10 minutes, the user rewinds it once, then the cumulative duration is 40 + 10 + 10 = 60 minutes;
[0084] And acquiring the video segments of slow playback, rewind, and pause when the first historical user watches the first course video, denoted as difficult course segments, and acquiring the subtitle text corresponding to the difficult course segments according to the course subtitle text, denoted as difficult subtitle text; the places where the user slows down, rewinds, and pauses when watching the course video are often where the knowledge content is more complex and abstract, and the user slows down, rewinds, and pauses to better think, understand, and learn;
[0085] Then acquiring the percentage of the score of the first historical user in the course quiz corresponding to the first course video in the total score, denoted as the first quiz score; the total scores of the course quizzes corresponding to different course videos may be inconsistent, and acquiring the percentage of the score in the total score facilitates subsequent processing and analysis under the same standard;
[0086] Marking the first learning duration, difficult course segments, difficult subtitle text, and the first quiz score as the learning-related data of the first historical user in the first course video;
[0087] Repeating to acquire the learning-related data of multiple historical users in the first course video and classifying them according to the corresponding historical users, denoted as the playback-related data of the first course video;
[0088] The user classification unit is configured with a user classification strategy, which includes: arranging the first learning durations of multiple historical users in ascending order, and removing the first learning durations that are less than the first course video duration. After completion, a first duration sequence is obtained; setting the first threshold ratio to k1%, and removing the largest k1% of the first learning durations in the first duration sequence to obtain a second duration sequence; in this embodiment, k1% is 10%. There may be some abnormal learning duration data in the learning duration data, which may be extreme values caused by system errors, user misoperations, or users' repeated viewings; these data have no reference value; removing the largest k1% of the first learning durations can effectively exclude the influence of these abnormal values on the overall data distribution and make the subsequent analysis results more accurate and reliable;
[0089] Obtain the first test score of the historical user to which any one of the first learning durations in the second duration sequence belongs; and sort them in the order of the users corresponding to the second duration sequence, denoted as the first score sequence;
[0090] For any historical user corresponding to the second duration sequence, denoted as the second user, if the second user meets any one of the screening conditions, then remove the data of the second user corresponding in the second duration sequence and the first score sequence; the screening conditions include: Condition 1, the first learning duration of the second user is equal to the first course video duration, and the first test score of the second user is equal to 100%; Condition 2, the first test score of the second user is less than k2%; Condition 3, the first learning duration of the second user is greater than k4; where k2% is the set second threshold ratio, and k4 is the set threshold duration; after completion, a third duration sequence and a second score sequence are obtained; in this embodiment, k2% is 30%; k4 can be set according to the first course video duration, generally more than 5 times the first course video duration;
[0091] Condition 1 represents users with extremely high learning levels or knowledge bases; Condition 2 represents users with extremely low learning levels or knowledge bases; Condition 3 represents users with extremely large first learning durations due to special reasons. The data of these three types of users have no reference value for most normal users, so they need to be removed;
[0092] Arrange the second score sequence in descending order, denoted as the third score sequence; divide the third score sequence according to the ratio of p1:p2:p3:p4 to obtain the first-level sequence, second-level sequence, third-level sequence, and fourth-level sequence in sequence; mark the historical users corresponding to the first-level sequence, second-level sequence, third-level sequence, and fourth-level sequence as first-level users, second-level users, third-level users, and fourth-level users in sequence; in this embodiment, p1:p2:p3:p4 is 10:15:15:40; the first-level users, second-level users, third-level users, and fourth-level users are users with excellent grades, good grades, average grades, and poor grades in sequence;
[0093] In the specific implementation process, when counting the cumulative duration of the first historical user watching the first course video, although the user may pause the video to think about the content just heard or seen, record important knowledge points, formulas, cases, etc.; these states can be counted in the cumulative duration; but it may also be that the user needs a short break, adjust the learning state, relieve learning fatigue, or has to pause the video due to external factors interference, and these states cannot be counted in the cumulative duration and need to be distinguished.
[0094] The knowledge extraction module performs hierarchical classification processing on the course-related knowledge point data based on relevant user data to obtain the course hierarchical knowledge point data;
[0095] The knowledge extraction module is configured with a knowledge extraction strategy, and the knowledge extraction strategy includes: calculating the proportions of AC1, AC2,..., ACn in AZ respectively, marked as the subtitle proportion, denoted as AD1, AD2,..., ADn in sequence; and calculating the proportions of BC1, BC2,..., BCm in BZ respectively, marked as the comment proportion, denoted as BD1, BD2,..., BDn in sequence;
[0096] Set the subtitle weight q1 and the comment weight q2, q1 + q2 = 1, 0 < q1, 0 < q2; in this embodiment, q1 = q2 = 0.5, which can be set according to the actual application scenario;
[0097] For any course professional word in the subtitle text and the comment text, denoted as the first professional word, obtain the subtitle proportion and the comment proportion corresponding to the first professional word, denoted as ADi and BDi in sequence; and calculate the comprehensive proportion corresponding to the first professional word through the comprehensive proportion formula, and the comprehensive proportion formula is as follows: ; where ABi represents the comprehensive proportion of the first professional word pair;
[0098] Repeatedly obtain the comprehensive proportion of all course professional terms in the subtitle word text and the comment word text, and arrange them in descending order, denoted as the first proportion sequence; obtain the largest k3% of the comprehensive proportions in the first proportion sequence, denoted as the core proportion sequence, and obtain the smallest k4% of the comprehensive proportions in the first proportion sequence, denoted as the general proportion sequence. Then, denote the remaining comprehensive proportions in the first proportion sequence as the important proportion sequence; in this embodiment, k4% = k3% = 30%, that is, the first 30% is the core proportion sequence, the middle 40% is the important proportion sequence, and the last 30% is the general proportion sequence;
[0099] Sequentially mark the course professional terms corresponding to the core proportion sequence, the important proportion sequence, and the general proportion sequence as core knowledge terms, important knowledge terms, and general knowledge terms respectively; and obtain the core knowledge terms, important knowledge terms, general knowledge terms, and their corresponding word explanations, denoted as important knowledge point information;
[0100] The subtitle text completely presents the explanatory content in the course video, and the proportion of course professional terms in it directly reflects the basic status and core degree of these terms; the comment text is the subjective feedback of students on the course content, and the proportion of course professional terms in the comments reflects the degree of attention and discussion heat of students on these knowledge points; comprehensively considering the proportion of course professional terms in the subtitle word text and the comment word text, and obtaining the comprehensive proportion through weighted summation to divide the importance, this method can more objectively reflect the importance of professional terms in the course content and students' discussions;
[0101] Respectively obtain the averages of the secondary sequence and the tertiary sequence, denoted as PU1 and PU2 in sequence, and calculate PU0, PU0 = q3 * PU1 + q4 * PU2, where q3 and q4 are set weight coefficients, q3 + q4 = 1, 0 < q3, 0 < q4; in this embodiment, q3 = 0.6, q4 = 0.4; the secondary sequence and the tertiary sequence respectively represent users with good grades and average grades; PU0 is used to judge the difficulty of the course video in the follow-up. The reason for choosing the data of users with good grades and average grades for judgment is that the user groups with good grades and average grades usually account for the majority of the total number of students, and their scores are more sensitive to changes in difficulty, which can improve the accuracy of subsequent judgments;
[0102] Please refer to Figure 3 As shown, obtain the course professional terms in the difficult subtitle text corresponding to first-level users, second-level users, third-level users, and fourth-level users respectively, and denote them as first-level word text, second-level word text, third-level word text, and fourth-level word text in sequence;
[0103] Perform duplicate filtering on the first-level word text, second-level word text, third-level word text, and fourth-level word text in sequence, including: for any course major word, if it exists in the first-level word text, remove the same course major words in the second-level word text, third-level word text, and fourth-level word text; if it exists in the second-level word text, remove the same course major words in the third-level word text and fourth-level word text; if it exists in the third-level word text, remove the same course major words in the third-level and fourth-level word texts; after completion, obtain the first-level major text, second-level major text, third-level major text, and fourth-level major text in sequence; that is, ensure that the same course major word only exists in one text.
[0104] Set score thresholds PV1 and PV2, where PV2 > PV1; mark the course major words in the fourth-level major text as simple knowledge words; PV1 and PV2 can be set according to the actual application scenario.
[0105] If PV1 < PU0 < PV2, mark the course major words in the first-level major text as difficult knowledge words, and mark the course major words in the second-level major text and third-level major text as normal knowledge words.
[0106] If PV2 ≤ PU0, mark the course major words in the first-level major text as difficult knowledge words, mark the course major words in the second-level major text as normal knowledge words, and mark the course major words in the third-level major text as simple knowledge words.
[0107] If PU0 ≤ PV1, mark the course major words in the first-level major text and second-level major text as difficult knowledge words, and mark the course major words in the third-level major text as normal knowledge words.
[0108] Because there are differences in the difficulty levels of different course videos themselves, most of the course major words in some course videos may be relatively simple, while in others they may be relatively complex. If the overall difficulty level of the course videos is not judged first and the difficulty levels of the major words are directly divided, the difficulty differences between courses may be ignored; for example, for a professional course that is inherently difficult, some of the major words in it may generally be of a certain difficulty level for students, but if the overall course difficulty is not considered, these words may be wrongly classified into a lower difficulty level; by first judging the difficulty level of the course videos and then combining the viewing behavior of students to divide the difficulty levels of the major words, the difficulty differences between courses can be fully considered, and the major words can be classified more reasonably.
[0109] Obtain difficult knowledge words, normal knowledge words, simple knowledge words, and their corresponding word explanations, denoted as difficulty knowledge point information.
[0110] Merge the difficult and easy knowledge point information and the important knowledge point information according to the corresponding course professional terms and store them, which is recorded as the course hierarchical knowledge point data;
[0111] In the specific implementation process, by analyzing the behaviors of slow playback, rewind, and pause of users with different performance levels when watching videos, the division of the difficulty levels of course professional terms is refined; these behaviors can intuitively reflect the users' understanding and acceptance of specific knowledge points during the learning process; for example, if users with excellent grades slow down, rewind, or pause the content corresponding to certain course professional terms, it means that these terms are relatively difficult for them to understand; thus, they can be classified as difficult knowledge terms.
[0112] Example 2, please refer to Figure 2 As shown, the present application provides a method for automatically extracting online teaching information based on cloud computing, including the following steps:
[0113] Step S1, collect the video subtitle file of the course video and the comment data of the course video to obtain the course-related text data; Step S1 includes the following sub-steps:
[0114] Step S101, for any course video in the online teaching process, denoted as the first course video; obtain the video subtitle file of the first course video and convert it into an editable text file format, marked as the course subtitle text;
[0115] Step S102, obtain the comments of users under the first course video, including: for any comment under the first course video, denoted as the first comment, obtain and record the text content of the first comment and the user ID of the user who posted the first comment, repeat to obtain and record all comments under the first course video, and convert them into the same text file format as the course subtitle text, denoted as the course comment text;
[0116] Step S103, denote the course subtitle text and the course comment text as the course-related text data, and denote the online course to which the first course video belongs as the first course.
[0117] Step S2, perform keyword extraction based on the course-related text data to obtain the course-related knowledge point data; Step S2 includes the following sub-steps:
[0118] Step S201, denote any comment in the course comment text as the second comment, and denote the user ID to which the second comment belongs as the second user. If there is a comment in all the comments posted by the second user that is exactly the same as the second comment in text content, then mark the second comment as the first repeated comment and remove it. Process all comments in the course comment text in turn. After completion, obtain the preliminary screened comment text;
[0119] Step S202, remove special characters from the course subtitle text and the preliminary screening review text respectively. After completion, obtain the first subtitle text and the first review text in sequence;
[0120] Step S203, use a word segmentation tool to divide continuous text sentences in the first subtitle text and the first review text into multiple independent words respectively, and use a part-of-speech tagging tool to tag the part-of-speech of the divided words in the corresponding text sentences. After completion, obtain the second subtitle text and the second review text in sequence;
[0121] Step S204, set stop words according to common and meaningless words, and establish a stop word list. Traverse each word in the second subtitle text and the second review text respectively. For any word, if there is the same word in the stop word list, remove it. After completion, obtain the third subtitle text and the third review text;
[0122] Step S205, record any comment in the third review text as the third comment, and record the user ID to which the third comment belongs as the third user. If there is a comment in all comments corresponding to the third user in the third review text that is exactly the same as the third comment in terms of text content, mark the third comment as the second duplicate comment and remove it. Process all comments in the third review text in turn. After completion, obtain the re-screened review text;
[0123] Step S206, traverse each word in the third subtitle text and the re-screened review text, and record the professional words related to the first course in the third subtitle text and the re-screened review text as course professional words;
[0124] Step S207, repeatedly obtain all course professional words in the third subtitle text and the re-screened review text respectively, and record them as the subtitle word text and the review word text in sequence;
[0125] Step S208, obtain the total number of course professional words in the subtitle word text and the re-screened review text respectively, and record them as AZ and BZ in sequence; obtain the number of occurrences of any course professional word in the subtitle word text, and record them as AC1, AC2,..., ACn respectively; obtain the number of occurrences of any course professional word in the subtitle word text, and record them as BC1, BC2,..., BCm respectively;
[0126] Step S209, record the subtitle word text and the review word text as course-related knowledge point data.
[0127] Step S3, obtain the playback-related data of historical users learning course videos, and perform classification and screening processing on historical users to obtain relevant user data; Step S3 includes the following sub-steps:
[0128] Step S301: For any historical user of the first course video, denoted as the first historical user, obtain the cumulative duration of the first historical user watching the first course video, denoted as the first learning duration.
[0129] Step S302: Also obtain the video segments of slow playback, rewind, and pause when the first historical user watches the first course video, denoted as the difficult course segments, and obtain the subtitle text corresponding to the difficult course segments according to the course subtitle text, denoted as the difficult subtitle text.
[0130] Step S303: Then obtain the percentage of the score of the first historical user in the course quiz corresponding to the first course video in the total score, denoted as the first quiz score.
[0131] Step S303: Mark the first learning duration, the difficult course segments, the difficult subtitle text, and the first quiz score as the learning-related data of the first historical user in the first course video.
[0132] Step S304: Repeatedly obtain the learning-related data of multiple historical users in the first course video, and classify them according to the corresponding historical users, denoted as the playback-related data of the first course video.
[0133] Step S305: Arrange the first learning durations of multiple historical users in ascending order, and remove the first learning durations that are less than the duration of the first course video. After completion, obtain the first duration sequence; Set the first threshold ratio as k1%, and remove the largest k1% of the first learning durations in the first duration sequence to obtain the second duration sequence.
[0134] Step S306: Obtain the first quiz score of the historical user to which any first learning duration in the second duration sequence belongs; And sort them according to the order of the users corresponding to the second duration sequence, denoted as the first score sequence.
[0135] Step S307: For any corresponding historical user in the second duration sequence, denoted as the second user, if the second user meets any one of the screening conditions, then remove the data of the second user in the second duration sequence and the first score sequence; The screening conditions include: Condition 1, the first learning duration of the second user is equal to the duration of the first course video, and the first quiz score of the second user is equal to 100%; Condition 2, the first quiz score of the second user is less than k2%; Condition 3, the first learning duration of the second user is greater than k4; where k2% is the set second threshold ratio and k4 is the set threshold duration; After completion, obtain the third duration sequence and the second score sequence.
[0136] Step S308: Arrange the second score sequence in descending order, denoted as the third score sequence; divide the third score sequence according to the ratio of p1:p2:p3:p4, and obtain the first-level sequence, second-level sequence, third-level sequence, and fourth-level sequence in sequence; mark the historical users corresponding to the first-level sequence, second-level sequence, third-level sequence, and fourth-level sequence as first-level users, second-level users, third-level users, and fourth-level users respectively.
[0137] Step S4: Perform a grading process on the course-related knowledge point data based on relevant user data to obtain the course-graded knowledge point data; Step S4 includes the following sub-steps:
[0138] Step S401: Calculate the proportions of AC1, AC2, ……, ACn in AZ respectively, marked as the subtitle proportion, denoted as AD1, AD2, ……, ADn in sequence; and calculate the proportions of BC1, BC2, ……, BCm in BZ respectively, marked as the comment proportion, denoted as BD1, BD2, ……, BDn in sequence;
[0139] Step S402: Set the subtitle weight q1 and the comment weight q2, where q1 + q2 = 1, 0 < q1, 0 < q2;
[0140] Step S403: For any course professional term in the subtitle word text and the comment word text, denoted as the first professional term, obtain the subtitle proportion and the comment proportion corresponding to the first professional term, denoted as ADi and BDi in sequence; and calculate the comprehensive proportion corresponding to the first professional term through the comprehensive proportion formula. The comprehensive proportion formula is as follows: ; where ABi represents the comprehensive proportion of the first professional term pair;
[0141] Step S404: Repeatedly obtain the comprehensive proportions of all course professional terms in the subtitle word text and the comment word text, and arrange them in descending order, denoted as the first proportion sequence; obtain the largest k3% of the comprehensive proportions in the first proportion sequence, denoted as the core proportion sequence, and obtain the smallest k4% of the comprehensive proportions in the first proportion sequence, denoted as the general proportion sequence. Then, denote the remaining comprehensive proportions in the first proportion sequence as the important proportion sequence;
[0142] Step S405: Mark the course professional terms corresponding to the core proportion sequence, important proportion sequence, and general proportion sequence as core knowledge terms, important knowledge terms, and general knowledge terms respectively; and obtain the core knowledge terms, important knowledge terms, general knowledge terms, and their corresponding word explanations, denoted as important knowledge point information;
[0143] Step S406: Obtain the averages of the secondary sequence and the tertiary sequence respectively, denoted as PU1 and PU2 in sequence, and calculate PU0, where PU0 = q3 * PU1 + q4 * PU2, with q3 and q4 being set weight coefficients, q3 + q4 = 1, 0 < q3, 0 < q4;
[0144] Step S407: Obtain the course-specific terms in the difficult subtitle texts corresponding to the primary users, secondary users, tertiary users, and quaternary users respectively, denoted as the primary term text, secondary term text, tertiary term text, and quaternary term text in sequence;
[0145] Step S408: Perform duplicate filtering on the primary term text, secondary term text, tertiary term text, and quaternary term text in sequence, including: for any course-specific term, if it exists in the primary term text, remove the same course-specific terms in the secondary term text, tertiary term text, and quaternary term text; if it exists in the secondary term text, remove the same course-specific terms in the tertiary term text and quaternary term text; if it exists in the tertiary term text, remove the same course-specific terms in the tertiary and quaternary term texts; after completion, obtain the primary professional text, secondary professional text, tertiary professional text, and quaternary professional text in sequence;
[0146] Step S409: Set score thresholds PV1 and PV2, where PV2 > PV1; mark the course-specific terms in the quaternary professional text as simple knowledge terms;
[0147] Step S410: If PV1 < PU0 < PV2, mark the course-specific terms in the primary professional text as difficult knowledge terms, and mark the course-specific terms in the secondary professional text and tertiary professional text as normal knowledge terms;
[0148] Step S411: If PV2 ≤ PU0, mark the course-specific terms in the primary professional text as difficult knowledge terms, mark the course-specific terms in the secondary professional text as normal knowledge terms, and mark the course-specific terms in the tertiary professional text as simple knowledge terms;
[0149] Step S412: If PU0 ≤ PV1, mark the course-specific terms in the primary professional text and secondary professional text as difficult knowledge terms, and mark the course-specific terms in the tertiary professional text as normal knowledge terms;
[0150] Step S413: Obtain the difficult knowledge terms, normal knowledge terms, simple knowledge terms, and their corresponding word explanations, denoted as the difficult and easy knowledge point information;
[0151] Step S414: Merge and store the easy and difficult knowledge point information and the important knowledge point information according to the corresponding course professional terms, and record it as the course hierarchical knowledge point data.
[0152] Example 3, please refer to Figure 4 as shown in Figure 4 A schematic structural diagram of an electronic device is illustrated. The electronic device may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The memory stores computer-readable instructions, and the processor can call the instructions in the memory. When the computer-readable instructions are executed by the processor, the steps in the method for automatically extracting network teaching information based on cloud computing are run to implement the following functions: collecting the video subtitle file of the course video and the comment data of the course video to obtain the course-related text data; extracting keywords based on the course-related text data to obtain the course-related knowledge point data; obtaining the playback-related data of the historical users learning the course video, and performing classification and screening processing on the historical users to obtain the relevant user data; performing a hierarchical classification process on the course-related knowledge point data based on the relevant user data to obtain the course hierarchical knowledge point data.
[0153] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0154] Embodiment 4. The present application further provides a computer-readable storage medium. The present application provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method for automatically extracting network teaching information based on cloud computing are run to implement the following functions: collecting the video subtitle file of the course video and the comment data of the course video to obtain course-related text data; extracting keywords based on the course-related text data to obtain course-related knowledge point data; obtaining the playback-related data of historical users learning the course video, and performing classification and screening processing on the historical users to obtain relevant user data; performing level division processing on the course-related knowledge point data based on the relevant user data to obtain course-level knowledge point data.
[0155] Through the description of the above embodiments, the embodiments of the present invention can be provided as a method, a system or a computer program product. Based on such an understanding, the above technical solution, in essence, or the part that makes a contribution to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0156] In the embodiments provided by the present application, it should be understood that the disclosed system or method can be implemented in other ways. The above-described embodiments are merely illustrative. For example, the division of modules or units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some communication interfaces. The indirect coupling or communication connection of systems, modules, and units can be in an electrical, mechanical, or other form.
[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An automatic extraction system for network teaching information based on cloud computing, characterized in that, It includes a text collection module, a text processing module, a learning collection module, and a knowledge extraction module; The text collection module is used to collect the video subtitle files of the course videos and the comment data of the course videos to obtain course-related text data; The text processing module extracts keywords based on the course-related text data to obtain course-related knowledge point data; The learning collection module includes a data acquisition unit and a user classification unit. The data acquisition unit is used to acquire the playback-related data of historical users learning course videos, and the user classification unit is used to classify and screen historical users to obtain relevant user data; The knowledge extraction module performs a grading process on the course-related knowledge point data based on the relevant user data to obtain course-graded knowledge point data; The knowledge extraction module is configured with a knowledge extraction strategy, and the knowledge extraction strategy includes: Respectively obtain the total number of course professional words in the subtitle text and the re-screened comment text, and record them as AZ and BZ in sequence; obtain the number of any course professional word in the subtitle text, and record them as AC1, AC2,..., ACn respectively; obtain the number of any course professional word in the subtitle text, and record them as BC1, BC2,..., BCm respectively; calculate the proportions of AC1, AC2,..., ACn in AZ respectively, marked as subtitle proportions, and record them as AD1, AD2,..., ADn in sequence; and calculate the proportions of BC1, BC2,..., BCm in BZ respectively, marked as comment proportions, and record them as BD1, BD2,..., BDn in sequence; Set the subtitle weight q1 and the comment weight q2, q1 + q2 = 1, 0 < q1, 0 < q2; For any course professional term in the subtitle word text and the comment word text, denoted as the first professional term, obtain the subtitle proportion and the comment proportion corresponding to the first professional term, denoted as ADi and BDi in sequence; and calculate the comprehensive proportion corresponding to the first professional term through the comprehensive proportion formula. The comprehensive proportion formula is as follows: ; where ABi represents the comprehensive proportion of the first professional term pair; Repeatedly obtain the comprehensive proportions of all course professional words in the subtitle text and the comment text, and arrange them in descending order, recorded as the first proportion sequence; obtain the largest k3% of the comprehensive proportions in the first proportion sequence, recorded as the core proportion sequence, and obtain the smallest k4% of the comprehensive proportions in the first proportion sequence, recorded as the general proportion sequence, and then record the remaining comprehensive proportions in the first proportion sequence as the important proportion sequence; Sequentially mark the course professional words corresponding to the core proportion sequence, the important proportion sequence, and the general proportion sequence as core knowledge words, important knowledge words, and general knowledge words respectively; and obtain the core knowledge words, important knowledge words, general knowledge words, and the corresponding word explanations, recorded as important knowledge point information; The knowledge extraction strategy also includes: Respectively obtain the averages of the secondary sequence and the tertiary sequence, and record them as PU1 and PU2 in sequence, and calculate PU0, PU0 = q3 * PU1 + q4 * PU2, where q3 and q4 are set weight coefficients, q3 + q4 = 1, 0 < q3, 0 < q4; Obtain the course professional words in the difficult subtitle text corresponding to the first-level users, second-level users, third-level users, and fourth-level users respectively, and record them as the first-level word text, second-level word text, third-level word text, and fourth-level word text in sequence; Perform duplicate filtering on the first-level word text, second-level word text, third-level word text, and fourth-level word text in sequence, including: for any course major word, if it exists in the first-level word text, remove the same course major words in the second-level word text, third-level word text, and fourth-level word text; if it exists in the second-level word text, remove the same course major words in the third-level word text and fourth-level word text; if it exists in the third-level word text, remove the same course major words in the third-level and fourth-level word texts; after completion, obtain the first-level professional text, second-level professional text, third-level professional text, and fourth-level professional text in sequence; Set score thresholds PV1 and PV2, where PV2 > PV1; mark the course major words in the fourth-level professional text as simple knowledge words; If PV1 < PU0 < PV2, mark the course major words in the first-level professional text as difficult knowledge words, and mark the course major words in the second-level professional text and third-level professional text as normal knowledge words; If PV2 ≤ PU0, mark the course major words in the first-level professional text as difficult knowledge words, mark the course major words in the second-level professional text as normal knowledge words, and mark the course major words in the third-level professional text as simple knowledge words; If PU0 ≤ PV1, mark the course major words in the first-level professional text and second-level professional text as difficult knowledge words, and mark the course major words in the third-level professional text as normal knowledge words; Obtain difficult knowledge words, normal knowledge words, simple knowledge words, and their corresponding word explanations, denoted as difficult and easy knowledge point information; Merge and store the difficult and easy knowledge point information and the important knowledge point information according to the corresponding course major words, denoted as course classification knowledge point data.
2. The automatic extraction system for network teaching information based on cloud computing according to claim 1, wherein The text collection module is configured with a text collection strategy, and the text collection strategy includes: For any course video in the online teaching process, denoted as the first course video; obtain the video subtitle file of the first course video and convert it into an editable text file format, marked as the course subtitle text; Obtain the comments of the user under the first course video, including: for any comment under the first course video, denoted as the first comment, obtain and record the text content of the first comment and the user ID that posted the first comment, repeat to obtain and record all comments under the first course video, and convert them into the same text file format as the course subtitle text, denoted as the course comment text; Denote the course subtitle text and the course comment text as course-related text data, and denote the online course to which the first course video belongs as the first course.
3. The automatic extraction system for network teaching information based on cloud computing according to claim 2, wherein The text processing module is configured with a text processing strategy, and the text processing strategy includes: Denote any comment in the course comment text as the second comment, and denote the user ID to which the second comment belongs as the second user. If there is a comment in all the comments posted by the second user that is exactly the same as the second comment in text content, then mark the second comment as the first duplicate comment and remove it. Process all comments in the course comment text in sequence. After completion, obtain the preliminary screened comment text; Remove special characters from the course subtitle text and the initially screened review text respectively. After completion, obtain the first subtitle text and the first review text in sequence; Use a word segmentation tool to divide the continuous text sentences in the first subtitle text and the first review text into multiple independent words respectively, and use a part-of-speech tagging tool to tag the part of speech of the divided words in the corresponding text sentences. After completion, obtain the second subtitle text and the second review text in sequence.
4. The automatic extraction system for network teaching information based on cloud computing according to claim 3, characterized in that, The text processing strategy also includes: Set stop words and establish a stop word list. Traverse each word in the second subtitle text and the second review text respectively. For any word, if there is the same word in the stop word list, remove it. After completion, obtain the third subtitle text and the third review text; Denote any review in the third review text as the third review, and denote the user ID to which the third review belongs as the third user. If there is a review in all the reviews corresponding to the third user in the third review text that is exactly the same as the third review in terms of text content, mark the third review as the second repeated review and remove it. Process all the reviews in the third review text in turn. After completion, obtain the re-screened review text.
5. The automatic extraction system for network teaching information based on cloud computing according to claim 4, wherein, The text processing strategy also includes: Traverse each word in the third subtitle text and the re-screened review text, and denote the professional words related to the first course in the third subtitle text and the re-screened review text as course professional words; Repeat to obtain all the course professional words in the third subtitle text and the re-screened review text respectively, and denote them as the subtitle word text and the review word text in sequence; Denote the subtitle word text and the review word text as the course-related knowledge point data.
6. The automatic extraction system for network teaching information based on cloud computing according to claim 5, characterized in that The data acquisition unit is configured with a data acquisition strategy, and the data acquisition strategy includes: For any historical user of the first course video, denote it as the first historical user, and obtain the cumulative duration of the first historical user watching the first course video, denoted as the first learning duration; And obtain the video segments of slow playback, rewind, and pause when the first historical user watches the first course video, denoted as the difficult course segments, and obtain the subtitle text corresponding to the difficult course segments according to the course subtitle text, denoted as the difficult subtitle text; Then obtain the percentage of the score of the first historical user in the course quiz corresponding to the first course video in the total score, denoted as the first quiz score; Mark the first learning duration, the difficult course segments, the difficult subtitle text, and the first quiz score as the learning-related data of the first historical user in the first course video; Repeat to obtain the learning-related data of multiple historical users in the first course video, and classify them according to the corresponding historical users, denoted as the playback-related data of the first course video.
7. The automatic extraction system for network teaching information based on cloud computing according to claim 6, characterized in that The user classification unit is configured with a user classification strategy, and the user classification strategy includes: Arrange the first learning durations of multiple historical users in ascending order, and remove the first learning durations that are less than the duration of the first course video. After completion, obtain the first duration sequence; Set the first threshold ratio as k1%, and remove the largest k1% of the first learning durations in the first duration sequence to obtain the second duration sequence; Obtain the first quiz score of the historical user to which any one of the first learning durations in the second duration sequence belongs; and sort them in the order of the users corresponding to the second duration sequence, denoted as the first score sequence; For any historical user corresponding to the second duration sequence, denoted as the second user, if the second user meets any one of the screening conditions, then remove the data of the second user corresponding in the second duration sequence and the first score sequence; The screening conditions include: Condition 1, the first learning duration of the second user is equal to the first course video duration, and the first quiz score of the second user is equal to 100%; Condition 2, the first quiz score of the second user is less than k2%; Condition 3, the first learning duration of the second user is greater than k5; where k2% is the set second threshold ratio and k5 is the set threshold duration; After completion, obtain the third duration sequence and the second score sequence; Arrange the second score sequence in descending order, denoted as the third score sequence; Divide the third score sequence according to the ratio of p1:p2:p3:p4, and obtain the first-level sequence, second-level sequence, third-level sequence and fourth-level sequence in order; Mark the historical users corresponding to the first-level sequence, second-level sequence, third-level sequence and fourth-level sequence as first-level users, second-level users, third-level users and fourth-level users in order.
8. A method for automatically extracting online teaching information based on cloud computing, which is used to implement the system for automatically extracting online teaching information based on cloud computing according to any one of claims 1-7, characterized in that, Include the following steps: Collect the video subtitle file of the course video and the comment data of the course video to obtain the course-related text data; Extract keywords based on the course-related text data to obtain the course-related knowledge point data; Obtain the playback-related data of historical users learning the course video, and perform classification and screening processing on historical users to obtain relevant user data; Based on the relevant user data, perform hierarchical classification processing on the course-related knowledge point data to obtain the course hierarchical knowledge point data.
Citation Information
Patent Citations
Automatic knowledge extraction method for network teaching system
CN113641716A
Language learning data recommendation method and device, electronic equipment and readable storage medium
CN118673221A
Learning recommendation method and device for occupational development and storage medium
CN118966858A
Micro course generation method based on OCR (Optical Character Recognition) and server
CN119580289A
Method for establishing knowledge repository for online courses
WO2023010514A1