Teaching material library similar resource extension method based on network resource capture
The method addresses the limitations of traditional educational resource management by using network crawlers to construct semantic graphs and personalize recommendations, enhancing educational content relevance and quality.
Patent Information
- Application Number
- CN202510806518.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing textbook library expansion method relies on manual updates and static resource classification, resulting in poor synchronization of textbook content and educational needs, insufficient personalized recommendations, and affecting the coverage of educational resources and teaching effect.
Through network resource capture technology, a knowledge point-level path structure set, semantic connection expansion map, and path node semantic matching resource list are constructed, and the core knowledge points in the textbook are accurately extracted and utilized to realize dynamic updates and personalized recommendations of textbook content.
It improves the discovery and accessibility of textbook resources, enhances the depth and breadth of textbook content, enriches the learning experience, and improves the practicality and quality of educational resources.
Smart Images

Figure CN120318043A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of textbook library expansion, and particularly to a method for expanding similar resources in a textbook library based on web resource scraping. Background Art
[0002] The technical field of textbook library expansion includes methods for managing and enhancing educational resources, focusing on how to effectively expand the textbook resource library to support teaching and learning. The core content of this technical field is to achieve dynamic update and enrichment of textbook content, ensuring that textbook resources keep pace with current educational needs and technological progress. In the systematic introduction, it involves the classification, storage, update, and personalized recommendation of textbook resources, as well as how to use information technology to improve the access efficiency and usage effect of the textbook library.
[0003] Among them, the method for expanding similar resources in a textbook library by web resource scraping refers to using web crawler technology to automatically collect textbook-related resources from the Internet. The technical matters targeted by this patent theme cover scraping textbook content from the web, analyzing the relevance and applicability of textbook resources, and integrating the resources into the existing textbook library. Specifically, it automatically accesses education-related websites through a web crawler, retrieves new resources that match the existing textbook content, screens and classifies the resources according to preset criteria, and updates them into the textbook library for teachers and students to use.
[0004] Although the prior art covers the management and enhancement of textbook resources, it has limitations in dynamically updating textbook content and matching current educational needs. Traditional methods for expanding textbook libraries rely on manual updates and static resource classification, which limit the timely synchronization and in-depth integration between textbook content and educational needs. The prior art also shows deficiencies in personalized recommendation and resource utilization efficiency, such as the lack of an effective mechanism to identify and recommend textbook content most relevant to learners' needs. These deficiencies result in the inability of textbook content to fully meet the diverse and developmental needs of students, affecting the coverage of educational resources and the personalized development of teaching. The access efficiency and usage effect of the textbook library are restricted, resulting in the potential value of learning resources not being fully explored, affecting the improvement of educational quality. Summary of the Invention
[0005] The purpose of the present invention is to solve the deficiencies existing in the prior art, and to propose a method for expanding similar resources in a textbook library based on web resource scraping.
[0006] To achieve the above purpose, the present invention adopts the following technical solution. A method for expanding similar resources in a textbook library based on web resource scraping includes the following steps: S1: Obtain the textbook chapter text, paragraph structure table, and glossary. Compare the positions of chapter titles and term definition paragraphs, extract the concept frequencies and context positions, and construct the concept attribution hierarchical paths based on the concept co-occurrence frequencies and paragraph order to generate a set of knowledge point hierarchical path structures; S2: According to the arrangement order of path nodes in the set of knowledge point hierarchical path structures, collect the co-occurrence sentence blocks and reference sentences between terms in the textbook, and call the up and down reference frequencies, sentence block distance positions, and co-occurrence ranges of each pair of nodes to generate a knowledge point semantic connection expansion graph; S3: Based on the node term groups in the knowledge point semantic connection expansion graph, obtain the text resources of educational resource websites, and judge whether the number of graph term nodes covered by the keywords and the term distribution positions cover the path starting point and extended nodes to generate a list of path node semantic matching resources; S4: According to the list of path node semantic matching resources, call the step information of students' answering questions and the formula call order, compare with the problem-solving path structure corresponding to the questions in the standard textbook, and judge whether there are missing nodes, misplaced orders, or incorrect replacement operations in the path to generate a set of knowledge point path error positionings;
[0007] As a further solution of the present invention, the set of knowledge point hierarchical path structures includes a concept attribution path table, a knowledge point theme index library, and a term hierarchical identification group. The knowledge point semantic connection expansion graph includes a term reference network, a node coverage range graph, and a semantic extension direction table. The list of path node semantic matching resources includes term matching resource entries, a path coverage comparison table, and a node distribution position set. The set of knowledge point path error positionings includes a problem-solving deviation node list, a path offset number group, and an abnormal knowledge point label set.
[0008] As a further solution of the present invention, the specific steps for obtaining the set of knowledge point hierarchical path structures are as follows: S111: Obtain the textbook chapter text, paragraph structure table, and glossary. Extract the starting positions of paragraph titles, the paragraph numbers and sentence positions of terms, calculate the occurrence frequencies and position distribution intervals of terms within the chapter, and analyze the term density of the term intervals covered by the titles and the subordinate paragraphs to generate a chapter term density distribution; S112: According to the dense section numbers in the chapter term density distribution, call the defined positions, term co-occurrence pairs, and the number of sentences in the paraphrase paragraphs in the glossary, and judge the sequential positions of co-occurring term pairs within the same section to generate a term order path combination table; S113: Call the term group path structures in the term order path combination table, compare the chapter indexes, paragraph numbers, and term reference quantities in the paragraph structure table, filter the cross-paragraph paths of consecutive term nodes in the path, and record the chapter positions to which they belong to generate a set of knowledge point hierarchical path structures.
[0009] As a further solution of the present invention, the steps for obtaining the knowledge point semantic connection expansion graph are specifically as follows: S211: Using the arrangement order of multiple path nodes in the knowledge point hierarchical path structure, correspondingly search for the common sentence block areas, term definition segment numbers, and term reference statement positions of term pairs in the textbook, call the definition segment number to compare the front and back definition orders of the terms, and generate the term co-occurrence sentence block difference value; S212: Based on the term co-occurrence sentence block difference value, call the definition positions and citation frequencies of the upper and lower term nodes in the term pair in the textbook, obtain the paragraph number sequence of the term reference statement and the paragraph number range of the co-occurrence range, make a directional judgment according to whether the upper term in the term pair appears before the lower term reference position, calculate the directional offset intensity value, and sort and screen the node direction connection relationship according to the offset intensity to generate the term node direction connection degree; S213: Call the term pair structure in the term node direction connection degree, assemble the paths with continuous connection relationships into a directed path graph structure, remove the path paragraph numbers with path jumps exceeding the chapter span, and establish the knowledge point semantic connection expansion graph.
[0010] As a further solution of the present invention, the steps for obtaining the path node semantic matching resource list are specifically as follows: S311: Based on the node term group in the knowledge point semantic connection expansion graph, obtain the text resources of educational resource websites, and extract the paragraph keywords, topic titles, and text sentence blocks in the text to generate a key element set of text resources; S312: Call the key element set of the text resources, extract the keyword associated node values for the number of graph term nodes covered in the paragraph keywords, mark the distribution positions of the starting node and extended nodes of the graph corresponding to the keywords, calculate the keyword node matching degree, and screen the keywords that cover the starting node and extend to the node to obtain the keyword node coverage distribution index set; S313: Based on the keyword node coverage distribution index set, screen the keyword content with the node coverage index reaching the set node coverage benchmark value, and combine the corresponding topic titles and text sentence blocks to generate the path node semantic matching resource list.
[0011] As a further solution of the present invention, the steps for obtaining the knowledge point path error positioning set are specifically as follows: S411: Based on the semantic similarity between the node information in the path node semantic matching resource list, extract the operation steps in the student's answering process, number the node order, call the numbered order and the path node order corresponding to the questions in the standard textbook for difference matching, obtain the numbered information corresponding to node missing, order misplacement, and substitution operations, and generate an abnormal node number set; S412: Use the set of abnormal node numbers, the total number of path structure nodes in the standard textbook, and the node order relationship, combine multiple node order comparison items and structure matching items in the abnormal node numbers to calculate the path sequence structure difference degree, compare it with the path structure deviation judgment threshold, obtain the abnormal node sequence that meets the deviation threshold condition, and get the path abnormal structure positioning number set; S413: Call the path abnormal structure positioning number set, match the node number and the classification path number according to the knowledge point node number table and the knowledge point path classification table in the textbook, extract the corresponding knowledge point labels, and generate a knowledge point path misunderstanding positioning set.
[0012] As a further solution of the present invention, the method further includes step S5: S5: Call the knowledge point path misunderstanding positioning set, match the textbook chapter and the explanation paragraph, extract the core explanation sentences and paragraphs in the chapter where the path node is located, check the explanation length of the matching sentences and the occurrence frequency of examples, and generate a recommended sequence of similar textbook resources; The recommended sequence of similar textbook resources includes matching textbook fragment indexes, resource priority identification lists, and example coverage density annotation sets.
[0013] As a further solution of the present invention, the steps for obtaining the recommended sequence of similar textbook resources are specifically as follows: S511: Call the knowledge point path misunderstanding positioning set, perform character matching on the knowledge point keywords marked by the path nodes and the textbook content chapter titles, extract the textbook chapter numbers and title information where the path nodes are located, and obtain the corresponding list of path node chapters; S512: Based on the corresponding list of path node chapters, locate all the explanation paragraphs in the chapter corresponding to the path nodes, extract the explanation paragraphs containing the path node keywords, call all the sentence contents in the explanation paragraphs, calculate the total number of words in the sentences and the length of the path node matching sentences, and obtain the path node explanation density distribution; S513: According to the path node explanation density distribution, screen the chapter paragraphs with density values exceeding the set explanation density threshold, extract the example content of the path node keywords, and count the total frequency of examples in the chapter paragraphs to generate a recommended sequence of similar textbook resources.
[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In the present invention, by acquiring and analyzing teaching materials content, extracting key information, and combining big data and artificial intelligence technologies, a multi-dimensional knowledge structure system is effectively constructed. Using text analysis and semantic understanding technologies, a hierarchical path of knowledge points and a semantic connection map are constructed, which not only enhances the structured information of the teaching material library but also improves the discoverability and accessibility of resources. Through meticulous comparison of chapters and terms, the core knowledge points in the teaching materials can be accurately extracted and utilized, making the update and enrichment of teaching resources more targeted and systematic. The semantic matching resource list enables teachers and students to obtain additional resources closely related to the learning content. Such matching increases the depth and breadth of the teaching materials content and enriches the learning experience. By real-time updating and precisely matching the teaching materials content, the practicality of educational resources and the educational quality are effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a schematic diagram of the main steps of the present invention; Figure 2 is a flowchart for obtaining the hierarchical path structure set of knowledge points in the present invention; Figure 3 is a flowchart for obtaining the semantic connection expansion map of knowledge points in the present invention; Figure 4 is a flowchart for obtaining the semantic matching resource list of path nodes in the present invention; Figure 5 is a flowchart for obtaining the error location set of knowledge point paths in the present invention; Figure 6 is a flowchart for obtaining the recommended sequence of similar resources of teaching materials in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0017] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality" is two or more unless otherwise specifically defined.
[0018] Please refer to Figure 1, the present invention provides a technical solution, a method for expanding similar resources in a teaching material library based on web resource scraping, including the following steps: S1: Obtain the teaching material chapter text, paragraph structure table and glossary, compare the chapter title position with the term definition section, extract the concept frequency and context position, and construct a concept attribution hierarchical path based on the concept co-occurrence frequency and paragraph order to generate a set of knowledge point hierarchical path structures; S2: According to the arrangement order of the path nodes in the set of knowledge point hierarchical path structures, collect the co-occurrence sentence blocks, adjacent definition sections and reference sentences between terms in the teaching material, call the upper and lower reference frequencies, sentence block distance positions and co-occurrence ranges of each pair of nodes, calculate the semantic coverage range and the upper and lower position extension directions between concept pairs, and construct node extension connections guided by the upper and lower position directions to generate a knowledge point semantic connection expansion map; S3: Based on the node term groups in the knowledge point semantic connection expansion map, obtain the text resources of educational resource websites, extract paragraph keywords, topic titles and body sentence blocks, and judge whether the number of map term nodes covered in the keywords and the term distribution positions cover the path starting point and extended nodes to generate a list of path node semantic matching resources; S4: According to the list of path node semantic matching resources, call the step information of students' answering questions and the formula call order, compare with the problem-solving path structure corresponding to the questions in the standard teaching material, judge whether there are missing nodes, misplaced orders or wrong substitution operations in the path, locate the abnormal node numbers, and generate a set of knowledge point path error positionings; S5: Call the set of knowledge point path error positionings, match the teaching material chapters and explanation paragraphs, extract the core explanation sentence segments in the chapters where the path nodes are located, check the explanation length of the matching sentences and the example occurrence frequency, and generate a recommended sequence of similar resources in the teaching material; The set of knowledge point hierarchical path structures includes a concept attribution path table, a knowledge point theme index library, and a term hierarchical identification group. The knowledge point semantic connection expansion map includes a term reference network, a node coverage range map, and a semantic extension direction table. The list of path node semantic matching resources includes term matching resource entries, a path coverage comparison table, and a node distribution position set. The set of knowledge point path error positionings includes a list of problem-solving deviation nodes, a group of path offset numbers, and a set of abnormal knowledge point labels. The recommended sequence of similar resources in the teaching material includes a matching teaching material fragment index, a resource priority identification list, and an example coverage density annotation set.
[0019] Please refer to Figure 2 , the specific steps for obtaining the set of knowledge point hierarchical path structures are as follows: S111: Obtain the textbook chapter text, paragraph structure table, and glossary. Extract the starting positions of paragraph titles, the paragraph numbers and sentence positions of terms, calculate the occurrence frequency and position distribution range of terms within the chapter, analyze the term coverage range of the title and the term density of the subordinate paragraphs, and generate the chapter term density distribution. Split the original textbook text by chapter number, using the identifier starting with "Chapter X" as the chapter boundary. Extract the total number of paragraphs in each chapter and number them to form a paragraph structure table. At the same time, retrieve the positions where the terms appear in the text item by item based on the glossary content, mark their paragraph numbers and sentence position indexes. The sentence position indexes are numbered in the order of sentences within the paragraph to establish an index table. Further, count the total number of terms in each chapter, calculate the term density value in combination with the number of paragraphs. Suppose there are 15 paragraphs in Chapter 1 and the terms appear 58 times, then the density value is 58 / 15 = 3.87. Calculate and record the term density values for each chapter in this way. Based on the starting paragraph number of the paragraph where the title is located in the chapter, extract the sequence of term occurrence density values of the adjacent paragraphs after the title. Determine whether this paragraph belongs to the title coverage area by comparing the difference between the density value and the overall chapter term density mean. If the density of adjacent paragraphs continuously exceeds the chapter average term density and is greater than the set benchmark density threshold (such as 3.0), then this paragraph is classified as a "term-intensive section". Suppose Chapter 3 contains 10 paragraphs, the total number of terms appears 43 times, and the average density is 4.3. Then the paragraphs with a density higher than 4.3 in the chapter are candidate sections. If their density continuity reaches more than 3 paragraphs, they are confirmed as high-intensive blocks. Construct a distribution structure with the chapter number, intensive section number, and their term average density values to obtain the chapter term density distribution.
[0020] S112: According to the intensive section numbers in the chapter term density distribution, call the defined positions of terms, term co-occurrence pairs, and the number of sentences in the explanatory paragraphs in the glossary, judge the sequential positions of co-occurring term pairs within the same section, and generate a term sequential path combination table. Call the defined positions of terms, term co-occurrence pairs, and the number of sentences in the explanatory paragraphs. Extract the set of terms that appear in each intensive section to form a paragraph term group and perform co-occurrence pairing operations. For each pair of term pairs, record the sentence position index and relative interval distance of the starting term according to their co-occurrence position relationship in the same paragraph. Further, call the defined position numbers of these term pairs in the glossary and the number of sentences in the corresponding explanatory paragraphs. Suppose term A is defined in paragraph 6 with 5 sentences, and term B is defined in paragraph 7 with 4 sentences. If they both appear in the 2nd and 4th sentences of paragraph 6, then record their sentence position difference as 2, relative distance as 0.5, and calculate their co-occurrence rate. , then calculate whether the directional order of its sentence position shows an upward extension trend. If term A is in the front sentence position and B is in the back in the explanatory paragraph, the order is considered established; then, for the multiple occurrences of each term pair in this dense section, frequency merging processing is performed. For term pairs with the number of occurrences greater than 3, a connection weight of 1 is assigned, and for those less than or equal to 3, a connection weight of 0.5 is assigned. A term co-occurrence connection table is constructed and numbered and grouped according to the section number; further, it is judged whether the term pairs in the connection have semantic order consistency within the same section, that is, at least two groups of term pairs have the same order before and after the definition section and the order of appearance within the section, then it is confirmed as an order path node pair. For example, in the dense section of Chapter 2, the terms "particle", "velocity", and "acceleration" appear simultaneously. Among them, the definition section of "particle" is in Paragraph 5, "velocity" is in Paragraph 6, and "acceleration" is in Paragraph 7. If the three appear in the corresponding order in the paragraph, and each appears more than 3 times between any two terms, then the order path "particle - velocity - acceleration" is formed; the term node combination structures of each path are integrated according to the order path number to generate a term order path combination table.
[0021] S113: Call the term group path structure in the term order path combination table, compare the chapter index, paragraph number, and term reference quantity in the paragraph structure table, filter the cross-paragraph paths of consecutive term nodes in the path, and record the chapter position to which it belongs to generate a knowledge point hierarchical path structure set; Compare the chapter index, paragraph numbers, and the number of term references in the paragraph structure table. For each term path, sequentially obtain the paragraph numbers corresponding to the term nodes of the path, and establish a path paragraph index sequence. Analyze whether the path paragraphs have the characteristic of continuity, that is, determine whether the term nodes in the path appear sequentially in adjacent paragraphs. If the corresponding section numbers for the term path "particle - velocity - acceleration" are 5, 6, and 7 respectively, it is recorded as a continuous path segment index group "[5, 6, 7]". If there is a break in the middle (such as the path segments are 5, 6, 9), it is recorded as a discontinuous path, and the path type is determined to be a broken type. Path groups that do not meet the continuity requirements need to be excluded. The continuity judgment criterion is set such that the cross - segment difference of the path nodes does not exceed 1. After confirming the continuous path, it is necessary to determine whether its cross - segment span exceeds the maximum teaching segment range set for the chapter. If the path segment span is greater than 60% of the total number of paragraphs in the chapter (assuming a chapter has 10 paragraphs, then the path span should not exceed 6 paragraphs), it is determined to be a long - winded path and is also excluded. For the term paths that meet the paragraph continuity and span limit conditions, further obtain the number of term references in the paragraphs, and sum up the total number of times the nodes in the path appear in their respective paragraphs. If "particle" appears 4 times in section 5, "velocity" appears 3 times in section 6, and "acceleration" appears 5 times in section 7, then the total number of path term references is 12. Sort the paths that meet the requirements according to the total number of references, and record the chapter number to which each path belongs, the paragraph number range within the path, the node sequence, and the total sum of reference frequencies. Output the top N path structures from high to low to construct a knowledge point system tree, which reflects the extension order, paragraph association, and content density among the terms of each knowledge point in the textbook in terms of paths, providing path support for the construction of the subsequent knowledge point semantic map, and obtaining a knowledge point hierarchical path structure set.
[0022] Please refer to Figure 3 , and the steps for obtaining the knowledge point semantic connection expansion map are specifically as follows: S211: Using the arrangement order of the nodes of multiple paths in the knowledge point hierarchical path structure set, correspondingly search for the common sentence block areas, term definition section numbers, and term reference statement positions of the term pairs in the textbook, and call the definition section numbers to compare the front - and - back definition order of the terms to generate a term co - occurrence sentence block difference value; The term pairs are sequentially extracted according to the node sequence to form a term path pair combination. For each group of term pairs, the path "acceleration-force-mass" is set, and "acceleration-force" and "force-mass" are treated as independent term pairs. The sentence block area where the term pairs first co-occur is searched in the original textbook one by one, and the paragraph number and sentence number index in the paragraph are obtained at the same time. A sentence index table is established to record the position difference of the term appearance. Taking "acceleration" and "force" as an example, if the two terms first co-occur in the third and fifth sentences of the 12th paragraph, the interval is 2 sentences. The co-occurrence of the term pair is marked as "existent" and the sentence block difference is 2. Subsequently, the multiple co-occurrences of the same term pair in the entire textbook are traversed, and the sentence block number and sentence position difference are recorded. The number of sentence blocks in each group of term pairs is counted to obtain the co-occurrence frequency of the term pairs. The definition segment number of each term node in the textbook is continued to be extracted to determine whether the definition segment of term A is earlier than the definition segment of term B. If true, it is marked as "positive order", otherwise it is marked as "reverse order". The order of the positions of A and B in the term pair in the quoted sentence is processed in the same way. If the average sentence position of term A is less than the average sentence position of term B, it is recorded as a "continuation structure". If the number of times this structure appears in the same term pair is greater than or equal to 3, the order is considered to be the main order. It is assumed that the term "force" is defined in the 5th paragraph and the term "quality" is defined in the 8th paragraph. At the same time, they co-occur in the 10th, 12th, and 15th paragraphs respectively. Among them, the term "force" is in the 2nd sentence and the term "quality" is in the 4th sentence. The average difference is 2 sentences, and the total number of occurrences is 3 times, which is recorded as a valid sequence structure. The co-occurrence frequency, sentence block position difference and definition segment order of the term pair are combined to construct a term pair attribute vector, and a comparison is performed within the path. If the number of occurrences of the term pair is less than 2 times or the average sentence position difference exceeds 5 sentences, the term pair is excluded from the path sequence, and the retained term path pair structure is recorded as the path number, term pair number, co-occurrence sentence block position, definition order status, and sentence position average difference, which are used for subsequent directionality judgment and connection relationship calculation to generate term co-occurrence sentence block difference value.
[0023] S212: Based on the difference value of the term co-occurrence sentence block, call the definition position and citation frequency of the upper and lower term nodes in the term pair in the textbook, obtain the paragraph number sequence and co-occurrence range paragraph number range of the term citation sentence, and make a directional judgment based on whether the upper term in the term pair appears in the position before the lower term is cited, using the formula: ; Calculate the directional offset strength value, and sort the node directional connectivity relationship by offset strength to generate the term node directional connectivity; in, represents the directional shift strength value from term i to j, and represents the citation frequency of term i to j, Indicates the paragraph spacing between the definition segments of terms i to j. and are the sentence positions of terms i and j in the k-th citation respectively, and n is the number of co-occurring citations. Parameter meaning and formula calculation derivation process: Parameter and represent the citation frequency values of the two terms in the term pair that appear in the textbook respectively. This value is obtained by counting the number of times each term is explicitly cited in the teaching content paragraphs in the full text of the textbook. The statistical standard is that the term appears as an independent professional term in the text and is counted based on the paragraph number. The citation frequency of term T3 is 9 times, and the citation frequency of term T4 is 7 times, which are respectively from the 9 valid term marker recognition results and 7 valid term marker recognition results in paragraphs 3 to 7 of chapter 2. Parameter represents the inter-paragraph distance between the paragraph numbers where the definition segments of terms i and j are located. The difference is calculated based on the paragraph numbers of the term definitions. If term T3 is defined in paragraph 5 of the textbook and term T4 is defined in paragraph 8, then the inter-paragraph distance ; Parameter and represent the sentence number positions of terms i and j respectively when they co-occur for the k-th time. This position number is marked in the order of sentences within the paragraph and is automatically numbered by the text processing program and recorded uniformly with the sentence number. In the 3 co-occurring citations of terms T3 and T4, their sentence position combinations are (4, 5), (6, 6), and (7, 8) respectively. By calculating the differences of these three groups of positions respectively to get 1, 0, and 1, and then summing and dividing by the number of co-occurrences n = 3 to obtain the average sentence position difference. Substitute the actual values for calculation as follows: ; ; ; ; The sentence position differences for the three co-occurrences: |4 - 5| = 1, |6 - 6| = 0, |7 - 8| = 1, the sum is 2, and the total number of times is 3, so the second term is ; Substitute into the formula for calculation: ; The results show that the terms T3 and T4 show a relatively stable upward and downward extended connection relationship in the textbook content, and their directional shift strength value is below the preset directional threshold of 3.5. The directional shift strength value is an indicator used to quantify the position shift of the superordinate and subordinate terms in the term pair in the textbook, reflecting the citation order, dependency and importance. In textbook analysis, it can help understand the logical connection and structure between terms, indicating that the two terms can be judged as a forward connected path structure, which can be included in the graph structure node group as a valid directed edge when constructing the semantic connection expansion graph in the subsequent construction.
[0024] S213: calling the term pair structure in the term node directional connectivity, assembling the paths of the continuous connection relationship into a directed path graph structure, removing the path paragraph numbers whose path jumps exceed the chapter span, and establishing a knowledge point semantic connection expansion graph; Call the term pair data in the term node directional connectivity value, screen the term pairs with offset strength values less than the directional connectivity benchmark value of 3.5, assemble adjacent term nodes according to the path connection rules, take the node that appears earlier in the term pair as the path starting point, and connect the subsequent term pairs that meet the directional continuity in sequence to form a term directed path chain. If the path span spans more than 6 paragraph numbers or the path is interrupted by more than 2 term pairs, it is determined to be an unstable path and removed from the construction scope. The remaining term paths are constructed as node groups in the graph structure, and edge sets are constructed for each path structure according to the order of term appearance to form a directed path graph structure map. The path map information is integrated, including node groups, edge sets, directional mapping dictionaries and chapter annotation mapping tables, to establish a knowledge point semantic connection expansion map.
[0025] See also Figure 4 , the specific steps for obtaining the path node semantic matching resource list are: S311: Expand the node term group in the graph based on the semantic connection of knowledge points, obtain text resources from educational resource websites, and extract paragraph keywords, topic titles and body sentence blocks in the text to generate a set of key elements of text resources; By specifically collecting data sets in various fields of educational resource websites, which are subdivided into categories such as primary and secondary schools, universities, and adult education. During the collection process, according to the directory classification standard of educational resources, node term group keywords are matched one by one, and the corresponding paragraph keyword sets, topic title sets, and body text block sets are extracted respectively. Among them, the paragraph keyword set uses word segmentation technology to obtain the high-frequency terms in each paragraph that are highly relevant to the node term group. The topic title set extracts the chapter titles of each document, and the body text block set extracts the continuous text blocks of each paragraph. For example, if the node term group is {physical mechanics, electromagnetics, thermodynamics}, then "Summary of High School Physics Knowledge Points" is selected as the text on the educational resource website. After word segmentation, paragraph keywords such as {Newton's laws, Coulomb's law, heat energy conduction} are extracted, topic titles such as {Fundamentals of Mechanics, Principles of Electromagnetics} are extracted, and the body text blocks are the content paragraphs under each title. At the same time, the frequencies of occurrence and the positions of the paragraphs where the extracted keywords appear are recorded respectively. The recording example is shown in Table 1, see Table 1; Table 1: Record Table of Keyword Extraction and Frequencies Keywords Frequency of occurrence Paragraph position Newton's law 5 Section 2, Chapter 1 Coulomb's law 3 Section 3, Chapter 2 Thermal energy conduction 2 Section 1, Chapter 3 As shown above, keywords and their frequencies are extracted from educational resources, and their corresponding positions are recorded. After completion, a key element set of text resources is generated; S312: Call the key element set of text resources. For the number of graph term nodes covered by the paragraph keywords, extract the keyword-related node values, and mark the distribution positions of the starting nodes and extended nodes of the graph paths corresponding to the keywords. Use the formula: ; Calculate the keyword node matching degree, filter the keywords that cover the starting nodes and extend to the nodes, and obtain the keyword node coverage distribution index set; Among them, AK represents the keyword node matching degree, represents the number of starting nodes of the path covered by the th keyword, represents the number of extended nodes covered by the th keyword, N represents the total number of paragraph keywords, and AN represents the total number of node terms in the extended graph; Use the keyword node distribution statistical method to mark item T and item P respectively, and record their respective coverage quantities. Set the example where the keyword "Newton's laws" corresponds to the graph node terms {Fundamentals of Mechanics, Laws of Motion}. Item T covers Fundamentals of Mechanics, and item P covers Laws of Motion. After statistics, for the keywords, use the comprehensive calculation method of keyword node coverage degree, By setting the total number of paragraph keywords to 10, and setting that a certain keyword "Coulomb's law" covers 1 starting node of the path and 2 extended nodes, then 、 ,then its sub-item value is ; After summing up the item values of the keywords and dividing by the number of keywords and the total number of node terms. For example, if the total number of node terms is 50, an approximate value of the overall AK can be obtained. According to the calculation, the AK value obtained comprehensively is 0.048. Then, taking 0.03 as the node coverage benchmark value for judgment, if it is greater than the benchmark value, the keyword is determined to be effective; otherwise, it is excluded. By using this method, the keyword content with the node coverage meeting the requirements is screened to generate a keyword node coverage distribution index set. The formula derivation process is as follows: Given 、 , Then: ; ; The value of a single keyword item is ; Set the total number of keywords to 10, the sum of keyword items to 24.14, and the total number of node terms to 50. Then: ; Comparing with the benchmark value of 0.03, 0.04828 > 0.03. Therefore, it is determined to be a keyword with effective coverage, indicating that the keyword node coverage ability meets the standard.
[0026] S313: Based on the keyword node coverage distribution index set, screen the keyword content whose node coverage index reaches the set node coverage benchmark value, and combine the corresponding topic titles and text chunks to generate a path node semantic matching resource list; Select the set of keywords whose node coverage index is greater than the node coverage benchmark value, and combine the corresponding topic titles and text chunks one by one. During the combination process, use the topic title as the index item to connect the content of the text chunks to form a complete semantic block. At the same time, check whether the combined content contains the terms of the starting point and extended nodes in the corresponding path node. By comparing the term occurrence frequency and distribution range, if the coverage rate reaches the set coverage benchmark value (such as 80%), then mark the content of this semantic block as an effective resource block. For example, for the topic title "Fundamentals of Mechanics", the text chunk covers the terms {Newton's laws, laws of motion, law of inertia}, and the total number of terms is 4, with an actual coverage of 3, so the coverage rate is 75%, which is less than 80% and needs to be excluded. However, if another topic title "Principles of Electromagnetism" covers the terms {Coulomb's law, electric field, electric potential, magnetic field theory}, the total number of terms is 4, and the actual coverage is 4, with a coverage rate of 100%, then it is marked as an effective resource block. Integrate the effective resource blocks to generate a path node semantic matching resource list.
[0027] Please refer to Figure 5 , and the specific steps for obtaining the knowledge point path mispositioning set are as follows: S411: Based on the semantic similarity between the node information in the resource list and the path nodes, extract the operation steps in the student's answering process, number the node order, call the path node order corresponding to the question in the standard textbook for difference matching with the numbered order, obtain the numbered information corresponding to node missing, order misplacement, and substitution operations, and generate an abnormal node number set; Extract path nodes from the student's answering record. This path consists of various problem-solving operations sequentially selected by the student during the answering process. Each node corresponds to a step action. In actual operation, the operation type, input object, and knowledge point association label clicked for each answer can be extracted from the problem-solving log file, and then the nodes in the standard problem-solving path can be extracted by calling the standard path node information stored in the resource library. The standard path is the ideal problem-solving order defined according to the textbook and question type settings. Calculate the semantic similarity between the student's nodes and the standard nodes in a node-by-node manner. Represent the semantic content of each operation using word vectors, and obtain its semantic distance value through the cosine similarity or Euclidean distance method. Set the distance between the student operation node "extract known quantities" and the standard operation "identify stem quantities" to 0.28. If the similarity threshold is set to 0.3 based on this value, it is considered that the operation matches successfully; otherwise, it is marked as unmatched. In this way, establish node matching pairs one by one, number the matching nodes in the order they appear in the student path to generate a student path order number vector, such as the node sequence [3, 5, 7]; at the same time, generate a corresponding standard number vector [2, 4, 6] for the standard path nodes according to their positions in the question type standard path. After numbering, compare the two numbered sequences one by one. If a node appears in the student path but not in the standard path, it is marked as a substitution operation node; if there is a situation of skipping a standard path node, it is marked as a missing node; if the numbered order is reversed compared to the standard order, for example, the 5th node appears before the 3rd node, it is recorded as an order misplacement node. Suppose in the solution of a question type "solve the area of plane geometry", the student skipped the "construct auxiliary line" step and directly entered "substitute into the formula for calculation". At this time, the standard number of the construct auxiliary line operation is 4, and there is no numbered 4 operation in the student path, so it is a missing node. Finally, collect and organize the abnormal node numbers of the above three types to generate an abnormal node number set.
[0028] S412: Use the abnormal node number set, the total number of path structure nodes in the standard textbook, and the node order relationship, combine multiple node order comparison items and structure matching items in the abnormal node numbers, and use the formula: ; Calculate the difference degree of the path sequence structure, compare it with the path structure deviation judgment threshold, obtain the abnormal node sequence that meets the deviation threshold condition, and get the path abnormal structure positioning number set; Among them, BL represents the difference degree of the path sequence structure, d represents the number of abnormal node numbers, represents the sequence number of the abnormal node r in the student's answer path, represents the sequence number of the abnormal node r in the standard path, represents the semantic offset distance between the operation corresponding to the abnormal node r and the standard structure, represents the number value of the abnormal node r, represents the formula number value corresponding to the abnormal node r in the standard path, represents the maximum formula level number of the abnormal node r in the question type; Based on the actual numbers of each node in the abnormal node set, extract the corresponding sequence numbers recorded in the student path, and retrieve the sequence numbers of the same nodes in the standard path. Set the 3rd node in the student path as the 2nd node in the standard path. At the same time, collect the distance offset values of the corresponding operations at the semantic level of the two. The offset value is obtained by back-calculating the cosine distance between the word embedding vectors. Set the offset distance between the operations of node 3 and node 2 to 1.2. In addition, for the formula operation part involved in the question, record the formula numbers called in the student path and the formula numbers that should be called by the corresponding nodes in the standard path respectively. The two numbers can be found from the knowledge point formula library. Set the numbers to 12 and 10. At the same time, extract the maximum formula level number in the path of this question type and set it to 5. And so on, execute this process for the abnormal nodes and summarize the data into a table, see the table below; for the above data, by taking the absolute value of the difference between the sequence numbers of each node in the student path and the standard path, obtain the degree of sequence deviation. Then, square and sum the semantic offset value and the formula number offset term of each node and take the square root to comprehensively measure the differences in structure, semantics, and formula call in three aspects. Then, accumulate the sequence difference and the square root difference to obtain the total deviation value of each node. After summing the deviation values of the nodes and taking the average, the path sequence structure difference degree BL value is calculated. The larger the BL, the more obvious the structural difference. The calculation process based on the example data in Table 2 is as follows: Table 2: Path Difference Parameter Table Abnormal node number Student path number Standard path number Semantic offset distance Student formula number Standard formula number Maximum level number 1 3 2 1.2 12 10 5 2 5 4 0.8 15 13 5 3 8 7 1.5 17 14 5 Lists the structure matching parameters of three typical abnormal nodes, including various indicators such as sequence offset, semantic difference, formula mismatch, etc., and calculates the total path structure deviation value of 2.231 based on the data. This value is used to judge whether the current path meets the structural rationality condition and provides a numerical basis for the extraction of the path abnormal structure positioning number set; Substitute the data in the table into the calculation process: For the first node: ; ; The value of this item is: 1 + 1.232 = 2.232; The second node: ; ; The value of this item is: 1 + 0.877 = 1.877 The third node: ; ; The value of this item is: 1 + 1.581 = 2.581; Substitute into the formula for calculation: ; This result indicates that the comprehensive difference degree of the current path structure is 2.231. It can be compared with the set path structure difference threshold to determine whether there is a structural deviation in this path. If the set threshold is 2.0, then there is a structural anomaly in the current path, and it is necessary to further mark and locate its number set to obtain the path anomaly structure location number set.
[0029] S413: Call the path anomaly structure location number set. According to the knowledge point node number table and the knowledge point path classification table in the textbook, match the node numbers with the classification path numbers, extract the corresponding knowledge point labels, and generate the knowledge point path misunderstanding location set; It is necessary to find the corresponding number information from the knowledge point node number table in the textbook. This number table has established a clear mapping relationship between each step of each question type and the corresponding knowledge points. The path node with the number set to 7 is mapped to the knowledge point "Determination of the direction of spatial vectors". Then, according to the number matching result, extract the classification path number where this number is located from the knowledge point path classification table. For example, the path number is "A3", and this number represents a specific problem-solving strategy path, setting the three-dimensional configuration determination path in solid geometry questions; it is necessary to call multiple location number items in sequence according to the number order, map them to specific knowledge point labels in turn, construct an abnormal knowledge point sequence, and further combine the existing knowledge point advanced order information in the path classification table to judge whether the location numbers are continuous and whether they all belong to the same path structure branch. If they are not continuous or cross path numbers, it indicates that there are behaviors such as confusion of different problem-solving paths and jumping cognition in the student's answer. If the number sequence contains "A2" and "C4", then it spans the solid geometry and sequence function paths; a knowledge point focus information set for path deviation can be formed. When implementing, it can be further screened in combination with the occurrence frequency of each node, the question type matching situation, and the module to which the corresponding knowledge point belongs. If the occurrence frequency of a certain node is as high as 85% and is concentrated on the "Function Monotonicity" path, it indicates that this knowledge point is an error-prone point in the answering process, and generate the knowledge point path misunderstanding location set.
[0030] Please refer to Figure 6 , and the specific steps for obtaining the recommended sequence of textbook similar resources are as follows: S511: Invoke the collection of knowledge point path error location, perform character matching between the knowledge point keywords marked by the path nodes and the chapter titles of the textbook content, extract the textbook chapter numbers and title information where the path nodes are located, and obtain the corresponding list of path node chapters; After disassembling the knowledge point keywords marked by the path nodes into independent keyword contents one by one, perform character-level equal-length position matching operations with the chapter titles of the textbook content. During the matching process, for each keyword, calculate the ratio between the number of matching characters and the total number of characters in the chapter title according to its position and length in the chapter title. When this ratio exceeds the preset chapter association determination threshold, extract the chapter number and title content of this chapter as the corresponding chapter information of the path node. Suppose when processing the keyword "Composition of Forces" of the path node numbered N1, the matching chapter title is "Composition of Forces", the total number of characters in the title is 6, and the number of completely matching characters is 4, then the corresponding matching ratio is 4 / 6, approximately 0.667. If the chapter association determination threshold is set to 0.6, this matching result is valid, and the chapter number 3 and the title "Composition of Forces" need to be recorded in the corresponding list of path nodes. In actual processing, perform such operations on the nodes in the path node set in sequence to obtain the chapter numbers and title data of all matching chapters, form a one-to-one corresponding set of path node chapter information, refer to the path node chapter matching information table, and obtain the corresponding list of path node chapters.
[0031] S512: Based on the corresponding list of path node chapters, locate all the explanatory paragraphs in the chapters corresponding to the path nodes, extract the explanatory paragraphs containing the path node keywords, call all the sentence contents in the explanatory paragraphs, calculate the total number of words in the sentences and the length of the sentences matching the path nodes, and obtain the path node explanatory density distribution; Check the textbook text content item by item, extract the explanatory paragraph text in each corresponding chapter, parse the content sentence by sentence in each explanatory paragraph, and retrieve whether there is a word expression consistent with the path node keywords. For the explanatory paragraphs with matching keywords, call all the sentences in this paragraph for sentence-level character statistics, respectively obtain the total number of words in the explanatory paragraph sentences and the number of words in the sentence where the path node keywords are located, and calculate the sentence density accordingly. The sentence density is defined as the percentage of the number of words in the sentence related to the path node keywords in the total number of words in all the sentences in this paragraph. Suppose in the chapter "Composition of Forces", an explanatory paragraph contains 5 sentences, with a total of 280 words, and the total number of words in the 2 sentences directly related to the path node "Composition of Forces" is 120, then the path node explanatory density value of this paragraph is , after processing the explanatory paragraphs under each chapter corresponding to each path node, record their explanatory density values respectively and establish a distribution vector set, and further archive and sort the explanatory density values of the path nodes in the chapters they correspond to to form a set array , where each Indicates the explanation density of a chapter paragraph, which is arranged by indexing with the chapter paragraph number in actual processing. Set the density of the paragraph with chapter number 3 to 42.86%, the paragraph with number 7 to 60.00%, and the paragraph with number 12 to 38.75%, then form a distribution density vector , where each density value is generated by merging after sentence-by-sentence calculation, obtaining the explanation density distribution of path nodes.
[0032] S513: According to the explanation density distribution of path nodes, screen out the chapter paragraphs whose density values exceed the set explanation density threshold, extract the example content of path node keywords, count the total frequency of examples in the chapter paragraphs, and generate a recommended sequence of textbook similar resources; The set of density values formed according to the explanation density distribution of path nodes , successively compare the density values of each explanation paragraph with the preset explanation density threshold for judgment. If the density value exceeds the set explanation density threshold, the chapter paragraph will be included in the subsequent example content extraction range. When the set explanation density threshold is 0.40, the chapter paragraphs numbered 3 and 7 will be retained, while the paragraph numbered 12 is excluded because the density is 0.3875 and does not meet the threshold condition. For the retained chapter paragraphs, extract the sentences containing path node keywords sentence by sentence, identify the expression content with example nature, and screen out and summarize the sentences containing typical identifier words such as "set", "as shown in", "provided with", "if known" in the judgment sentence structure. The example content extracted from each chapter paragraph is separately marked and its occurrence frequency is counted, and the frequency value represents the example usage intensity corresponding to the path node in this paragraph of the explanation. Set that in chapter number 3, there are 3 example sentences related to the path node keyword with the "set" structure, then its frequency value is 3, and in number 7, there are 5 example sentences of the same type, and the frequency is 5, obtaining a frequency array , in order to uniformly compare the example usage in each chapter, it is necessary to normalize the frequency array, using the normalization formula: ; where is the maximum frequency, is the minimum frequency; here the maximum frequency is 5 and the minimum frequency is 3, then chapter 3 is normalized to , chapter 7 is , obtaining a normalized example frequency array , using this as the sorting basis, arranging the corresponding chapter order from high to low, that is, giving priority to recommending chapter number 7, followed by 3, and chapter number 12 is not recommended because it does not meet the threshold condition, generating a recommended sequence of textbook similar resources.
[0033] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the relevant art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for expanding similar resources in a teaching material library based on web resource scraping, characterized in that It includes the following steps: S1: Obtain the textbook chapter text, paragraph structure table, and glossary. Compare the positions of chapter titles and term definition paragraphs, extract the concept frequencies and context positions, and construct the concept attribution hierarchical path based on the concept co-occurrence frequency and paragraph order to generate a set of knowledge point hierarchical path structures; S2: According to the arrangement order of path nodes in the set of knowledge point hierarchical path structures, collect the co-occurrence sentence blocks and reference sentences between terms in the textbook, and call the upper and lower reference frequencies, sentence block distance positions, and co-occurrence ranges of each pair of nodes to generate a knowledge point semantic connection expansion map; S3: Based on the node term groups in the knowledge point semantic connection expansion map, obtain the text resources of educational resource websites, and judge whether the number of map term nodes covered by the keywords and the term distribution positions cover the path starting point and extended nodes to generate a list of path node semantic matching resources; S4: According to the list of path node semantic matching resources, call the step information of students' answering questions and the formula call order, compare with the problem-solving path structure corresponding to the questions in the standard textbook, and judge whether there are missing nodes, misplaced orders, or incorrect replacement operations in the path to generate a set of knowledge point path error positionings; 2. The method for expanding similar resources in a teaching material library based on web resource scraping according to claim 1, wherein The set of knowledge point hierarchical path structures includes a concept attribution path table, a knowledge point theme index library, and a term hierarchical identification group. The knowledge point semantic connection expansion map includes a term reference network, a node coverage range map, and a semantic extension direction table. The list of path node semantic matching resources includes term matching resource entries, a path coverage comparison table, and a node distribution position set. The set of knowledge point path error positionings includes a problem-solving deviation node list, a path offset number group, and an abnormal knowledge point label set.
3. The method for expanding similar resources in a teaching material library based on web resource scraping according to claim 1, wherein The specific steps for obtaining the set of knowledge point hierarchical path structures are as follows: S111: Obtain the textbook chapter text, paragraph structure table, and glossary, extract the starting positions of paragraph titles, the paragraph numbers and sentence positions of terms, calculate the occurrence frequency and position distribution interval of terms in the chapter, and analyze the term density of the term intervals covered by the titles and the subordinate paragraphs to generate a chapter term density distribution; S112: According to the dense section numbers in the chapter term density distribution, call the defined positions, term co-occurrence pairs, and the number of sentences in the definition paragraphs in the glossary, and judge the sequential positions of co-occurring term pairs in the same section to generate a term order path combination table; S113: Call the term group path structures in the term order path combination table, compare with the chapter index, paragraph numbers, and term reference quantities in the paragraph structure table, screen the cross-paragraph paths of consecutive term nodes in the path, and record the chapter positions to which they belong to generate a set of knowledge point hierarchical path structures.
4. The method for expanding similar resources in a teaching material library based on web resource scraping according to claim 3, wherein The specific steps for obtaining the knowledge point semantic connection expansion map are as follows: S211: Use the arrangement order of multiple path nodes in the set of knowledge point hierarchical path structures to correspondingly find the common sentence block areas, term definition paragraph numbers, and term reference sentence positions of term pairs in the textbook, and call the definition paragraph numbers to compare the front and back definition orders of terms to generate a term co-occurrence sentence block difference value; S212: Based on the difference value of the term co-occurrence sentence block, call the definition position and citation frequency of the upper and lower term nodes in the term pair in the textbook, obtain the paragraph number sequence and co-occurrence range paragraph number range of the term citation sentence, make a directional judgment based on whether the upper term in the term pair appears in the position before the lower term is cited, calculate the directional offset strength value, and sort the node directional connectivity relationship according to the offset strength to generate the term node directional connectivity; S213: calling the term pair structure in the term node directional connectivity, assembling the paths of continuous connection relationships into a directed path graph structure, removing the path paragraph numbers whose path jumps exceed the chapter span, and establishing a knowledge point semantic connection expansion graph.
5. The method for expanding similar resources in a teaching material library based on web resource scraping according to claim 4, wherein The steps for obtaining the path node semantic matching resource list are specifically as follows: S311: Based on the node term group in the knowledge point semantic connection expansion graph, obtain text resources from educational resource websites, extract paragraph keywords, subject titles and text sentence blocks in the text, and generate a set of key elements of text resources; S312: calling the text resource key element set, extracting the keyword-related node values for the number of graph term nodes covered by the paragraph keywords, marking the distribution positions of the graph path starting nodes and extended nodes corresponding to the keywords, calculating the keyword node matching degree, screening the keywords that cover the starting nodes and extend to the nodes, and obtaining the keyword node coverage distribution indicator set; S313: Based on the keyword node coverage distribution index set, filter the keyword content whose node coverage index reaches the set node coverage benchmark value, and combine the corresponding subject title and text sentence blocks to generate a path node semantic matching resource list.
6. The method for expanding similar resources in a teaching material library based on web resource scraping according to claim 5, wherein, The steps for obtaining the knowledge point path error location set are specifically as follows: S411: Based on the semantic similarity between the node information in the path node semantic matching resource list, extract the operation steps in the student's answering process, and number the node sequence, call the number sequence and the path node sequence corresponding to the questions in the standard textbook for difference matching, obtain the number information corresponding to the node missing, sequence misplacement and replacement operation, and generate an abnormal node number set; S412: using the abnormal node number set and the total number of path structure nodes and node sequence relationship in the standard textbook, combining multiple items in the abnormal node number with node sequence comparison items and structure matching items, calculating the path sequence structure difference, comparing it with the path structure deviation judgment threshold, obtaining the abnormal node sequence that meets the deviation threshold condition, and obtaining the path abnormal structure positioning number set; S413: calling the path anomaly structure positioning number set, matching the node number with the classification path number according to the knowledge point node number table and the knowledge point path classification table in the textbook, extracting the corresponding knowledge point label, and generating a knowledge point path error positioning set.
7. The method for expanding similar resources in a teaching material library based on web resource scraping according to claim 1, wherein The method further comprises step S5: S5: calling the knowledge point path misunderstanding location set, matching the textbook chapters and explanation paragraphs, extracting the core explanation sentences in the chapter where the path node is located, checking the explanation length and example occurrence frequency of the matching sentences, and generating a recommendation sequence of similar resources of the textbook; The recommended sequence of similar resources for the teaching materials includes a matching teaching material fragment index, a list of resource priority identifiers, and an example coverage density annotation set.
8. The method for expanding similar resources in a teaching material library based on web resource scraping according to claim 7, wherein The specific steps for obtaining the recommended sequence of similar resources for the teaching materials are as follows: S511: Call the set of knowledge point path error positionings, perform character matching between the knowledge point keywords marked by the path nodes and the chapter titles of the teaching material content, extract the serial numbers and title information of the teaching material chapters where the path nodes are located, and obtain the corresponding list of path node chapters; S512: Based on the corresponding list of path node chapters, locate all the explanatory paragraphs in the chapters corresponding to the path nodes, extract the explanatory paragraphs containing the path node keywords, call all the sentence contents in the explanatory paragraphs, calculate the total number of words in the sentences and the length of the sentences matching the path nodes, and obtain the path node explanatory density distribution; S513: According to the path node explanatory density distribution, screen the chapter paragraphs with density values exceeding the set explanatory density threshold, extract the example contents of the path node keywords, count the total frequency of examples in the chapter paragraphs, and generate the recommended sequence of similar resources for the teaching materials.