Intelligent fusion and quantitative evaluation method and system of multi-source teaching resources based on large language model

CN122818210APending Publication Date: 2026-09-25NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610751878.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

现有的搜索引擎和检索系统主要基于关键词匹配,难以理解查询的语义含义,无法实现精准的跨资源知识点定位即便部分现有技术引入了向量化语义匹配方法,但由于缺乏上游阶段结构化知识点的精确锚定,匹配过程中难以有效利用知识点的逻辑层级和层次位置信息来辅助判定语义相关性,容易产生语义漂移和误匹配

Benefits of technology

[0023]一种计算机可读存储介质,所述计算机可读存储介质上存储有计算机程序,所述计算机程序被处理器执行时实现所述的基于大语言模型的多源教学资源智能融合与定量评价方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122818210A_ABST
    Figure CN122818210A_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-source teaching resource intelligent fusion and quantitative evaluation method and system based on large language model belongs to the application of artificial intelligence in educational technology.The method includes: benchmark knowledge point structured combing stage, adopts large language model to carry out knowledge point extraction to anchor resource and generates structured tuple;Cross resource precision matching stage, realize semantic matching by vector embedding and cosine similarity calculation;Multi-dimensional data extraction stage, extract core information point, logical level, effective details, evidence quantity, text word number and redundancy proportion and other indexes;Based on the richness quantization stage of multiple, calculate five indexes such as core information coverage multiple, logical level integrity multiple, detail supplement density multiple, evidence quantity multiple and redundancy correction coefficient, and obtain comprehensive multiple by weighted synthesis.The application constructs a kind of scientific objective quantitative evaluation system, and supports cross-language resource alignment and knowledge graph dynamic updating.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of artificial intelligence and educational technology, specifically involving a method and system for intelligent fusion and quantitative evaluation of multi-source teaching resources based on a large language model. Background Technology

[0002] With the deepening of the informatization and digital transformation of higher education, the global education field is undergoing profound changes. Digital teaching resources are experiencing explosive growth, and teaching resources for various courses are becoming increasingly abundant and diverse in form. A typical university course is usually equipped with multiple forms of teaching resources, including: carefully prepared presentation slides (PowerPoint / Keynote, etc.), a complete textbook (including print and electronic versions), cutting-edge academic papers and research reports, abundant supplementary reading materials, laboratory manuals and practical guides, exercise sets and test question banks, etc.

[0003] The widespread availability of Learning Management Systems (LMS) and Open Educational Resources (OER) platforms has made these materials more accessible than ever before, creating a vast knowledge base that spans multiple formats, languages, and teaching methods.

[0004] The internationalization of higher education has added an important dimension to this field. Universities are increasingly serving a diverse student body representing different linguistic and cultural backgrounds. To accommodate the needs of international learners, educational institutions typically maintain course materials in multiple languages ​​and create multilingual resource collections. This multilingual dimension is particularly prominent in courses designed for international students, where course materials may be available simultaneously in the host language (such as Chinese) and English or other widely used languages.

[0005] In recent years, the rapid development of artificial intelligence technology, especially large language models, has brought new opportunities to the field of educational technology. Large language models, represented by the GPT series and Claude series, have demonstrated powerful semantic understanding, text generation, and multilingual processing capabilities, providing a new technological path to solve the long-standing problem of resource integration in the education field.

[0006] Despite the increasing abundance of teaching resources, this apparent abundance actually presents new challenges for learners and educators. In-depth analysis reveals the following main shortcomings in existing technologies: The problem of knowledge fragmentation is severe: the same knowledge point is often scattered across different media such as presentation slides (PPT), textbooks, and academic papers, lacking an effective connection mechanism. This fragmentation forces students to switch back and forth between multiple resources, making it difficult to build a systematic knowledge framework. Although existing technologies attempt to use large language models to process textbooks into knowledge graphs, they generally use simple triples or flat keyword tagging systems to organize knowledge points, failing to establish a multi-dimensional structured representation that includes core concepts, related elements, logical hierarchy, and hierarchical position. As a result, the extracted knowledge points lack sufficient semantic and hierarchical information to support accurate cross-resource alignment and in-depth quantitative evaluation.

[0007] The lack of resource retrieval and comparison mechanisms means that learners need to manually search repeatedly across multiple resources when looking for and comparing specific knowledge points. Existing search engines and retrieval systems are mainly based on keyword matching, which makes it difficult to understand the semantic meaning of queries and achieve accurate cross-resource knowledge point positioning. Even though some existing technologies have introduced vectorized semantic matching methods, the lack of precise anchoring of structured knowledge points in the upstream stage makes it difficult to effectively utilize the logical hierarchy and positional information of knowledge points to assist in determining semantic relevance during the matching process, which easily leads to semantic drift and mismatches.

[0008] The lack of a scientific quantitative evaluation system is a significant issue. Existing technologies struggle to objectively quantify the differences in the richness of presentation of the same knowledge point across various teaching resources, failing to provide learners and educators with a scientific basis for resource evaluation. Traditional resource evaluation relies primarily on subjective judgment or simple statistical indicators (such as word count and page count), which cannot reflect the quality and depth of resource content. While existing research has proposed multidimensional evaluation methods for teaching programs based on knowledge graphs, the design of these evaluation dimensions fails to simultaneously consider the five orthogonal dimensions of breadth of knowledge coverage, structural depth and completeness, detail density per unit length, sufficiency of argumentation, and suppression of information redundancy. Furthermore, the lack of quantitative means for normalized comparison in the form of "multiples" makes it impossible to provide intuitive, interpretable, and traceable quantitative evaluation results.

[0009] Insufficient multilingual integration capabilities: Although the same course may have teaching resources in multiple languages, there is a lack of intelligent language conversion and resource matching mechanisms. Existing cross-language knowledge graph alignment technologies are mainly geared towards encyclopedic knowledge bases or general texts, failing to adapt and optimize for the specific characteristics of teaching resources. In particular, they fail to deeply couple cross-language alignment with structured knowledge extraction and multi-dimensional quantitative evaluation, resulting in the integration of teaching resources in cross-language scenarios remaining at a superficial translation and comparison level.

[0010] Lack of dynamic updating capability. Teaching resources are constantly being updated, but existing technology lacks an incremental update mechanism, making it impossible to reflect resource changes in real time, which affects the timeliness and usability of the system.

[0011] The lack of tightly coupled closed-loop connections between various technical stages: Existing technologies generally lack tightly coupled closed-loop connections between various technical stages such as knowledge structuring, semantic matching, and quantitative evaluation. Each stage operates independently, and the output structure and semantic information of the previous stage cannot be effectively transmitted to the subsequent stages as constraints and anchoring bases. This results in a systemic defect in the overall technical solution: "each stage is optimized separately, but the overall effect is poor." Existing technologies have not yet developed a technical solution that designs "structured extraction—precise matching—multidimensional quantification—multiple-fold evaluation" as a tightly coupled four-stage closed-loop whole, nor do they lack an integrated approach to uniformly handle multilingual alignment and dynamic updates within this closed-loop framework. Summary of the Invention

[0012] Purpose of the invention: To address the shortcomings of the existing technologies, this invention provides a method and system for intelligent fusion and quantitative evaluation of multi-source teaching resources based on a large language model. It can establish a closed-loop processing architecture with tight coupling of four stages: "structured extraction - precise matching - multi-dimensional quantification - multiple evaluation". It realizes intelligent fusion of multi-source teaching resources through the deep semantic understanding capability of the large language model, and realizes scientific quantitative evaluation of the richness of teaching resources through a five-dimensional orthogonal multiple index system. It also integrates cross-language alignment and dynamic knowledge graph update capabilities within the closed-loop architecture.

[0013] Technical solution: A method for intelligent fusion of multi-source teaching resources based on a large language model, comprising the following steps: S1. Structured Analysis of Baseline Knowledge Points: This involves format parsing and text extraction of anchored teaching resources, and standardizing knowledge points. This includes designing prompt templates to guide the extraction of knowledge points from a large language model, generating structured tuples for each knowledge point. ,in Indicates the core concept, Indicates related elements, Represents logical hierarchy, Indicates hierarchical position; This step involves representing the generated quadruple structure as corresponding encoded information, providing multi-dimensional alignment anchors for cross-resource matching; S2. Cross-resource precise matching: Preprocessing and content segmentation of the target teaching resources are performed. A pre-trained text embedding model is used to transform the baseline knowledge points and target resource text fragments into high-dimensional vectors. Semantic matching between the two is achieved through cosine similarity calculation, and a similarity threshold is set. The process involves screening, then synonym expansion to enhance recall, and purification of candidate results to achieve precise alignment. S3. Multi-dimensional data extraction: Extracting the number of core information points from benchmark resources. Number of logical levels Effective number of details Number of supporting evidence Text word count and redundancy ratio As an evaluation benchmark, the number of core information coverage points extracted from the content matched with the target resources is used. Number of logical levels covered Effective number of details Number of supporting evidence Text word count and redundancy ratio ; S4. Richness Measurement Based on Individual Component Multiples: Calculate the five individual component multiples and obtain the overall richness multiple through weighted summation. ; It is a multi-dimensional indicator used to calculate the coverage multiple of core information. It is a multiple of logical hierarchy integrity. It is a detail supplement density multiple, This is evidence of multiples in quantity. It is a redundancy correction factor; S5. Cross-language resource alignment: A multilingual pre-trained model is used to map teaching resource texts in different languages ​​to a unified multilingual semantic space. Cross-language similarity calculation is used to achieve semantic alignment between the source language baseline knowledge points and the target language resource content. After translation verification to ensure the alignment quality, a unified representation structure for cross-language knowledge points is established. S6. Dynamic Update of Knowledge Graph: Establish a resource monitoring mechanism to detect changes in teaching resources, extract incremental knowledge from the changed content, update entities and relationships in the knowledge graph based on the extraction results, and complete the consistency maintenance of the knowledge graph to ensure the timeliness and accuracy of the knowledge system.

[0014] Furthermore, the anchored teaching resources mentioned in step S1 include one or more of the following formats: PPT / PPTX, PDF, DOC / DOCX, and HTML. The standardization process includes terminology normalization, redundancy elimination, relation annotation, and index construction, wherein index construction involves transforming knowledge points into a combination of "short keywords + complete phrases". The prompt template includes a task description, output format specifications, extraction dimension description, and example guidance elements. The output format is JSON. The prompt template is based on a knowledge point granularity control strategy. By setting extraction granularity parameters, the abstraction level of knowledge points is controlled to ensure the consistency of semantic granularity of the extracted knowledge points and avoid deviations in subsequent matching and evaluation due to uneven granularity. The dimensions for extracting knowledge points include core concepts, key elements, and logical frameworks.

[0015] Furthermore, the segmentation strategy for dividing the target resource content described in step S2 comprehensively considers natural paragraph boundaries, chapter structure, semantic integrity, and length limitations of vectorization processing; The pre-trained text embedding model includes one or more of the following: general text embedding model, multilingual embedding model, and domain-specific embedding model. The formula for calculating the cosine similarity is:

[0016] in, The embedding vector of the benchmark knowledge points. The embedding vector of the target resource text fragment. Describes the L2 norm of a vector; The synonym expansion includes term synonym expansion, abbreviation expansion, and related concept association based on knowledge graphs; The candidate result purification includes using a large language model to evaluate semantic relevance, eliminating mismatched content, integrating scattered related content, and generating target resource content that corresponds one-to-one with the benchmark knowledge points.

[0017] Furthermore, in step S4, the number of words in the text... and It is used as a normalization factor in the calculation of detail supplement density to eliminate the impact of differences in the length of different resources on the evaluation of detail density. It is not used as an independent evaluation dimension in the weighted comprehensive calculation.

[0018] Furthermore, the method includes a factor based on the overall richness multiple. Classification of resource richness levels: The target resources are significantly more abundant than the benchmark resources; 1.2 < The target resources are relatively abundant; 0.8 < The target resources are comparable in abundance to the benchmark resources; 0.8, the target resources are relatively simple.

[0019] Furthermore, the multilingual pre-trained model mentioned in step S5 includes one or more of mBERT, XLM-RoBERTa, and multilingual-e5; The translation verification involves translating the matched target language content into the source language, comparing the semantic similarity with the original source language's baseline knowledge points, and identifying and marking translation deviations or conceptual inconsistencies. The unified representation structure for cross-language knowledge points is a unique node in the knowledge graph corresponding to each knowledge point. Different language versions are stored as multilingual attributes of the node, while preserving the mapping relationship and alignment confidence between language versions. It also includes the construction of a cross-language mapping table between technical terms and the target language.

[0020] Furthermore, the resource monitoring mechanism described in step S6 achieves change detection by establishing version identifiers and content hash values ​​for access resources. The identified change types include adding content, modifying content, and deleting content. The incremental knowledge extraction process only processes the changed resource fragments, reusing the existing knowledge point structure and relationships; The knowledge graph consistency maintenance includes verifying structural integrity, handling knowledge point version conflicts, recalculating affected quantitative indicators and multiplier scores, and updating the search index.

[0021] On the other hand, the present invention also provides an intelligent fusion and quantitative evaluation system for multi-source teaching resources based on a large language model to implement the method, comprising: The resource access module is used to receive multi-source heterogeneous teaching resources, perform format parsing, text extraction and preprocessing on the resources, and output analyzable text data. The knowledge extraction engine is built on a large language model, including the GPT series and Claude series models. It accurately extracts knowledge points from preprocessed anchored resources through preset prompt templates, generates knowledge points into structured tuples, and completes the standardization of knowledge points. The semantic matching engine, built on the vector space model and deep semantic understanding, realizes the vectorized representation of benchmark knowledge points and target resource text fragments, cosine similarity calculation, threshold filtering, synonym expansion and content purification, and completes accurate cross-resource matching. The text embedding models include text-embedding-ada-002, the sentence-transformers series, paraphrase-multilingual-mpnet-base-v2, or domain-specific embedding models finely tuned for specific disciplines. The quantitative evaluation engine is used to automatically extract multi-dimensional quantitative indicators from the matching content of benchmark resources and target resources, calculate five individual multiple indicators and the comprehensive richness multiple, and output quantitative evaluation results and resource richness level. The multilingual alignment module is built using a multilingual pre-trained model to achieve cross-lingual semantic alignment, translation verification, and unified representation of cross-lingual knowledge points for teaching resources in different languages. The dynamic update module is used to establish a teaching resource monitoring mechanism to complete resource change detection, incremental knowledge extraction, knowledge graph entity and relationship update, and knowledge graph consistency maintenance. The closed-loop scheduling controller is used to coordinate the data flow and execution order between modules to form a four-stage serial processing pipeline. It also reactivates the processing flow of the corresponding stage when the dynamic update is triggered. This includes providing a standardized inter-module communication interface through which each module realizes the structured transmission of data, ensuring the data consistency and atomicity of the four-stage closed-loop processing. The various modules of this system work together to achieve intelligent integration and quantitative evaluation of multi-source teaching resources, output a unified knowledge graph and quantitative evaluation results, and adopt a modular architecture design. It supports plug-and-play use of multiple large language models, flexible access to multiple document formats including PPT / PPTX, PDF, DOC / DOCX, and HTML, and seamless expansion to multiple languages.

[0022] A computer device includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program to implement the intelligent fusion and quantitative evaluation method for multi-source teaching resources based on a large language model.

[0023] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the intelligent fusion and quantitative evaluation method for multi-source teaching resources based on a large language model.

[0024] Beneficial Effects: Existing technologies do not design the aforementioned technical aspects as a tightly coupled four-stage closed-loop whole. Although existing technologies utilize large language models to process teaching materials into knowledge graphs, their knowledge representation structures differ fundamentally from the original four-tuple structured approach of this invention, failing to provide sufficient structured anchoring information for subsequent multi-dimensional accurate matching and quantitative evaluation. Similarly, while there are existing multi-dimensional dynamic evaluation methods for teaching programs based on knowledge graphs, their evaluation dimensions often suffer from information redundancy and dimensional coupling, failing to achieve true orthogonal decomposition. Furthermore, existing cross-language knowledge graph alignment technologies are primarily geared towards general domains, failing to deeply integrate cross-language alignment with the structured extraction and quantitative evaluation specific to teaching resources.

[0025] This invention addresses the pain points of existing multi-source teaching resource integration, such as low efficiency, vague evaluation standards, poor cross-language adaptability, and lagging knowledge system updates. Leveraging the deep semantic understanding and vector representation capabilities of a large language model, it achieves intelligent integration and precise quantitative evaluation of multi-source teaching resources. Its beneficial effects are mainly reflected in the following aspects: (1) Improve the efficiency and accuracy of integrating multi-source teaching resources and break down resource barriers. Structured knowledge extraction is more efficient: By guiding the large language model with preset prompt templates, it automatically extracts knowledge points from anchored resources in multiple formats such as PPT, PDF, and DOCX, and generates standardized structured tuples, replacing the traditional manual extraction method. This significantly reduces labor costs, improves the efficiency and accuracy of knowledge point extraction, and avoids omissions and errors caused by manual operation.

[0026] More accurate cross-resource matching: A pre-trained text embedding model is used to transform knowledge points and target resource text fragments into high-dimensional vectors. Combined with cosine similarity calculation and synonym expansion strategy, a reasonable similarity threshold is set to achieve accurate matching. At the same time, content purification is used to remove mismatched content, which solves the problems of semantic disconnect and low matching accuracy in traditional matching methods, and achieves efficient alignment and integration of teaching resources from different sources and in different formats.

[0027] Seamless integration of multilingual resources: By leveraging multilingual pre-trained models, teaching resources in different languages ​​are mapped to a unified semantic space. Through cross-language similarity calculation and translation verification, a unified representation structure for cross-language knowledge points is established, breaking down language barriers and achieving barrier-free integration of multilingual teaching resources, adapting to international teaching and cross-language knowledge dissemination scenarios.

[0028] (2) To achieve quantitative evaluation of the richness of teaching resources and provide objective decision-making basis.

[0029] The evaluation index system is more comprehensive: quantitative indicators are extracted from five dimensions: core information coverage, logical hierarchy completeness, detail supplementation density, supporting evidence quantity, and redundancy control. This constructs a multi-dimensional evaluation system, which makes up for the shortcomings of traditional evaluation methods that rely solely on subjective judgment and have a single evaluation dimension, making the evaluation results more objective and comprehensive.

[0030] The evaluation process is more scientific and traceable: the comprehensive richness multiple is obtained through multiple calculation and weighted summation, and the resource richness level is clearly defined. The entire evaluation process is based on quantifiable indicators and clear calculation formulas. The steps are clear and traceable, avoiding the randomness of subjective evaluation, and providing a scientific quantitative basis for the selection, optimization and integration of teaching resources.

[0031] The evaluation system is more adaptable: it can flexibly adjust the weight of each indicator according to the needs of different disciplines and teaching scenarios, adapt to the evaluation needs of different types of teaching resources (such as theoretical, practical and case-based), and improve the universality and flexibility of the evaluation scheme.

[0032] (3) To achieve dynamic updates of the knowledge system and ensure the timeliness and accuracy of teaching resources.

[0033] Real-time resource change detection: By establishing a resource monitoring mechanism through version identifiers and content hash values, changes such as the addition, modification, and deletion of teaching resources can be quickly identified, avoiding knowledge lag caused by untimely resource updates.

[0034] Incremental updates improve efficiency: Incremental knowledge extraction is performed on changed content, reusing existing knowledge point structures and relationships, eliminating the need to reprocess all resources, significantly improving the efficiency of knowledge graph updates and reducing system computational costs.

[0035] Knowledge system consistency maintenance: During the update process, the consistency verification and version conflict handling of the knowledge graph are completed simultaneously to ensure the accuracy and completeness of entities and relationships in the knowledge graph, providing reliable knowledge support for teaching and learning and avoiding misleading information.

[0036] (4) Enhance the system's practicality and scalability to adapt to diverse teaching scenarios.

[0037] Multi-format resource compatibility: Supports access and parsing of various mainstream document formats such as PPT / PPTX, PDF, DOC / DOCX, and HTML, adapting to the diverse storage formats of existing teaching resources without the need for resource format conversion, thus lowering the barrier to entry for users.

[0038] Modular architecture for easy expansion: The system adopts a modular design, with each functional module working together. It supports plug-and-play functionality for various large language models and text embedding models. Core modules can be flexibly upgraded according to technological developments and user needs, improving the system's maintainability and scalability.

[0039] Seamlessly adaptable to multiple scenarios: It can be integrated with existing learning management systems and open educational resource platforms, and is suitable for various teaching scenarios such as college teaching, online education, and vocational training. It can not only meet the resource integration and evaluation needs of teachers, but also provide learners with more accurate and richer learning resources, thereby improving teaching and learning effectiveness.

[0040] (5) Lower the barriers to technology use and promote the standardization and normalization of teaching resources.

[0041] High degree of automation: From resource access, knowledge point extraction, matching and integration, to quantitative evaluation and knowledge updating, the entire process is automated, requiring no professional technical skills from users, thus lowering the technical threshold for teaching resource management and evaluation, and facilitating widespread application.

[0042] Promote resource standardization: By standardizing knowledge points (terminology standardization, index construction, etc.), establish a unified knowledge point representation structure, promote the standardized management of multi-source teaching resources, avoid the problems of disorganized and unusable resources, and improve the reuse rate and value of teaching resources.

[0043] In summary, this invention, through its intelligent, quantitative, and dynamic design, effectively addresses the core pain points in the integration and evaluation of multi-source teaching resources, improves the efficiency and quality of teaching resource management, provides reliable technical support for teaching activities, and has significant practicality, innovation, and promotional value. Attached Figure Description

[0044] Figure 1 This is a flowchart illustrating the overall process of the method of the present invention.

[0045] Figure 2 This is a schematic diagram of the four-stage closed-loop processing architecture and data flow relationship.

[0046] Figure 3 This is a schematic diagram of the structured tuple quadruple representation and its transmission relationship at each stage.

[0047] Figure 4 This is a schematic diagram illustrating the relationship between multi-dimensional indicator extraction and multiple calculation.

[0048] Figure 5 A schematic diagram illustrating the process of aligning and unifying representations across language resources.

[0049] Figure 6 This is a schematic diagram of the dynamic updating and cascading re-evaluation process of a knowledge graph.

[0050] Figure 7 This is a diagram of the system's functional module architecture.

[0051] Figure 8 This is a schematic diagram illustrating the classification of resource richness levels and multidimensional diagnostic analysis. Detailed Implementation

[0052] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0053] Combination Figure 1 and Figure 2 As shown, the present invention provides an intelligent fusion and quantitative evaluation method for multi-source teaching resources based on a large language model. The method adopts a closed-loop processing architecture with four tightly coupled stages: "structured extraction - precise matching - multidimensional quantification - multiple evaluation". There are strict data flow dependencies between each stage, including the following steps.

[0054] S1. Structured organization of key knowledge points. For example... Figure 3As shown, the anchored teaching resources undergo format parsing and text extraction. After text cleaning and noise filtering, the knowledge points are standardized, including designing prompt templates to guide the large language model in extracting knowledge points and generating structured tuples for each knowledge point. This four-tuple structured representation uniformly encodes the semantic, structural, and hierarchical information of the knowledge points, providing multi-dimensional alignment anchors for subsequent cross-resource matching and forming the basic data structure of the entire closed-loop processing architecture.

[0055] S2. Cross-resource precise matching. Based on the structured knowledge point tuples output in step S1, the target teaching resources are preprocessed and segmented. A pre-trained text embedding model is used to transform the baseline knowledge points and target resource text fragments into high-dimensional vector representations. Semantic matching between the two is achieved through cosine similarity calculation. A similarity threshold τ is set for filtering, and synonym expansion is used to enhance recall and candidate results are purified to achieve precise alignment. This matching process fully utilizes the logical hierarchy information and hierarchical position information extracted in step S1 as auxiliary judgment criteria. On the basis of vectorized semantic matching, structured constraints are superimposed to effectively suppress semantic drift and improve matching accuracy.

[0056] S3, Multi-dimensional data extraction. Combined with... Figure 4 As shown, based on the cross-resource alignment results output in step S2, the number of core information points is extracted from the benchmark resource. Number of logical levels Effective number of details Number of supporting evidence Text word count and redundancy ratio As an evaluation benchmark, the number of core information coverage points extracted from the content matched with the target resources is used. Number of logical levels covered Effective number of details Number of supporting evidence Text word count and redundancy ratio Of the six original indicators, the number of words in the text is... and It is used as a normalization factor in the calculation of the detail supplement density multiple in the subsequent step S4 to eliminate the systematic impact of the difference in the length of different resources on the detail density evaluation. It is not used as an independent evaluation dimension in the weighted comprehensive calculation.

[0057] S4. Rich metric based on multiples. Combined with Figure 5As shown, based on the multi-dimensional quantitative data output in step S3, five individual multiple indicators are calculated and weighted to obtain the overall richness multiple. The five multiple indicators quantify the resource richness from five orthogonal dimensions: coverage breadth, structural depth, detail density, sufficiency of argumentation, and information purity, forming a complete evaluation indicator space.

[0058] S5, cross-language resource alignment. Combined with... Figure 5 As shown, a multilingual pre-trained model is used to map teaching resource texts in different languages ​​to a unified multilingual semantic space. Cross-lingual similarity calculation is used to achieve semantic alignment between source language baseline knowledge points and target language resource content. After translation verification to ensure alignment quality, a unified representation structure for cross-lingual knowledge points is established. The four-tuple structure and evaluation system established in steps S1 to S4 of cross-lingual alignment reuse ensure that cross-lingual resources are also included in a unified quantitative evaluation framework, achieving seamless application of the same evaluation criteria across different language resources.

[0059] S6. Knowledge graph is dynamically updated. Combined with... Figure 6 As shown, a resource monitoring mechanism is established to detect changes in teaching resources. Incremental knowledge extraction is performed on the changed content, and entities and relationships in the knowledge graph are updated based on the extraction results. Consistency maintenance of the knowledge graph is also completed to ensure the timeliness and accuracy of the knowledge system. After a dynamic update is triggered, the four-stage closed-loop processing steps S1 to S4 are re-executed for the affected knowledge points. A difference-driven cascading update strategy is adopted, re-evaluating only the affected knowledge points and their related knowledge points to ensure that the updated evaluation results are consistent and comparable with the original system.

[0060] The steps S1 to S4 described above constitute the core four-stage closed-loop processing architecture of this invention. The key technical feature of this architecture lies in the strict data dependencies and logical progression between each stage. The four-tuple structured output of step S1 provides multi-dimensional alignment anchors and structured constraints for semantic matching in step S2. The precise alignment result of step S2 defines the accurate data collection boundaries for multi-dimensional data extraction in step S3. The six original indicators of step S3 provide a complete data foundation for the multiple calculation in step S4, with the text word count indicator indirectly participating in the evaluation through the embedded detail density formula. This tightly coupled closed-loop design ensures that the information transmission throughout the entire chain from knowledge extraction to quantitative evaluation is distortion-free, enabling the final comprehensive richness multiple to truly, objectively, and traceably reflect the richness differences of teaching resources, overcoming the systemic defects of poor overall results caused by the independent operation of each technical link in existing technologies.

[0061] This invention can be widely applied to teaching resource management, knowledge system construction, and resource quality assessment in various teaching scenarios such as higher education, vocational education, online education, corporate training, and international education. The core technologies involved in this invention include: semantic understanding and generation technology of Large Language Models (LLMs), Natural Language Processing (NLP) technology, knowledge graph construction technology, vector space models and semantic matching technology, cross-lingual information retrieval and alignment technology, machine learning and deep learning technologies, etc.

[0062] Combination Figure 7 and Figure 8 As shown, the intelligent fusion and quantitative evaluation system for multi-source teaching resources based on a large language model provided by this invention includes the following functional modules: Resource Access Module: This module is responsible for receiving and preprocessing multi-source teaching resources. Supported resource types include presentation slides (PPT / PPTX), e-textbooks (PDF / Word), academic papers, supplementary materials, etc. PyPDF2 is used for PDF text extraction, python-pptx for PPT parsing, and jieba is used for Chinese word segmentation to prepare analyzable text data for subsequent processing.

[0063] Knowledge Extraction Engine: An intelligent knowledge point extraction module based on a large language model (such as GPT-4o or Claude 3). Through carefully designed prompt engineering, it guides the large language model to accurately identify and extract knowledge points from different types of teaching resources. The extraction results are output in a structured form, including four dimensions: core concepts, related elements, logical hierarchy, and hierarchical position.

[0064] Semantic matching engine: A cross-resource knowledge point matching module based on vector space models and deep semantic understanding. It employs sentence transformers or large language model embedding techniques to calculate semantic vectors and achieves accurate matching through cosine similarity calculation. A similarity threshold is set. Meanwhile, a large language model is used to generate synonym replacement words to avoid missed matches.

[0065] Quantitative Evaluation Engine: A module for quantifying the richness of teaching resources based on a five-dimensional indicator system. It automatically calculates the core information coverage multiple, logical hierarchy completeness multiple, detail supplementation density multiple, supporting evidence quantity multiple, and redundancy correction coefficient, and obtains the overall richness score through weighted synthesis.

[0066] Multilingual Alignment Model: Utilizing the multilingual processing capabilities of a pre-trained multilingual model, semantic-level alignment is achieved between teaching resources in different language versions. A cross-language knowledge point mapping mechanism is established to construct a unified knowledge representation space.

[0067] Dynamic update module: A knowledge graph maintenance module based on an incremental processing mechanism, which realizes the detection of changes in teaching resources, incremental knowledge extraction, and related updates.

[0068] Example 1: Verification was conducted using a university's "Computer Architecture" course as the experimental subject.

[0069] The four-stage closed-loop processing procedure is as follows.

[0070] The baseline knowledge point structuring stage involves format parsing and text extraction of 32 hours of PPT courseware. The python-pptx tool is used for PPT parsing, and after text cleaning to remove redundant formatting marks, a targeted prompt template is designed to guide the large language model in identifying and extracting knowledge points from the courseware content. The prompt template includes a task description, JSON format output specifications, explanations of the three extraction dimensions, and annotation examples. A four-tuple structured representation is generated for each extracted knowledge point. For example, the structured tuple for the knowledge point "pipeline hazard" is k = ("pipeline hazard", {"data hazard", "control hazard", "structural hazard", "forward technique"}, 3, "processor design / instruction pipelining / hazard handling"), where the core concept is "pipeline hazard," the related elements include three hazard types and forward technique, and the logical level of 3 indicates that it belongs to the third level of sub-concept. The hierarchical position indicates the path of this knowledge point in the overall knowledge structure. After standardization processing such as terminology normalization, redundancy elimination, relation annotation, and index construction, 45 structured knowledge point tuples are output.

[0071] In the cross-resource precise matching stage: based on the 45 structured knowledge point tuples output from the previous stage, preprocessing and content segmentation were performed on both Chinese and English textbooks. A series of sentence-transformer embedding models were used to transform the baseline knowledge points and textbook text fragments into high-dimensional vector representations. Initial screening was performed using cosine similarity calculation with a threshold τ≥0.85, followed by synonym expansion to enhance recall. In the candidate result purification stage, a large language model was used in conjunction with the logical hierarchy and hierarchical position information extracted in the previous stage for refined semantic relevance determination. Taking the knowledge point "assembly line adventure" as an example, mismatched content in the textbook that contained the keyword "adventure" but belonged to other contexts was excluded based on its logical hierarchy and hierarchical position information. Finally, textbook matching content corresponding one-to-one with the baseline knowledge points was generated, achieving a Chinese semantic matching accuracy of 0.94 and an English semantic matching accuracy of 0.90.

[0072] Multi-dimensional data extraction stage: Based on the cross-resource alignment results output in the previous stage, a large language model is used to automatically extract various quantitative indicators in structured JSON format, obtaining the number of core information points, number of logical levels, number of effective details, number of supporting evidence, text word count, and redundancy ratio of the benchmark and target resources. The text word count is used in subsequent steps to normalize the detail density calculation, eliminating the impact of length differences on the evaluation.

[0073] The rich quantification phase based on multiples: Based on the multi-dimensional quantitative data output from the previous phase, five individual multiple indicators are calculated, using the default weight configuration w. N =0.30、w D =0.25、w M =0.20、w E =0.15、w R =0.10 for weighted aggregation. The results show that the two textbooks exhibit higher richness than PPT in most dimensions, especially in the multiple of supporting evidence (greater than 2.0) and the multiple of detail density (greater than 1.8). The overall multiple of the English textbook is 1.62, and the overall multiple of the Chinese textbook is 1.52, both of which belong to the level of "target resources are significantly richer than benchmark resources".

[0074] Cross-language resource alignment stage: The XLM-RoBERTa multilingual pre-trained model is used to map the texts of Chinese and English textbooks to a unified multilingual semantic space. The four-tuple structure and evaluation system established in the previous four stages are reused. Semantic alignment of Chinese and English knowledge points is achieved through cross-language similarity calculation. Translation verification ensures alignment quality, and a unified representation structure for cross-language knowledge points is established. The cross-language alignment F1 score reaches 0.86.

[0075] Knowledge graph dynamic update phase: Establish a resource monitoring mechanism to detect changes to courseware and textbooks through version identifiers and content hash values. Perform incremental knowledge extraction on the changed content to update entities and relationships in the knowledge graph. After a change is triggered, re-execute the four-stage closed-loop processing on the affected knowledge points to ensure that the updated evaluation results remain comparable to historical evaluations.

[0076] Experimental dataset: 32 hours of lecture slides, approximately 450 pages of Chinese textbook, approximately 520 pages of English textbook, and 45 core knowledge points (covering 12 instruction set architectures, 10 processor designs, 9 memory systems, 7 input / output systems, and 7 parallel processing systems).

[0077] Experimental results: Knowledge extraction F1 score: 0.90; Chinese semantic matching accuracy: 0.94; English semantic matching accuracy: 0.90; Cross-language alignment F1 score: 0.86; Average processing time: 45 seconds per knowledge point; API cost: approximately 2.80 yuan per knowledge point; Complete course analysis: 68 minutes, total cost 252 yuan.

[0078] Multiplier analysis results: Both textbooks showed greater richness than PPT in most dimensions, especially in the number of supporting evidence ( ) and density of detail ( The advantages are obvious in this aspect. The overall multiple of English textbooks (1.62) is slightly higher than that of Chinese textbooks (1.52).

[0079] The experimental results fully verify the effectiveness of the method of the present invention in knowledge extraction, semantic matching, cross-language alignment and rich metrics.

[0080] This invention fundamentally solves the long-standing problem of knowledge fragmentation in education by using intelligent knowledge point extraction and cross-resource alignment technology driven by a large language model. The system can automatically identify and associate knowledge points scattered across different resources (PPTs, textbooks, papers, etc.), constructing a complete knowledge system framework and transforming fragmented teaching resources into an organic whole.

[0081] Example 2: Taking the "Introduction to Artificial Intelligence" course on an online education platform as an application scenario, this course provides transcripts of teacher-recorded video lectures as anchor resources. Target resources include three Chinese textbooks from different publishers and two English reference books. This example focuses on demonstrating the application of this invention in a scenario of horizontal comparison and evaluation of multiple target resources.

[0082] After performing a four-stage closed-loop process, the overall richness multiple of each target resource relative to the anchor resource is obtained. By comparing the overall multiple and individual multiples horizontally, a quantitative decision-making basis is provided for teachers to select the optimal supplementary teaching materials. For example, textbook A performs best in core information coverage multiple but has a low detail density multiple, textbook B excels in supporting evidence multiple, and textbook C is best in logical hierarchy completeness multiple. Teachers can selectively choose different textbooks as supplementary materials for different teaching stages according to the course teaching objectives.

[0083] This embodiment fully demonstrates the technical advantages of the four-stage closed-loop processing architecture in scenarios involving horizontal comparison of multiple target resources. Since all target resources undergo the same closed-loop processing flow, employing a unified four-tuple structured representation and a five-dimensional multiple evaluation system, the evaluation results between resources are strictly comparable and consistent, avoiding comparison biases caused by the use of different evaluation standards.

[0084] This invention innovatively proposes a five-dimensional quantitative evaluation index system, covering five key dimensions: core information coverage, logical hierarchy completeness, detail supplementation density, supporting evidence quantity, and redundancy correction coefficient. This achieves a scientific quantitative assessment of the richness of teaching resources. For the first time, it elevates the evaluation of teaching resources from qualitative judgment to quantitative analysis; it provides intuitive and interpretable evaluation results through a multiplier calculation method; and this evaluation system has strong universality and portability, and can be widely applied to different courses and subject areas.

[0085] This invention leverages the multilingual processing capabilities of a large language model to establish semantic bridges between multilingual teaching resources, enabling intelligent integration of teaching resources in different language versions. It achieves semantic-level alignment based on a unified multilingual semantic space; supports accurate mapping and translation verification of professional terminology; and constructs a cross-language knowledge graph, supporting multilingual retrieval and querying. This effectively reduces language barriers for international students and improves the efficiency of cross-language utilization of high-quality teaching resources. This invention establishes a complete incremental update mechanism that reflects changes in teaching resources in real time. Incremental processing based on change detection avoids full recalculation; intelligent conflict detection and resolution mechanisms; and automated consistency maintenance ensure the timeliness and accuracy of the knowledge graph and reduce system maintenance costs.

Claims

1. A method for intelligent fusion of multi-source teaching resources based on a large language model, characterized in that, Includes the following steps: S1. Structured Analysis of Baseline Knowledge Points: This involves format parsing and text extraction of anchored teaching resources, and standardizing knowledge points. This includes designing prompt templates to guide the extraction of knowledge points from a large language model, generating structured tuples for each knowledge point. ,in Indicates the core concept, Indicates related elements, Represents logical hierarchy, Indicates hierarchical position; This step involves representing the generated quadruple structure as corresponding encoded information, providing multi-dimensional alignment anchors for cross-resource matching; S2. Cross-resource precise matching: Preprocessing and content segmentation of the target teaching resources are performed. A pre-trained text embedding model is used to transform the baseline knowledge points and target resource text fragments into high-dimensional vectors. Semantic matching between the two is achieved through cosine similarity calculation, and a similarity threshold is set. The process involves screening, then synonym expansion to enhance recall, and purification of candidate results to achieve precise alignment. S3. Multi-dimensional data extraction: Extracting the number of core information points from benchmark resources. Number of logical levels Effective number of details Number of supporting evidence Text word count and redundancy ratio As an evaluation benchmark, the number of core information coverage points extracted from the content matched with the target resources is used. Number of logical levels covered Effective number of details Number of supporting evidence Text word count and redundancy ratio ; S4. Richness Measurement Based on Individual Component Multiples: Calculate the five individual component multiples and obtain the overall richness multiple through weighted summation. ; It is a multi-dimensional indicator used to calculate the coverage multiple of core information. It is a multiple of logical hierarchy integrity. It is a detail supplement density multiple, This is evidence of multiples in quantity. It is a redundancy correction factor; S5. Cross-language resource alignment: A multilingual pre-trained model is used to map teaching resource texts in different languages ​​to a unified multilingual semantic space. Cross-language similarity calculation is used to achieve semantic alignment between the source language baseline knowledge points and the target language resource content. After translation verification to ensure the alignment quality, a unified representation structure for cross-language knowledge points is established. S6. Dynamic Update of Knowledge Graph: Establish a resource monitoring mechanism to detect changes in teaching resources, extract incremental knowledge from the changed content, update entities and relationships in the knowledge graph based on the extraction results, and complete the consistency maintenance of the knowledge graph to ensure the timeliness and accuracy of the knowledge system.

2. The intelligent fusion method for multi-source teaching resources according to claim 1, characterized in that, The anchored teaching resources mentioned in step S1 include one or more of the following formats: PPT / PPTX, PDF, DOC / DOCX, and HTML. The standardization process includes terminology normalization, redundancy elimination, relation annotation, and index construction, wherein index construction involves converting knowledge points into a combination of "short keywords + complete phrases". The prompt template includes a task description, output format specifications, extraction dimension description, and example guidance elements. The output format is JSON. The prompt template is based on a knowledge point granularity control strategy. By setting extraction granularity parameters, the abstraction level of knowledge points is controlled to ensure the consistency of semantic granularity of the extracted knowledge points and avoid deviations in subsequent matching and evaluation due to uneven granularity. The dimensions for extracting knowledge points include core concepts, key elements, and logical frameworks.

3. The intelligent fusion method for multi-source teaching resources according to claim 1, characterized in that, The segmentation strategy for dividing the target resource content described in step S2 comprehensively considers natural paragraph boundaries, chapter structure, semantic integrity, and length limitations of vectorization processing; The pre-trained text embedding model includes one or more of the following: general text embedding model, multilingual embedding model, and domain-specific embedding model. The formula for calculating the cosine similarity is: in, The embedding vector of the benchmark knowledge points. The embedding vector of the target resource text fragment. Describes the L2 norm of a vector; The synonym expansion includes term synonym expansion, abbreviation expansion, and related concept association based on knowledge graphs; The candidate result purification includes using a large language model to evaluate semantic relevance, eliminating mismatched content, integrating scattered related content, and generating target resource content that corresponds one-to-one with the benchmark knowledge points.

4. The intelligent fusion method for multi-source teaching resources according to claim 1, characterized in that, In step S4, the number of words in the text and It is used as a normalization factor in the calculation of detail supplement density to eliminate the impact of differences in the length of different resources on the evaluation of detail density. It is not used as an independent evaluation dimension in the weighted comprehensive calculation.

5. The intelligent fusion method for multi-source teaching resources according to claim 1, characterized in that, This method includes methods based on comprehensive richness multiples. Classification of resource richness levels: The target resources are significantly more abundant than the benchmark resources; 1.2 < The target resources are relatively abundant; 0.8 < The target resources are comparable in abundance to the benchmark resources; 0.8, the target resources are relatively simple.

6. The intelligent fusion method for multi-source teaching resources according to claim 1, characterized in that, The multilingual pre-trained model mentioned in step S5 includes one or more of mBERT, XLM-RoBERTa, and multilingual-e5; The translation verification involves translating the matched target language content into the source language, comparing the semantic similarity with the original source language's baseline knowledge points, and identifying and marking translation deviations or conceptual inconsistencies. The unified representation structure for cross-language knowledge points is a unique node in the knowledge graph corresponding to each knowledge point. Different language versions are stored as multilingual attributes of the node, while preserving the mapping relationship and alignment confidence between language versions. It also includes the construction of a cross-language mapping table between technical terms and the target language.

7. The intelligent fusion method for multi-source teaching resources according to claim 1, characterized in that, The resource monitoring mechanism described in step S6 achieves change detection by establishing version identifiers and content hash values ​​for access resources. The types of changes identified include added content, modified content, and deleted content. The incremental knowledge extraction process only processes the changed resource fragments, reusing the existing knowledge point structure and relationships; The knowledge graph consistency maintenance includes verifying structural integrity, handling knowledge point version conflicts, recalculating affected quantitative indicators and multiplier scores, and updating the search index.

8. A multi-source teaching resource intelligent fusion and quantitative evaluation system based on a large language model, implementing the method of any one of claims 1-7, characterized in that, include: The resource access module is used to receive multi-source heterogeneous teaching resources, perform format parsing, text extraction and preprocessing on the resources, and output analyzable text data. The knowledge extraction engine is built on a large language model, including the GPT series and Claude series models. It accurately extracts knowledge points from preprocessed anchored resources through preset prompt templates, generates knowledge points into structured tuples, and completes the standardization of knowledge points. The semantic matching engine, built on the vector space model and deep semantic understanding, realizes the vectorized representation of benchmark knowledge points and target resource text fragments, cosine similarity calculation, threshold filtering, synonym expansion and content purification, and completes accurate cross-resource matching. The text embedding models include text-embedding-ada-002, the sentence-transformers series, paraphrase-multilingual-mpnet-base-v2, or domain-specific embedding models finely tuned for specific disciplines. The quantitative evaluation engine is used to automatically extract multi-dimensional quantitative indicators from the matching content of benchmark resources and target resources, calculate five individual multiple indicators and the comprehensive richness multiple, and output quantitative evaluation results and resource richness level. The multilingual alignment module is built using a multilingual pre-trained model to achieve cross-lingual semantic alignment, translation verification, and unified representation of cross-lingual knowledge points for teaching resources in different languages. The dynamic update module is used to establish a teaching resource monitoring mechanism to complete resource change detection, incremental knowledge extraction, knowledge graph entity and relationship update, and knowledge graph consistency maintenance. The closed-loop scheduling controller is used to coordinate the data flow and execution order between modules to form a four-stage serial processing pipeline. It also reactivates the processing flow of the corresponding stage when the dynamic update is triggered. This includes providing a standardized inter-module communication interface through which each module realizes the structured transmission of data, ensuring the data consistency and atomicity of the four-stage closed-loop processing. The various modules of this system work together to achieve intelligent integration and quantitative evaluation of multi-source teaching resources, output a unified knowledge graph and quantitative evaluation results, and adopt a modular architecture design. It supports plug-and-play use of multiple large language models, flexible access to multiple document formats including PPT / PPTX, PDF, DOC / DOCX, and HTML, and seamless expansion to multiple languages.

9. A computer device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the computer program to implement the intelligent fusion and quantitative evaluation method for multi-source teaching resources based on a large language model as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the intelligent fusion and quantitative evaluation method for multi-source teaching resources based on a large language model as described in any one of claims 1-7.