Special subject teaching material Agent construction method, system and device and medium
By constructing a specialized subject textbook Agent system, the system achieves the association of three-dimensional data: textbooks, knowledge, and questions. This solves the problem of insufficient multi-dimensional data integration, improves teaching and learning efficiency, and provides personalized learning support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-14
AI Technical Summary
Existing teaching resources lack the integration and utilization of multi-dimensional data such as subject textbooks, subject question banks, subject backgrounds, and subject applications, resulting in insufficient data processing, making it difficult to meet the needs of teaching and learning. Furthermore, the lack of an intelligent three-dimensional data association structure makes it impossible to provide personalized and high-quality learning support.
By constructing a specialized subject textbook Agent system, data preprocessing, Agent system construction, and regression learning are performed to achieve the association of three-dimensional data: textbooks, knowledge, and questions. A large model is used to optimize the content and structure, and a regression learning mechanism is introduced to automatically complete abnormal questions.
It improved the utilization rate of subject data, optimized the teaching and learning process, enhanced learning efficiency and accuracy, reduced computing resource consumption, and improved user experience.
Smart Images

Figure CN121860003A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of education and artificial intelligence technology, specifically to a method, system, device, and medium for constructing a special subject teaching material agent. Background Technology
[0002] In the field of education, with the rapid development of information technology, how to efficiently utilize subject textbook data to assist teaching and learning has become a research hotspot. Currently, although some teaching systems and learning tools exist, there are still shortcomings in integrating multi-dimensional data from specific subjects, optimizing data processing, and providing precise services for student learning.
[0003] Existing teaching resources are mostly presented in a single dimension, such as providing only textbook texts or simple exercises. They lack the integration and utilization of multi-dimensional data, including subject textbooks, subject question banks, subject background, and subject applications, making it difficult to fully meet the needs of teaching and learning. Moreover, the processing of this data is not refined enough, and effective preprocessing and in-depth processing have not been carried out, resulting in the failure to fully realize the usability and value of the data.
[0004] In terms of application in teaching scenarios, traditional teaching methods mainly rely on teachers to manually sort out textbook knowledge and design questions, lacking intelligent three-dimensional data association structure construction, which makes the connection between textbooks, knowledge and questions not close enough, which is not conducive to students' systematic mastery of knowledge.
[0005] Furthermore, existing learning tools lack effective feedback and optimization mechanisms during the student learning process. When students encounter problems, it is difficult to accurately pinpoint the reasons for abnormal answers and optimize the knowledge system in a timely manner, making it difficult to provide personalized, high-quality learning support.
[0006] Therefore, how to comprehensively analyze and cover the teaching content of specialized subject textbooks, improve the utilization rate of subject data, and optimize the teaching and learning process are technical problems that urgently need to be solved. Summary of the Invention
[0007] The technical objective of this invention is to provide a method, system, device, and medium for constructing a subject-specific textbook agent, in order to solve the problem of how to comprehensively analyze and cover the teaching content of subject-specific textbooks, improve the utilization rate of subject data, and optimize the teaching and learning process.
[0008] The technical objective of this invention is achieved as follows: a method for constructing a specialized subject textbook agent, the specific method of which is as follows:
[0009] Data preprocessing: Collect subject-specific textbook data including textbook texts, question banks, background information sets, and application cases. Clean the text data using regular expressions re.sub(r'https?: / / \S+|www\.\S+',",text) and re.sub(r'<.*?>',",text), removing interfering data including HTML tags and URL links. Define custom data processing prompts. The large model extracts keywords and knowledge fragments from the text data based on the data processing prompts and divides the content, organizing the subject-specific textbook data into a knowledge correspondence structure of original text, knowledge fragments, and keywords.
[0010] Constructing an Agent system: Based on the keyword-knowledge fragment-original text correspondence structure, construct a three-dimensional data association of textbooks, knowledge, and questions. At the same time, optimize the content and structure of textbooks, knowledge, and questions through a large model to form an Agent knowledge system with a closed-loop relationship of extracting knowledge from textbooks, associating knowledge with questions, and analyzing textbook knowledge points through questions.
[0011] Regression learning: Collect user evaluations of the Agent's knowledge system responses, filter out abnormal problem scenarios, combine large models with manual analysis to identify the causes of anomalies, automatically optimize the Agent's knowledge system, and complete the Agent's knowledge system with abnormal problems.
[0012] As a preferred approach, the knowledge correspondence structure of the subject-specific textbook data is organized into original text, knowledge fragments, and keywords. Specifically, based on the TF-IDF algorithm (TF-IDF = TF × IDF), the core keywords in documents and paragraphs are mined by comparing the term frequency (TF) and inverse document frequency (IDF). The keywords are used as knowledge retrieval points to associate the original knowledge text. The text structure is understood through a large model to segment important knowledge fragments and match learning problems and learning needs in the application process.
[0013] As a preferred approach, the Agent system is constructed as follows: Students' knowledge mastery and learning progress in a subject are analyzed and statistically assessed. A large-scale model is used to analyze students' current learning progress in the subject, dividing the learning process into three stages: pre-learning preparation, in-learning question answering, and post-learning review. Knowledge fragments are extracted from subject-specific textbook data, and questions from a question bank are matched based on these fragments. The corresponding knowledge points in the textbook are located through question analysis, and the large-scale model is used to optimize the textbook's chapter structure, knowledge levels, and question difficulty.
[0014] Ideally, the pre-learning preparation stage should be as follows:
[0015] Based on the subject knowledge system, a large model is used to divide subject knowledge points according to the course learning plan, generate a course knowledge point topology map, and then obtain the hierarchical relationship between course knowledge points. This enables students to systematize their knowledge during the subject learning process, and the content of each lesson can be connected and linked together, promoting the construction of students' knowledge system in the subject learning process and improving students' learning efficiency. Based on the subject course learning arrangement, the learning content is synchronized with students in advance, enabling students to preview course knowledge, and the large model provides an AI learning assistant in the preview process.
[0016] Based on subject-specific textbook data, we screened learning data of the corresponding subjects in various formats (such as text, images, videos, etc.) from online platforms including educational websites and forums. After review and verification, low-quality and inaccurate information was removed to obtain the learning data of the corresponding subjects.
[0017] The acquired subject-specific learning data in various formats is converted into a unified format to facilitate subsequent processing and storage, and duplicate data records are removed, errors and inconsistencies in the data are corrected, and the data quality is improved. At the same time, the subject-specific learning data is labeled and classified according to subject, knowledge point or difficulty level to obtain pre-processed subject-specific learning data.
[0018] By setting up a block-based approach, the preprocessed subject-specific learning data is vectorized, and the vectorized preprocessed subject-specific learning data is stored in a vector database. At the same time, a subject-specific data knowledge base is established, and data keywords are automatically generated in the subject-specific data knowledge base. Semantic clustering and frequency statistics are performed based on the data keywords for subsequent text retrieval and semantic matching.
[0019] Custom prompts are used to accurately describe the characteristics and requirements of the required data. Leveraging the semantic understanding and matching capabilities of large models, based on vector space models or deep learning models, data content related to the prompts is retrieved from the subject data knowledge base. The retrieved data content is then integrated with authoritative subject data systems to supplement and improve the detailed content in the subject data knowledge system.
[0020] Based on the subject-specific data knowledge system, and according to the course learning schedule, learning content including course outlines, knowledge point introductions, and pre-study materials are synchronized to students in advance through the learning platform or mobile client. Based on students' learning history, learning ability, and interests, personalized learning content is pushed to students to improve the pertinence and effectiveness of pre-study.
[0021] Ideally, the answering stage during learning should be as follows:
[0022] Based on the course content, answer students' learning questions during the lecture, and provide detailed answers and comprehensive explanations of the knowledge points based on the knowledge points presented in the course.
[0023] Extracting core elements from the background information, such as textbook version, assessment focus (chapter / percentage), resource channels (past exam questions / online courses), and interaction methods; adopting a progressive logic of "textbook → application → resources → interaction," with bullet points (levels 1-4 headings) and key points highlighted (bold data / red prompts) to improve readability; and introducing quantitative indicators (60% of exam points come from textbooks, homework accounts for 40%, and error correction improves pass rate by 32%) to strengthen the scientific nature of the recommendations, and linking them to specific school resources to increase the depth of integration into business scenarios;
[0024] By combining school textbooks, assessment rules, and on-campus resources, personalized learning paths are obtained, and the prompts are designed with point-by-point logic and key point annotations to make complex suggestions clear and easy to understand; at the same time, on-campus data is integrated to support the effectiveness of the method through quantitative data (pass rate, test point ratio);
[0025] An adaptive routing mechanism based on data scale, combined with questions raised by students in class, optimizes the output content and provides answers to teaching content that is instructive and inspiring. Specifically, the adaptive routing mechanism works as follows: the input string is split by semicolons, each substring is stripped (leading and trailing whitespace characters are removed), and all empty strings are filtered out to obtain a list of valid questions, valid_questions; the analysis strategy identifier is determined based on the correspondence between the number of valid questions and a threshold.
[0026] As a preferred approach, semantic clustering and frequency statistics employ a two-dimensional analysis algorithm, as detailed below:
[0027] Semantic vectorization processing: The get_semantic_embeddings function converts the input list of questions into a numerical representation in the semantic vector space, providing a foundation for subsequent clustering;
[0028] Dynamic adaptive clustering: The `calculate_optimal_clusters` function is used to dynamically calculate the optimal number of clusters based on the number of questions (e.g., through adaptive algorithms such as the elbow rule), achieving intelligent matching between the number of clusters and the data size;
[0029] Frequency statistics with sequential records: The key operations completed synchronously through the dual dictionary structure are: the frequency_map dictionary counts the frequency of each issue; the order_map dictionary records the index position of the first occurrence of each issue (implemented by enumerate index during traversal). The dual dictionary structure design ensures that while counting frequencies, the first occurrence order information of issues in the original list is completely preserved.
[0030] Two-dimensional comprehensive sorting: The problem list is reordered using a composite sorting strategy: the primary sorting dimension is frequency (descending order, with high-frequency problems given priority); the secondary sorting dimension is the order of first appearance (ascending order, with earlier-appearing problems among those with the same frequency given priority).
[0031] As a preferred approach, regression learning is specifically as follows:
[0032] By optimizing the data-driven model capabilities and individual student learning efficiency, the large model learns based on the collected student learning data, establishes a data view of student learning status, and adds satisfaction feedback to the overall service process, collects student satisfaction evaluation comments, and supplements the large model's solution library for similar problems based on the content of student satisfaction evaluation comments.
[0033] Establish student learning portfolios, integrate multi-source data such as pre-learning preparation, in-learning quizzes, post-learning review and exam scores, and generate a visualized learning trajectory according to the three logics of time axis, knowledge point and ability dimension. Establish a time axis through semester-month-week to generate learning progress curves for different subjects.
[0034] By breaking down student learning into individual knowledge points within each subject area, a knowledge point mapping algorithm is used to create a radar chart of knowledge point mastery, clearly identifying students' weaknesses. The formula for the knowledge point mapping algorithm is as follows:
[0035]
[0036] Among them, M k C represents the mastery level of the k-th knowledge point; k E represents the number of correct answers (including homework and quizzes) for the k-th knowledge point; k The number of incorrect answers to the k-th knowledge point (including homework, quizzes, and wrong answers); T k Number of corrections completed for the kth knowledge point (correction of incorrect questions); T total,k The total number of incorrect answers for the kth knowledge point (the number of questions that need to be corrected); α and β are weighting coefficients (α+β=1, such as α=0.7 focusing on the correct answer rate, β=0.3 focusing on the effect of correcting incorrect answers);
[0037] The core competencies of a subject are identified from its knowledge points. Students' subject learning progress is used to infer their overall mastery of the subject knowledge. A weakness index algorithm is used to identify weak areas in the knowledge points. Furthermore, the knowledge points are ranked by difficulty (weighted by the number of incorrect answers and the percentage of points lost) to select the most challenging areas. k The top 20% of knowledge points are identified as weaknesses, and a subject ability assessment is generated accordingly. The formula for the weakness index algorithm is as follows:
[0038]
[0039] Among them, W k This represents the "weakness index" of the k-th knowledge point (the higher the value, the weaker the knowledge point). A represents the error rate (basic metric) of the k-th knowledge point; k : Represents the set of abilities associated with the k-th knowledge point (e.g., "geometric proof" associated with "logical reasoning" and "spatial imagination"); S a The weakness coefficient representing ability a (derived from teacher evaluations or behavioral data, S) a ∈[0,1], the higher the value, the weaker the ability); γ is the weight (e.g., γ = 0.6 focuses on error rate, 0.4 focuses on the weakness of association ability).
[0040] A specialized subject textbook agent construction system is provided, which implements the specialized subject textbook agent construction method described above; the system includes:
[0041] The data preprocessing unit collects subject-specific textbook data, including textbook texts, question banks, background information sets, and application cases. It cleans the text data using regular expressions `re.sub(r'https?: / / \S+|www\.\S+',”,text)` and `re.sub(r'<.*?>',”,text)`, removing interfering data such as HTML tags and URL links. It also defines custom data processing prompts. The large model extracts keywords and knowledge fragments from the text data based on these prompts and divides the content, organizing the subject-specific textbook data into a knowledge-correspondence structure of original text, knowledge fragments, and keywords.
[0042] The Agent system building unit is used to construct the three-dimensional data association of textbooks, knowledge, and questions based on the keyword-knowledge fragment-original text correspondence structure. At the same time, it optimizes the content and structure of textbooks, knowledge, and questions through a large model, forming an Agent knowledge system with a closed-loop relationship of extracting knowledge from textbooks, associating knowledge with questions, and analyzing textbook knowledge points through questions.
[0043] The regression learning unit is used to collect user evaluations of the Agent's knowledge system responses, filter out abnormal problem scenarios, combine large models with manual judgment to analyze the causes of anomalies, automatically optimize the Agent's knowledge system, and complete the Agent's knowledge system with abnormal problems.
[0044] An electronic device includes: a memory and at least one processor;
[0045] The memory contains computer programs;
[0046] The at least one processor executes the computer program stored in the memory, causing the at least one processor to execute the subject-specific textbook agent construction method as described above.
[0047] A computer-readable storage medium storing a computer program that can be executed by a processor to implement the subject-specific textbook agent construction method described above.
[0048] The method, system, device, and medium for constructing specialized subject teaching material agents according to the present invention have the following advantages:
[0049] (I) At the technical level, this invention constructs a subject agent system based on specialized subject textbooks. It uses data preprocessing technology to perform format processing, deduplication, and keyword extraction on multi-dimensional subject data. At the same time, it uses a large model combined with multi-modal parameter identification to process multi-modal data in subject teaching. At the application level, it focuses on the subject teaching scenarios of educational textbooks, constructs a three-dimensional data association structure of textbooks, knowledge, and questions, and optimizes the content and structure with the help of a large model. In addition, it introduces a regression learning mechanism. By collecting abnormal questions and answers through user evaluation, it combines the large model with manual judgment and the original textbook to perform regression analysis, automatically completing abnormal questions into the knowledge system and improving system performance.
[0050] (II) This invention constructs a three-dimensional data association structure of teaching materials, knowledge, and questions by preprocessing and secondary processing of special data, and designs a regression learning mechanism to improve the utilization efficiency of subject data, optimize the teaching and learning process, and provide more intelligent and efficient services for the education field. Attached Figure Description
[0051] The invention will be further described below with reference to the accompanying drawings.
[0052] Appendix Figure 1 A flowchart illustrating the method for constructing agents for specialized subject textbooks. Detailed Implementation
[0053] The following detailed description of the method, system, device, and medium for constructing a specialized subject textbook agent according to the present invention is provided with reference to the accompanying drawings and specific embodiments.
[0054] Example 1:
[0055] As attached Figure 1 As shown in the figure, this embodiment describes a method for constructing a specialized subject textbook agent, which is as follows:
[0056] S1. Data Preprocessing: Collect subject-specific textbook data, including textbook texts, question banks, background information sets, and application cases. Clean the text data using regular expressions re.sub(r'https?: / / \S+|www\.\S+',",text) and re.sub(r'<.*?>',",text), removing interfering data including HTML tags and URL links. Define custom data processing prompts. The large model extracts keywords and knowledge fragments from the text data based on the data processing prompts and divides the content, organizing the subject-specific textbook data into a knowledge correspondence structure of original text, knowledge fragments, and keywords.
[0057] S2. Constructing the Agent System: Based on the keyword-knowledge fragment-original text correspondence structure, construct the three-dimensional data association of textbooks, knowledge, and questions. At the same time, optimize the content and structure of textbooks, knowledge, and questions through a large model to form an Agent knowledge system with a closed-loop relationship of extracting knowledge from textbooks, associating knowledge with questions, and analyzing textbook knowledge points through questions.
[0058] S3. Regression Learning: Collect user evaluations of the Agent's knowledge system responses, filter out abnormal problem scenarios, combine large models with manual analysis to identify the causes of anomalies, automatically optimize the Agent's knowledge system, and complete the Agent's knowledge system with abnormal problems.
[0059] In this embodiment, step S1, which organizes the subject-specific textbook data into a knowledge correspondence structure of original text, knowledge fragments, and keywords, specifically involves: based on the TF-IDF algorithm (TF-IDF = TF × IDF), by comparing term frequency (TF) and inverse document frequency (IDF), core keywords in documents and paragraphs are mined. These keywords are used as knowledge retrieval points to associate with the original knowledge text. Furthermore, the text structure is understood through a large model to segment out important knowledge fragments and match learning problems and learning needs in the application process.
[0060] The subject-specific teaching material data in this embodiment includes not only traditional textbooks and supplementary books, but also a wealth of online course materials, teaching video scripts, and other forms of educational resources.
[0061] The construction of the Agent system in step S2 of this embodiment is specifically as follows: Based on the students' learning situation in the subject, and on the basis of sorting out subject-specific data, intelligent learning of students' full-cycle subject knowledge is provided based on keywords and high-quality knowledge point fragments; the students' knowledge mastery and learning progress in the subject are sorted out and statistically analyzed, and the students' learning progress in the current subject is analyzed through a large model, dividing the students' corresponding subject learning process into three stages: pre-learning preview, in-learning answering questions, and post-learning review; knowledge fragments are extracted from subject-specific textbook data, and questions are matched with questions in the question bank based on the knowledge fragments. The corresponding knowledge points in the textbook are located through question analysis, and the textbook chapter structure, knowledge level, and question difficulty are optimized using a large model.
[0062] In this embodiment, the pre-learning preparation stage is specifically as follows:
[0063] (1) Based on the subject knowledge system, the subject knowledge points are divided according to the course learning plan through a large model (first, a subject data knowledge base is established, and then, based on the authoritative subject data system, the knowledge and thinking patterns of the large model with a large number of parameters are used to recall the knowledge base data content of the data system, collect the data content to supplement the detailed content of the subject data system, and design prompt words, and sort out the knowledge points in the data system based on the AI capabilities of the large model). This generates a course knowledge point topology map, thereby obtaining the hierarchical relationship between course knowledge points, so that students can realize the systematization of knowledge in the subject learning process, and the content of each course can be connected and linked, promoting the construction of the knowledge system of students in the subject learning process and improving students' learning efficiency; and based on the subject course learning arrangement, the learning content is synchronized with students in advance to realize students' course knowledge preview, and the large model provides an AI learning assistant in the preview process.
[0064] (2) Based on the subject-specific textbook data, various formats (such as text, images, videos, etc.) of corresponding subject learning data are screened from online platforms including educational websites and forums. After review and verification, low-quality and inaccurate information is removed to obtain the corresponding subject learning data.
[0065] (3) Convert the acquired learning data of various subjects in a unified format to facilitate subsequent processing and storage, remove duplicate data records, correct errors and inconsistencies in the data, and improve the quality of the data; at the same time, label and classify the learning data of the corresponding subjects according to the subject, knowledge point or difficulty level to obtain the preprocessed learning data of the corresponding subjects.
[0066] (4) Vectorize the preprocessed subject learning data by setting up a block method, and store the vectorized preprocessed subject learning data in a vector database; at the same time, establish a subject data knowledge base, and automatically generate data keywords in the subject data knowledge base. Based on the data keywords, perform semantic clustering and frequency statistics for subsequent text retrieval and semantic matching.
[0067] (4) Custom prompts are used to accurately describe the characteristics and requirements of the required data. Utilizing the semantic understanding and matching capabilities of large-scale models, and based on vector space models or deep learning models, data content related to the prompts is retrieved from the subject-specific knowledge base. The retrieved data content is then integrated with authoritative subject-specific data systems to supplement and improve the detailed content within the subject-specific knowledge system. For example, for a specific knowledge point, relevant teaching cases and exercises are found in the knowledge base to enrich the teaching resources for that knowledge point. The reasoning capabilities of large-scale models are used to expand and extend the knowledge points in the data system, generating new knowledge content and further improving the subject-specific data system.
[0068] (6) Based on the subject data knowledge system, according to the course learning arrangement, the learning content including the course outline, knowledge point introduction and pre-study materials are synchronized to students in advance through the learning platform or mobile client. Based on the students' learning history, learning ability and interests, personalized learning content is pushed to students to improve the pertinence and effectiveness of pre-study.
[0069] In this embodiment, the learning and answering stage is specifically as follows:
[0070] (1) In conjunction with the course learning content, answer students’ learning questions during the course (provide students with model question and answer function during the course learning process, and answer each student’s learning questions in real time based on the current learning content), and provide detailed answers to questions and comprehensive explanations of knowledge points based on the knowledge points in the course lecture.
[0071] (2) Extract the core elements from the background: textbook version, assessment focus (chapter / percentage), resource channels (past exam questions / online courses), and interaction methods; adopt the progressive logic of "textbook → application → resources → interaction", mark points (levels 1-4 headings) and highlight key points (bold data / red prompts) to improve readability; and introduce quantitative indicators (60% of the test points come from the textbook, homework accounts for 40%, and the review of wrong questions improves the pass rate by 32%) to strengthen the scientific nature of the suggestions and link them with specific resources within the school to improve the depth of business scenario integration;
[0072] (3) Combine school textbooks, assessment rules and school resources to obtain personalized learning paths, and design point-by-point logic and key point marking in the prompt words to make complex suggestions clear and easy to understand; at the same time, access school data and support the effectiveness of the method through quantitative data (pass rate, test point ratio);
[0073] (4) An adaptive routing mechanism based on data scale, combined with the questions raised by students in class, reasonably optimizes the output content and provides teaching content answers with guiding and inspiring significance; the adaptive routing mechanism is as follows: the input string is split by semicolons, the strip operation (removing leading and trailing whitespace characters) is performed on each split substring, and all empty strings are filtered out to obtain a list of valid questions, valid_questions; the analysis strategy identifier is determined according to the correspondence between the number of valid questions and the threshold.
[0074] In this implementation, semantic clustering and frequency statistics employ a two-dimensional analysis algorithm, as detailed below:
[0075] ① Semantic vectorization processing: The get_semantic_embeddings function is used to convert the input list of questions into a numerical representation in the semantic vector space, providing a foundation for subsequent clustering;
[0076] ② Dynamic adaptive clustering: Using the calculate_optimal_clusters function, the optimal number of clusters is dynamically calculated based on the number of questions (e.g., through adaptive algorithms such as the elbow rule), to achieve intelligent matching between the number of clusters and the data size;
[0077] ③ Frequency statistics with sequential records: The key operations completed synchronously through the dual dictionary structure are: the frequency_map dictionary counts the frequency of each question; the order_map dictionary records the index position of the first occurrence of each question (implemented by enumerate index during traversal). The dual dictionary structure design ensures that while counting the frequency, the first occurrence order information of the questions in the original list is completely preserved.
[0078] ④ Two-dimensional comprehensive sorting: The problem list is reordered using a composite sorting strategy: the primary sorting dimension is frequency (descending order, with high-frequency problems given priority); the secondary sorting dimension is the order of first appearance (ascending order, with earlier-appearing problems among those with the same frequency given priority).
[0079] In this embodiment, during the post-learning stage, in conjunction with the course schedule, specialized review support for course knowledge points is provided during students' course review process. An AI knowledge assistant is provided to answer students' learning questions by combining a large model and subject knowledge. Based on the students' questions, the system analyzes the students' mastery of knowledge points and generates post-class assessments to comprehensively improve the effectiveness of students' post-class review and enhance their knowledge mastery.
[0080] This embodiment is based on a large model, designs prompts, and conducts a specific analysis of students' learning status and problems to rank students' learning behaviors and provide guidance on students' learning problems.
[0081] Establish a multi-dimensional academic data integration and quantitative comparison system, integrating data such as course grades (major / elective / required / theoretical / practical), credit progress, and GPA, and mining characteristics through comparative analysis: such as elective course GPA being higher than required course GPA (reflecting self-learning motivation), completed credits accounting for 22% of graduation requirements, and current GPA being in the middle range of the major.
[0082] Enhancing scientific rigor by incorporating educational theories: using the theory of multiple intelligences to explain the advantages of physical education courses (bodily-kinesthetic intelligence), constructivist theory to connect the active learning characteristics of elective courses, and behaviorist theory to guide the decomposition learning method for weak theoretical courses.
[0083] Establish a diagnostic mechanism for individual weaknesses and provide precise intervention suggestions. For issues such as failing core theoretical courses or insufficient language skills, customize solutions that match professional backgrounds: such as using sports training plans as an analogy for course design to assist theoretical understanding, and recommending scenario simulation training to improve oral skills.
[0084] By aligning with the school's student development program, abstract suggestions are translated into actionable on-campus resources: for example, the "Oral Communication Workshop for Teacher Trainees" held every Wednesday afternoon on designated campuses provides action guidelines with specific times and locations.
[0085] It provides students with post-study development predictions and incentive guidance. Based on the average GPA data of the major, it predicts that after improving 1-2 weak courses, they can enter the top 50% of the major, and uses growth feedback to strengthen their learning motivation.
[0086] Establish a closed-loop analysis process that combines "data-driven approach, theoretical support, personalized implementation, and incentive guidance" to help improve students' learning efficiency.
[0087] The regression learning in step S3 of this embodiment is as follows:
[0088] S301. Through targeted optimization of data-driven model capabilities and individual student learning efficiency, the large model learns based on the collected student learning data, establishes a data view of student learning status, adds satisfaction feedback during the overall service process, collects student satisfaction evaluation comments, and supplements the large model's solution library for similar problems based on the content of student satisfaction evaluation comments.
[0089] This system integrates multi-source data on students' subject learning, including pre-learning preparation, in-learning mastery, post-learning review, and exam scores. Based on this diverse data, it categorizes and summarizes information by subject, time, and ability to create a traceable learning trajectory. It outputs visualized profiles to assist students in self-reflection, teachers in providing targeted guidance, and parents in understanding their children's learning progress.
[0090] S302. Establish student learning portfolios, integrate multi-source data such as pre-learning preparation, in-learning quizzes, post-learning review and exam scores, and generate a visualized learning trajectory according to the three logics of time axis, knowledge point and ability dimension. Establish a time axis through semester-month-week to generate learning progress curves for different subjects.
[0091] S303. By breaking down student learning into individual knowledge points within each subject area, a knowledge point mapping algorithm is used to create a radar chart of knowledge point mastery, clearly identifying students' weak areas. The formula for the knowledge point mapping algorithm is as follows:
[0092]
[0093] Among them, M k C represents the mastery level of the k-th knowledge point; k E represents the number of correct answers (including homework and quizzes) for the k-th knowledge point; k The number of incorrect answers to the k-th knowledge point (including homework, quizzes, and wrong answers); T k Number of corrections completed for the kth knowledge point (correction of incorrect questions); T total,k The total number of incorrect answers for the kth knowledge point (the number of questions that need to be corrected); α and β are weighting coefficients (α+β=1, such as α=0.7 focusing on the correct answer rate, β=0.3 focusing on the effect of correcting incorrect answers);
[0094] S304. Organize core subject competencies within subject knowledge points. Based on students' subject learning progress, infer their overall subject knowledge mastery. Identify weak areas using a weakness index algorithm. Then, rank knowledge points by difficulty (weighted by the number of incorrect answers and the percentage of points lost), and select the appropriate level (W). k The top 20% of knowledge points are identified as weaknesses, and a subject ability assessment is generated accordingly. The formula for the weakness index algorithm is as follows:
[0095]
[0096] Among them, W k This represents the "weakness index" of the k-th knowledge point (the higher the value, the weaker the knowledge point). A represents the error rate (basic metric) of the k-th knowledge point; k : Represents the set of abilities associated with the k-th knowledge point (e.g., "geometric proof" associated with "logical reasoning" and "spatial imagination"); S a The weakness coefficient representing ability a (derived from teacher evaluations or behavioral data, S) a ∈[0,1], the higher the value, the weaker the ability); γ is the weight (e.g., γ = 0.6 focuses on error rate, 0.4 focuses on the weakness of association ability).
[0097] This embodiment combines subject-specific textbook data. During the construction of the subject agent, it preprocesses the subject-specific data (including format processing, content classification, and data deduplication), and then, based on a large model and multimodal parameter recognition, processes, transforms, classifies, and extracts core data from the multimodal data within the subject-specific data, completing secondary processing and refinement of the subject-specific data. It also outlines the subject agent scenario, constructing a three-dimensional data association structure of textbooks, knowledge, and questions within the application scenario of subject teaching using educational textbooks. Knowledge is extracted from textbooks, questions are associated with knowledge, and textbook knowledge points are analyzed from questions. The large model optimizes the content and structure of textbooks, knowledge, and questions respectively. Simultaneously, a regression learning mechanism is added to the subject agent service and student learning process. By evaluating the effectiveness of agent question answers during user interaction, it collects question-and-answer scenarios with problems and poor answers. Based on the large model, combined with manual judgment and the original textbook, a statistical algorithm for subject learning is designed to perform regression analysis on questions, clarify the reasons for abnormal question answers, and establish automatic optimization logic to automatically complete abnormally answered questions into the subject agent's knowledge system, achieving the following effects:
[0098] ① Improved efficiency: Processing time reduced by 30%-50% (through adaptive strategies);
[0099] ② Improved accuracy: Scene classification accuracy increased by 25%;
[0100] ③ Resource optimization was achieved: computing resource consumption was reduced by 40%;
[0101] ④ Improved user experience: Response speed increased by 60%.
[0102] Meanwhile, this embodiment solves the problems of difficulty in meeting students' personalized learning needs, excessive teaching burden on teachers, untimely updates and insufficient expansion of teaching materials, and incomplete and untimely evaluation of learning outcomes.
[0103] Example 2:
[0104] This embodiment provides a specialized subject textbook agent construction system, which is used to implement the specialized subject textbook agent construction method as described in Embodiment 1; the system includes:
[0105] The data preprocessing unit collects subject-specific textbook data, including textbook texts, question banks, background information sets, and application cases. It cleans the text data using regular expressions `re.sub(r'https?: / / \S+|www\.\S+',”,text)` and `re.sub(r'<.*?>',”,text)`, removing interfering data such as HTML tags and URL links. It also defines custom data processing prompts. The large model extracts keywords and knowledge fragments from the text data based on these prompts and divides the content, organizing the subject-specific textbook data into a knowledge-correspondence structure of original text, knowledge fragments, and keywords.
[0106] The Agent system building unit is used to construct the three-dimensional data association of textbooks, knowledge, and questions based on the keyword-knowledge fragment-original text correspondence structure. At the same time, it optimizes the content and structure of textbooks, knowledge, and questions through a large model, forming an Agent knowledge system with a closed-loop relationship of extracting knowledge from textbooks, associating knowledge with questions, and analyzing textbook knowledge points through questions.
[0107] The regression learning unit is used to collect user evaluations of the Agent's knowledge system responses, filter out abnormal problem scenarios, combine large models with manual judgment to analyze the causes of anomalies, automatically optimize the Agent's knowledge system, and complete the Agent's knowledge system with abnormal problems.
[0108] Example 3:
[0109] This embodiment also provides an electronic device, including: a memory and a processor;
[0110] The memory stores the instructions executed by the computer.
[0111] The processor executes the computer execution instructions stored in the memory, causing the processor to execute the subject-specific textbook agent construction method in any embodiment of the present invention.
[0112] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can be a microprocessor or any conventional processor.
[0113] Memory is used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart memory cards (SMC), secure digital cards (SD cards), flash memory cards, at least one disk storage device, flash memory devices, or other volatile solid-state storage devices.
[0114] Example 4:
[0115] This embodiment also provides a computer-readable storage medium storing multiple instructions, which are loaded by a processor to cause the processor to execute the subject-specific textbook agent construction method in any embodiment of the present invention. Specifically, a system or apparatus equipped with a storage medium may be provided, on which software program code implementing the functions of any of the above embodiments is stored, and the computer (or CPU or MPU) of the system or apparatus may read and execute the program code stored in the storage medium.
[0116] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of the present invention.
[0117] Storage media embodiments for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.
[0118] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby achieving the function of any of the embodiments described above.
[0119] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion unit connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion unit execute some and all of the actual operations, thereby realizing the function of any of the embodiments described above.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a specialized subject textbook agent, characterized in that, The method is as follows: Data preprocessing: Collect subject-specific textbook data including textbook texts, question banks, background information sets, and application cases. Clean the text data using regular expressions re.sub(r'https?: / / \S+|www\.\S+',",text) and re.sub(r'<.*?>',",text), removing interfering data including HTML tags and URL links. Define custom data processing prompts. The large model extracts keywords and knowledge fragments from the text data based on the data processing prompts and divides the content, organizing the subject-specific textbook data into a knowledge correspondence structure of original text, knowledge fragments, and keywords. Constructing an Agent system: Based on the keyword-knowledge fragment-original text correspondence structure, construct a three-dimensional data association of textbooks, knowledge, and questions. At the same time, optimize the content and structure of textbooks, knowledge, and questions through a large model to form an Agent knowledge system with a closed-loop relationship of extracting knowledge from textbooks, associating knowledge with questions, and analyzing textbook knowledge points through questions. Regression learning: Collect user evaluations of the Agent's knowledge system responses, filter out abnormal problem scenarios, combine large models with manual analysis to identify the causes of anomalies, automatically optimize the Agent's knowledge system, and complete the Agent's knowledge system with abnormal problems.
2. The method for constructing a specialized subject textbook agent according to claim 1, characterized in that, The specific knowledge correspondence structure of organizing subject-specific textbook data into original text, knowledge fragments, and keywords is as follows: Based on the TF-IDF algorithm (TF-IDF = TF × IDF), by comparing word frequency and inverse document frequency, core keywords in documents and paragraphs are mined. Keywords are used as knowledge retrieval points to associate knowledge with the original text. Through a large model, the text structure is understood, important knowledge fragments are segmented, and learning problems and learning needs in the application process are matched.
3. The method for constructing a specialized subject textbook agent according to claim 1, characterized in that, The Agent system is constructed by sorting out and statistically analyzing students' knowledge mastery and learning progress in a subject, analyzing students' current learning progress through a large model, and dividing the students' learning process in the corresponding subject into three stages: pre-learning preparation, in-learning question answering, and post-learning review. Furthermore, knowledge fragments are extracted from the textbook data of specific subjects, and questions in the question bank are matched based on the knowledge fragments. The corresponding knowledge points in the textbook are located through question analysis, and the textbook chapter structure, knowledge level and question difficulty are optimized using a large model.
4. The method for constructing a specialized subject textbook agent according to claim 3, characterized in that, The pre-learning preparation stage is as follows: Based on the subject knowledge system, the subject knowledge points are divided according to the course learning plan through a large model, generating a course knowledge point topology map, thereby obtaining the hierarchical relationship between course knowledge points, so that students can realize the knowledge system in the subject learning process, and the content of each course can be connected. Based on the subject course learning arrangement, the learning content is synchronized with students in advance, enabling students to preview the course knowledge, and providing an AI learning assistant in the preview process through the large model. Based on subject-specific textbook data, we screened learning data of the corresponding subjects in various formats from online platforms including educational websites and forums. After review and verification, we removed low-quality and inaccurate information and obtained the learning data of the corresponding subjects. The acquired subject-specific learning data in various formats is converted into a unified format to facilitate subsequent processing and storage, and duplicate data records are removed, and errors and inconsistencies in the data are corrected; at the same time, the subject-specific learning data is labeled and classified according to subject, knowledge point or difficulty level to obtain pre-processed subject-specific learning data; By setting up a block-based approach, the preprocessed subject-specific learning data is vectorized, and the vectorized preprocessed subject-specific learning data is stored in a vector database. At the same time, a subject-specific data knowledge base is established, and data keywords are automatically generated in the subject-specific data knowledge base. Semantic clustering and frequency statistics are then performed based on the data keywords. Custom prompt words are used to accurately describe the characteristics and requirements of the required data. Leveraging the semantic understanding and matching capabilities of large models, based on vector space models or deep learning models, data content related to the prompt words is retrieved from the subject data knowledge base. The recalled data will be integrated with authoritative subject data systems to supplement and improve the detailed content in the subject data knowledge system. Based on the subject-specific data knowledge system, and according to the course learning schedule, learning content including course outlines, knowledge point introductions, and pre-study materials are synchronized to students in advance through the learning platform or mobile client. Personalized learning content is also provided to students based on their learning history, learning ability, and interests.
5. The method for constructing a specialized subject textbook agent according to claim 3, characterized in that, The specific steps for answering questions during the learning process are as follows: Based on the course content, answer students' learning questions during the lecture, and provide detailed answers and comprehensive explanations of the knowledge points based on the knowledge points presented in the course. Extract the core elements from the background, such as textbook version, assessment focus, resource channels, and interaction methods; adopt a progressive logic of "textbook → application → resources → interaction", with point-by-point annotation and emphasis on key points; Quantitative indicators were introduced to enhance the scientific rigor of the recommendations; By combining school textbooks, assessment rules, and on-campus resources, personalized learning paths are obtained, and the prompts are designed with point-by-point logic and key point annotations to make complex suggestions clear and easy to understand; at the same time, on-campus data is integrated to support the effectiveness of the methods through quantitative data. An adaptive routing mechanism based on data scale, combined with questions raised by students in class, optimizes the output content and provides answers to teaching content that is instructive and inspiring. Specifically, the adaptive routing mechanism works as follows: the input string is split by semicolons, a strip operation is performed on each substring, and all empty strings are filtered out to obtain a list of valid questions, valid_questions; the analysis strategy identifier is determined based on the correspondence between the number of valid questions and a threshold.
6. The method for constructing a specialized subject textbook agent according to claim 1, characterized in that, Semantic clustering and frequency statistics employ a two-dimensional analysis algorithm, as detailed below: Semantic vectorization processing: The get_semantic_embeddings function converts the input list of questions into a numerical representation in the semantic vector space; Dynamic adaptive clustering: The `calculate_optimal_clusters` function is used to dynamically calculate the optimal number of clusters based on the number of questions, achieving intelligent matching between the number of clusters and the data size; Frequency statistics with sequential records: The key operations synchronized through a dual-dictionary structure are: the frequency_map dictionary counts the frequency of each issue; the order_map dictionary records the index position of the first occurrence of each issue. Two-dimensional comprehensive sorting: The problem list is reordered using a composite sorting strategy: the primary sorting dimension is frequency; the secondary sorting dimension is the order of first appearance.
7. The method for constructing a specialized subject textbook agent according to claim 1, characterized in that, The specifics of regression learning are as follows: By optimizing the data-driven model capabilities and individual student learning efficiency, the large model learns based on the collected student learning data, establishes a data view of student learning status, and adds satisfaction feedback to the overall service process, collects student satisfaction evaluation comments, and supplements the large model's solution library for similar problems based on the content of student satisfaction evaluation comments. Establish student learning portfolios, integrate multi-source data such as pre-learning preparation, in-learning quizzes, post-learning review and exam scores, and generate a visualized learning trajectory according to the three logics of time axis, knowledge point and ability dimension. Establish a time axis through semester-month-week to generate learning progress curves for different subjects. By breaking down student learning into individual knowledge points within each subject area, a knowledge point mapping algorithm is used to create a radar chart of knowledge point mastery, clearly identifying students' weaknesses. The formula for the knowledge point mapping algorithm is as follows: Among them, M k C represents the mastery level of the k-th knowledge point; k E represents the number of correct answers to the k-th knowledge point; k The number of incorrect answers for the k-th knowledge point; T k The number of times the correction for the kth knowledge point has been completed; T total,k The total number of incorrect answers for the k-th knowledge point; α and β are weighting coefficients; The core competencies of a subject are identified from its knowledge points. Students' subject learning progress is used to infer their overall knowledge mastery. A weakness index algorithm is used to identify weak areas in the knowledge points. Finally, a challenging ranking of knowledge points is implemented to select the most suitable candidates. k The top 20% of knowledge points are identified as weaknesses, and a subject ability assessment is generated accordingly. The formula for the weakness index algorithm is as follows: Among them, W k This represents the "weakness index" of the k-th knowledge point; A represents the error rate of the k-th knowledge point; k S represents the set of abilities associated with the k-th knowledge point; a γ represents the bottleneck coefficient of capability a; γ is the weight.
8. A specialized subject textbook agent construction system, characterized in that, This system is used to implement the subject-specific textbook agent construction method as described in any one of claims 1 to 1; the system includes: The data preprocessing unit collects subject-specific textbook data, including textbook texts, question banks, background information sets, and application cases. It cleans the text data using regular expressions `re.sub(r'https?: / / \S+|www\.\S+',”,text)` and `re.sub(r'<.*?>',”,text)`, removing interfering data such as HTML tags and URL links. It also defines custom data processing prompts. The large model extracts keywords and knowledge fragments from the text data based on these prompts and divides the content, organizing the subject-specific textbook data into a knowledge-correspondence structure of original text, knowledge fragments, and keywords. The Agent system building unit is used to construct the three-dimensional data association of textbooks, knowledge, and questions based on the keyword-knowledge fragment-original text correspondence structure. At the same time, it optimizes the content and structure of textbooks, knowledge, and questions through a large model, forming an Agent knowledge system with a closed-loop relationship of extracting knowledge from textbooks, associating knowledge with questions, and analyzing textbook knowledge points through questions. The regression learning unit is used to collect user evaluations of the Agent's knowledge system responses, filter out abnormal problem scenarios, combine large models with manual judgment to analyze the causes of anomalies, automatically optimize the Agent's knowledge system, and complete the Agent's knowledge system with abnormal problems.
9. An electronic device, characterized in that, include: Memory and at least one processor; The memory contains computer programs; The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the subject-specific textbook agent construction method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed by a processor to implement the subject-specific textbook agent construction method as described in any one of claims 1 to 7.