Humanities learning multi-modal adaptive companion learning method and system based on large language model
By adopting a multimodal adaptive learning companion method for humanities learning based on a large language model, the problem of insufficient deep understanding and interest perception in the humanities and social sciences fields of existing systems is solved, and personalized content push and dynamic adjustment are realized, thereby improving the intelligence level of the learning system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DALIAN HOUREN EDUCATION TECH CO LTD
- Filing Date
- 2025-10-31
- Publication Date
- 2026-07-21
AI Technical Summary
Existing learning systems struggle to deeply understand the connotations of humanities knowledge and accurately perceive students' interests in the humanities and social sciences. They also lack flexible adaptive adjustment capabilities, resulting in mechanical and rigid push strategies that fail to achieve individualized instruction and interest guidance.
We adopt a multimodal adaptive learning companion method for humanities learning based on a large language model. We generate a knowledge graph through deep semantic analysis, analyze students' multimodal interaction behavior, construct a dynamic interest tag set and interest degree model, calculate the knowledge gravity value, generate a personalized content push sequence, and optimize the model parameters based on student feedback.
It achieves a deep understanding of humanities knowledge and accurate perception of interests, generates scientific and reasonable personalized learning paths, dynamically adjusts content delivery, and improves learning efficiency and interest guidance.
Smart Images

Figure CN121684097B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of educational recommendation technology, and in particular to a multimodal adaptive learning companion method and system for humanities learning based on a large language model. Background Technology
[0002] With the deep integration of artificial intelligence technology into education, various intelligent learning systems and applications are becoming increasingly widespread. These systems aim to assist the teaching process and improve learning efficiency through technological means, and have made significant progress, particularly in areas such as exercise recommendation and path planning. This signifies an important trend in the transformation of educational informatization from tool-based applications to intelligent services.
[0003] However, most current learning systems, especially in the humanities and social sciences where deep semantic understanding is required, remain at a relatively rudimentary stage of intelligence, facing several fundamental technical bottlenecks. First, at the knowledge representation level, existing systems largely rely on pre-built, static structured knowledge bases, lacking the ability to deeply understand and automatically construct knowledge from massive amounts of unstructured text resources. This results in slow updates to their knowledge base, narrow coverage, and difficulty in capturing the complex logical connections and implicit relationships within humanities knowledge. Second, at the student perception level, most systems employ a single interaction mode, failing to fully utilize multimodal data to comprehensively and three-dimensionally depict students' dynamic learning states, particularly struggling to accurately and quantitatively capture and track students' evolving interests over time. Finally, at the core of adaptive decision-making, existing recommendation algorithms often remain at the level of linear difficulty adjustment based on answer accuracy or simple content association based on keyword matching, lacking a computable model that integrates dynamic interests, semantic connections of knowledge, and cognitive development patterns. This leads to mechanical and rigid recommendation strategies, failing to achieve true personalized learning and interest guidance, thus hindering both the stimulation of students' sustained intrinsic motivation for learning and the effective breaking of information silos to expand students' cognitive boundaries.
[0004] In conclusion, existing technologies are insufficient to construct an intelligent learning companion system that can deeply understand the connotations of humanities knowledge, accurately perceive students' interest levels, and make scientific and flexible dynamic adjustments based on this understanding. A solution to these problems is urgently needed. Summary of the Invention
[0005] This disclosure provides a multimodal adaptive learning companion method and system for humanities learning based on a large language model, in order to solve the technical problem in the existing technology of it being difficult to build an intelligent learning companion system that can deeply understand the connotation of humanities knowledge, accurately perceive students' interest status, and make scientific and flexible dynamic adjustments on this basis.
[0006] According to the first aspect of this disclosure, a multimodal adaptive learning companion method for humanities learning based on a large language model is provided, including: Input humanities texts from primary and secondary schools into a large language model, perform in-depth semantic analysis and structured processing, and generate a knowledge graph for the humanities field; Acquire student interaction behavior, analyze student interaction behavior through multimodal technology, extract student interest tags, calculate initial interest level, and construct dynamic interest tag set; Construct an interest update model, combine dynamic interest tag set to calculate the interest of relevant knowledge points, and obtain a comprehensive interest set; Traverse the knowledge graph of the humanities field, calculate the knowledge gravity value of each knowledge point by combining the interest degree comprehensive set, integrate all knowledge gravity values, and construct a knowledge gravity set; A push sequence is generated based on the knowledge gravity set to deliver personalized content. The push sequence includes key knowledge points and exploratory knowledge points. By tracking the interaction results between students and the pushed sequences, and adjusting the parameters in the interest update model based on the students' feedback behavior, the interest calculation is optimized.
[0007] According to a second aspect of this disclosure, a multimodal adaptive learning companion system for humanities learning based on a large language model is provided, comprising: The knowledge graph construction module is used to input humanities texts from primary and secondary schools into a large language model, perform deep semantic analysis and structured processing, and generate a humanities knowledge graph. An interest tag construction module is used to acquire student interaction behavior, analyze student interaction behavior through multimodal technology, extract student interest tags, calculate initial interest level, and construct dynamic interest tag set. An interest score calculation module is used to construct an interest score update model, combine a dynamic interest tag set to calculate the interest score of relevant knowledge points, and obtain an interest score comprehensive set. The knowledge gravity calculation and integration module is used to traverse the knowledge graph in the humanities field, calculate the knowledge gravity value of each knowledge point by combining the interest degree comprehensive set, integrate all knowledge gravity values, and construct a knowledge gravity set. A personalized content push sequence generation module is used to generate a push sequence based on a knowledge gravity set for personalized content push. The push sequence includes main knowledge points and exploratory knowledge points. The parameter adaptive optimization module is used to track the interaction results between students and the pushed sequence, and adjust the parameters in the interest update model according to the students' feedback behavior to optimize the interest calculation.
[0008] One or more technical solutions provided in this disclosure have at least the following technical effects or advantages: Inputting humanities texts from primary and secondary schools into a large language model, performing deep semantic analysis and structured processing to generate a humanities knowledge graph; acquiring student interaction behavior, analyzing student interaction behavior through multimodal technology, extracting student interest tags, calculating initial interest levels, and constructing a dynamic interest tag set; constructing an interest level update model, combining the dynamic interest tag set to calculate the interest level of relevant knowledge points, and obtaining a comprehensive interest level set; traversing the humanities knowledge graph, combining the comprehensive interest level set to calculate the knowledge gravity value of each knowledge point, integrating all knowledge gravity values, and constructing a knowledge gravity set; generating a push sequence based on the knowledge gravity set for personalized content push, the push sequence including main knowledge points and exploratory knowledge points; tracking the interaction results between students and the push sequence, adjusting the parameters in the interest level update model according to student feedback behavior, and optimizing interest level calculation. This solves the technical problem in existing technologies of the difficulty in constructing an intelligent learning companion system that can deeply understand the connotation of humanities knowledge, accurately perceive students' interest status, and make scientific and flexible dynamic adjustments based on this. It has achieved the technical effect of constructing a dynamic knowledge base with deep semantic understanding capabilities, realizing refined perception and quantitative tracking of students' interest status, and generating scientific and reasonable personalized learning paths based on this.
[0009] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0011] Figure 1 A flowchart illustrating the multimodal adaptive learning companion method for humanities learning based on a large language model provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of a multimodal adaptive learning companion system for humanities learning based on a large language model, provided in an embodiment of this application.
[0012] Figure labeling: Knowledge graph construction module 11, interest tag construction module 12, interest degree comprehensive calculation module 13, knowledge gravity calculation and integration module 14, personalized content push sequence generation module 15, parameter adaptive optimization module 16. Detailed Implementation
[0013] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0014] Example 1, the multimodal adaptive learning companion method for humanities learning based on a large language model provided in this disclosure, is referred to below. Figure 1 The methods include: S1: Input humanities texts from primary and secondary schools into a large language model, perform in-depth semantic analysis and structured processing, and generate a knowledge graph in the humanities field; Furthermore, step S1 also includes: Texts in the humanities field from primary and secondary schools are collected to construct an unstructured initial corpus. The initial corpus is then preprocessed to obtain a standardized corpus. The preprocessing includes format standardization and primary structure parsing. Select a large language model and perform domain-adaptive fine-tuning on it; The large language model, after domain-adaptive fine-tuning, is used to process the text of a standardized corpus to generate a set of knowledge point relationships. The text processing includes knowledge element extraction, semantic relationship mining, entity alignment, and semantic disambiguation. Generate a knowledge graph for the humanities field based on the set of relationships between knowledge points.
[0015] Specifically, we collected humanities texts from primary and secondary schools, including textbooks, curriculum standards, and authoritative academic literature, to construct an unstructured initial corpus. The initial corpus was then preprocessed to obtain a standardized corpus. The core preprocessing operations included format standardization and basic structure parsing. The main purpose was to transform heterogeneous document formats into a unified, machine-readable plain text sequence and to initially identify document metastructures such as chapter titles, paragraphs, and figure / table annotations, laying a data foundation for subsequent deep semantic analysis.
[0016] After acquiring the standardized corpus, the domain-adaptive fine-tuning of the large language model began. First, an existing DeepSeek general-purpose large language model was selected as the base model, and then domain-adaptive fine-tuning was performed. This was done because while the general-purpose large language model possesses broad general knowledge, its knowledge system has gaps with the specific contexts of the humanities domain in primary and secondary schools. Therefore, domain-adaptive fine-tuning was necessary. This process is equivalent to a targeted, professional intensive training of the large language model. Using portions of the standardized corpus and through instruction-based fine-tuning techniques, the large language model deeply internalizes the specialized terminology, conceptual systems, and internal logic of the humanities domain. Through this process, the large language model can accurately understand the complex knowledge within the humanities domain of primary and secondary schools, providing cognitive support for subsequent high-precision knowledge parsing tasks.
[0017] Subsequently, all data from the standardized corpus is input into the large language model to begin the core knowledge point relation set construction phase. This phase aims to deconstruct coherent natural language text into discrete, semantically rich knowledge units. The primary task involves knowledge element extraction steps, including entity recognition and concept extraction. The large language model performs scanning parsing on a sentence-by-sentence basis, utilizing its built-in named entity recognition and semantic role labeling capabilities to accurately locate and classify core knowledge units in the text. For example, when processing the sentence "Du Fu wrote 'Spring View' during the An Lushan Rebellion," the model can accurately identify "Du Fu" as a historical figure entity, "An Lushan Rebellion" as a historical event entity, and "'Spring View'" as a literary work entity. More importantly, it can extract core conceptual entities from abstract expressions, such as extracting the concept entity "feudal society" from "Dream of the Red Chamber reveals the decline of feudal society."
[0018] Building upon the identification of entities and concepts, the large language model further executes semantic relationship mining steps encompassing relation extraction and attribute induction. This step aims to construct semantic links between entities and concepts and enrich their connotative features. The model infers and formalizes the relationships between entities by analyzing syntactic structure and contextual semantics. Taking the aforementioned sentence as an example, it will construct structured triples such as "Du Fu - created during the An Lushan Rebellion" and "Du Fu - created - 'Spring View'". At a deeper level, it can infer implicit, high-value semantic associations such as "'Spring View' - background of creation - the An Lushan Rebellion" and "'Spring View' - expresses emotions - sorrow for the country and the times". In parallel, the large language model performs attribute induction on entity nodes, constructing feature profiles for nodes from scattered descriptions. For example, from narrative texts about Du Fu, it extracts key attributes such as "Zimei" and "Shaoling Yela" (his pen name is Shaoling Yela).
[0019] Finally, entity alignment and semantic disambiguation are performed. Knowledge elements extracted from massive amounts of text are fragmented and must undergo a process of knowledge fusion and unification to form a consistent and clean knowledge network. This stage mainly addresses two issues: First, entity alignment, which merges descriptions of the same real-world object from different data sources, eliminating information redundancy and forming unified and more information-rich knowledge nodes. Second, semantic disambiguation, which distinguishes between objects with the same name but different meanings based on context. For example, accurately determining whether the term "Spring and Autumn" refers to a "historical period" or the Confucian classic "Spring and Autumn Annals" in a specific context.
[0020] After completing knowledge element extraction, semantic relationship mining, entity alignment, and semantic disambiguation, a knowledge point relationship set is generated for the obtained knowledge points and the relationships between them. Based on the knowledge point relationship set, the structured data produced through the above-mentioned processing is serialized into a standard format recognizable by the graph database. Each knowledge point is assigned a globally unique identifier as a node, and each relationship is defined as an edge, with the source node, relationship type, and target node of each edge explicitly defined. Subsequently, the IRT method is used to perform difficulty analysis on all obtained knowledge point nodes to obtain the inherent difficulty attributes of the knowledge points and associate them with the knowledge points.
[0021] Ultimately, this vast, interconnected semantic network was systematically imported into a graph database, completing the final transformation from unstructured text to a structured knowledge graph of the humanities domain. Thus, a domain knowledge foundation rich in semantic relationships and pedagogical attributes was constructed. This transformed the "Battle of Red Cliffs" from an isolated term into a knowledge hub deeply connected to countless nodes such as "the Three Kingdoms period," "Su Shi's 'Nian Nu Jiao: Reminiscences of Red Cliffs,'" and "ancient warfare tactics," providing a solid and intelligent knowledge infrastructure for subsequent personalized learning.
[0022] S2: Obtain student interaction behavior, analyze student interaction behavior through multimodal technology, extract student interest tags, calculate initial interest level, and construct dynamic interest tag set; Furthermore, step S2 also includes: Collect students' multimodal interaction data, construct a multimodal dataset, preprocess the multimodal dataset to obtain a structured feature set, the multimodal dataset includes unstructured text dialogues, continuous speech streams, discrete image snapshots, and implicit behavior logs within the system; Based on the structured feature set, cross-modal semantic alignment is performed using a domain-adaptive fine-tuned large language model to generate a semantic representation of student intent. Based on the semantic representation of student intent, the large language model with domain-adaptive fine-tuning is used to perform interest tag recognition and initial interest degree calculation, and a dynamic interest tag set is constructed.
[0023] Specifically, the system continuously collects multimodal, asynchronous, time-series data generated by students during their interactions with the learning companion system. This includes unstructured text dialogues, continuous audio streams, discrete image snapshots, and implicit behavioral logs within the system. All data is integrated, assigned high-precision timestamps, and buffered and synchronized using a message queue, ultimately constructing a unified, time-series multimodal dataset. This dataset represents raw observational data characterizing students' interactive behaviors, cognitive states, emotional tendencies, and interest preferences.
[0024] Subsequently, standardized multimodal data preprocessing operations are performed on the multimodal dataset to transform the raw data into machine-readable, structured data. For integrated continuous speech stream data, noise reduction, pre-emphasis, and frame-segmentation windowing operations are first performed. Then, an automatic speech recognition engine based on an end-to-end deep learning architecture transcribes the temporal audio signal into a symbolic text sequence, and simultaneously extracts paralinguistic features such as fundamental frequency, energy, and speech rate from the audio as auxiliary parameters for sentiment intensity analysis. For discrete image snapshot data, contrast normalization and size normalization operations are performed, followed by calling an object detection and scene classification model based on a convolutional neural network pre-trained to parse the pixel matrix into structured semantic descriptors. For unstructured text dialogue data and implicit behavior log data, data cleaning, removal of irrelevant symbols, and serialization labeling are performed to form a standardized input vector. Finally, the preprocessed data is integrated and stored in a structured feature set. The structured feature set includes two types of data: one is a symbolic text sequence obtained from speech, image conversion and text input, and the other is a numerical feature vector composed of speech paralinguistic features, image embedding vectors, text embedding vectors and behavior log vectors.
[0025] Based on a domain-adaptive fine-tuned large language model, cross-modal semantic alignment and deep interest intent parsing are performed on data in the structured feature set. The goal of this step is to deeply parse the multimodal data stored in the structured feature set into interest tags and initial interest scores.
[0026] First, a central semantic fusion and alignment process is performed. At this stage, a domain-adaptive, fine-tuned large language model acts as the central semantic fusion unit, receiving all data, including symbolic text sequences and numerical feature vectors from the structured feature set. The large language model's data reception is not a simple information splicing, but rather a deep cross-modal semantic alignment and representation learning process. Through its internal cross-attention mechanism, the model weightedly fuses and cross-verifies information from different modalities, thereby generating a unified and dense semantic representation of student intent. It is important to note that this semantic representation of student intent is not a natural language paragraph, but a high-dimensional, dense numerical vector that encapsulates rich, fused semantic information, including core topics, specific entities, intentional actions, emotional tone, and cognitive depth.
[0027] Building upon this foundation, the large language model performs two key tasks in parallel: interest tag identification and initial interest degree calculation. Interest tag identification is handled by a multi-label classification head. Based on previously obtained student intent semantic representations, it guides the extraction of refined interest tags from text sequences within a structured feature set. The model first performs explicit entity extraction. Through a conditional random field layer or pointer network, it performs sequence labeling, examining the input text sequence word by word and assigning a BIO tag to each word. This ultimately pinpoints explicitly mentioned entities in the text, obtaining explicit interest tags. For example, when processing "Why did Shang Yang's reforms ultimately succeed?", explicit interest tags such as "Shang Yang" and "Shang Yang's reforms" are extracted. After explicit entity extraction, the model begins implicit interest inference based on cue engineering. This process uses student intent semantic representations incorporating multimodal information as input. The large language model performs step-by-step inference based on built-in cue words, analyzing the semantic roles of the text to obtain inference tags, including "who, to whom, what did, under what circumstances, and why". Next, the large language model maps the inference tags to the built-in humanities knowledge graph, ultimately using keywords from the humanities knowledge graph as the final implicit interest tags. Taking the text "Why did Shang Yang's reforms ultimately succeed, but Shang Yang was persecuted?" as an example, the large language model analyzes the semantic structure of the sentence using prompts, identifying the subject, predicate, and object, and inferring the actions and relationships within. It outputs inference tags including "Core Action: Investigating the Reasons," "Focus Subject: Shang Yang's Reforms," "Focus of Contradiction: The Success of the Reforms and the Personal Tragedy," and "Implicit Relationship: Shang Yang's Reforms - Possessing Attributes - Success." These inference tags are then mapped to higher-level concepts or topics in the knowledge graph. For example, "personal tragedy" is associated with "evaluation of historical figures," and "the contradiction between success and personal fate" is associated with "historical background." Finally, the large language model outputs implicit interest tags for "evaluation of historical figures" and "historical background." Through this multi-round reasoning, the large language model achieves context awareness and infers implicit interest orientations and cognitive needs.
[0028] The initial interest score is calculated by a multi-output regression head that works in parallel with the classification head. Each output channel is strictly aligned with an interest label, meaning that the calculation of the initial interest score and the label recognition are two synchronous and independent pathways. The regression head's calculation does not depend on the classification head's judgment; the two are then paired. Each channel is an independent regressor that specifically learns how to extract and weight the most relevant multidimensional signals to its corresponding label from a unified semantic representation of student intent. These signals include the semantic clarity, sentiment polarity, interaction initiative, and interaction depth of the interest label. Finally, each channel outputs an independent initial interest score for its corresponding label, which is normalized to the [0,1] interval.
[0029] Finally, the outputs of the classification and regression heads are aligned and paired to generate an <interest label, initial interest level> pair for each identified interest label. This result, along with the current timestamp and decay factor, is stored in a tuple. The decay factor is a customizable parameter used to simulate the human memory and interest decay patterns described by the Ebbinghaus forgetting curve. The resulting series of tuples constitutes the direct output of this interaction, quantifying the student's knowledge preferences and exploration tendencies at any given time, providing crucial and computable cognitive state input for subsequent processing. All tuples are integrated to generate a dynamic set of interest labels. This dataset is not a static set of labels, but rather a set of interest state variables that decay over time and can be reactivated.
[0030] S3: Construct an interest update model, combine dynamic interest tag set to calculate the interest of relevant knowledge points, and obtain a comprehensive interest set; Furthermore, step S3 also includes: Traverse the dynamic interest tag set, and for interest tags that have not been repeatedly activated, calculate their real-time interest level at the current moment according to the time decay model. The specific calculation formula is as follows: ; Where i represents the interest tag, and t represents the current time. This represents the real-time interest level of interest tag i at the current moment. This represents the interest level of the interest tag at the time of the last record. If there is no previous record, this value is the initial interest level in the tuple. e is the exponential function, λ represents the decay coefficient, and Δt is the precise interval between the current time and the last intensity update timestamp. When the interest tag is identified again, the existing interest level is enhanced and updated by combining the initial interest level and semantic relevance of this interaction. The specific calculation formula is as follows: ; in, This represents the real-time interest level after the interest tag i is dynamically updated. This represents the current interest level of the interest tag after time decay correction. A is the initial interest level given to the interest tag by the large language model when parsing this interaction, while S is a key semantic relevance coefficient. Represents a nonlinear adjustment factor. Represents the inhibition factor based on the current level of interest, where It controls based on the current level of interest. The hyperparameter of compressive strength; Integrate all updated interest tags and their interest scores to construct a comprehensive set of interest scores.
[0031] Specifically, after obtaining a dynamic set of interest tags that includes a time dimension and an initial level of interest, the system begins to calculate the student's real-time interest level for each interest tag. It's important to clarify that the real-time interest level is a dynamic variable obtained based on the initial interest level and a decay factor. It reflects the student's evolving interest state over time; its intensity naturally decays over time, but it can also be activated and enhanced by new relevant interactions.
[0032] In the dynamic interest tag set, each interest tag is associated with an independent state record tuple, which contains the interest tag, initial interest level, current timestamp, and decay factor. When real-time interest level calculation is required, the system first iterates through all records in the dynamic interest tag set and performs real-time correction on the current interest level of each interest tag based on time decay. Its mathematical core is an exponential decay model, which simulates the natural decay of memory and interest in human cognition. By obtaining the current precise time, the system calculates the time interval since the last update for the interest tag, and substitutes this interval along with the decay factor parameter personalized for the student into the exponential decay function to calculate the real-time interest level of the interest tag at this very moment. The specific calculation formula is as follows: ; Where i represents the interest tag, and t represents the current time. This represents the real-time interest level of interest tag i at the current moment. This represents the interest level of the interest tag at the time of the last record. If there is no previous record, this value is the initial interest level in the tuple. e is an exponential function, λ represents the decay coefficient, which is an adjustable parameter greater than zero that controls the rate at which the interest intensity decays over time. Its value can be personalized according to the student's overall learning habits. Δt is the precise interval between the current time and the timestamp of the last intensity update. This mathematical model simulates the psychological process of the natural decay of interest intensity from a computational perspective, ensuring the timeliness of the interest model.
[0033] The final calculated real-time interest score will overwrite the initial interest score obtained from the tuple, thus achieving dynamic updates of the interest score of the interest tag.
[0034] The above calculation describes how the interest level of an interest tag naturally decays over time when it is retrieved only once. However, in actual interactions, an interest tag may be retrieved multiple times, indicating that the student has mentioned the knowledge points related to that tag multiple times. Therefore, when this happens, the interest level of the tag no longer follows the natural decay pattern but is instead activated and enhanced multiple times.
[0035] Specifically, when the system parses an interest tag from a student's new interaction, and this tag already exists in the dynamic interest tag set, the system first performs identification and matching, ensuring the newly parsed interest tag corresponds perfectly with an existing interest tag in the dynamic interest tag set. Next, it performs an update; the system does not create a new interest tag but instead strengthens the interest level of the existing interest tag in the model. The specific update formula is as follows: ; in, This represents the real-time interest level after the interest tag i is dynamically updated. The current interest level of the interest tag after time decay correction is represented by A, which is the initial interest level given by the large language model when parsing this interaction. It quantifies the intensity of interest reflected in this interaction itself. S is a key semantic relevance coefficient, which characterizes the degree of semantic association between the specific content of this interaction and the interest tag i. This value is obtained by calculating the cosine similarity between the two in the vector space. The γ value represents a non-linear modulating factor, typically ranging from 0.5 to 2.0, and is set to 1.2 here. This factor's effect is that when γ > 1, it means only semantically highly relevant interactions (S close to 1) can significantly strengthen interest, while weakly relevant interactions (S close to 0) have a drastically reduced impact. This accurately reflects reality: a deep, relevant discussion is far more effective in consolidating and enhancing interest than a general mention. This represents a suppression factor based on the current level of interest, used to achieve a gentle update for high values and a relatively amplified update for low values. It controls based on the current level of interest. The hyperparameter of compressive strength is typically set between 2 and 10; here it is set to 5. The significance of this parameter is that... The larger the value, the earlier the inhibition takes effect; that is, the smaller the value... It will then enter the inhibition zone. The smaller it is, the larger it can be. The overall logic of the formula is that when the current interest level of an interest tag is low, a relatively large increase is allowed; when the current interest level of a tag is already high, the increase is suppressed to prevent a surge. This update mechanism ensures that the student's interest model can respond sensitively and reasonably to their continuous and relevant exploratory behavior.
[0036] After calculating and updating the interest level of each interest tag, the system persistently associates and stores the interest tag with its corresponding real-time interest level. This associated data is then updated in the original state record tuple of the interest tag, replacing the old initial interest level, thus ensuring that the interest state maintained in the dynamic interest tag set always reflects the student's latest cognitive tendencies. Subsequently, the system iterates through all updated interest tags and their corresponding real-time interest levels within the dynamic interest tag set, systematically integrating these discrete and quantified interest state data points to construct a unified and structured comprehensive interest level set. This set, as the output of the entire interest level update model, not only fully represents the distribution of students' interest intensity across various dimensions within the knowledge domain at the current moment, but also provides accurate and computable data input for the next stage of calculating the knowledge attraction of specific nodes in the knowledge graph. It serves as a key conversion hub for realizing the transformation from student interests to knowledge delivery.
[0037] S4: Traverse the knowledge graph of the humanities field, calculate the knowledge gravity value of each knowledge point by combining the interest degree comprehensive set, integrate all knowledge gravity values, and construct a knowledge gravity set; Furthermore, step S4 also includes: Traverse the knowledge point nodes in the knowledge graph of the humanities domain and calculate the semantic association strength between them and each interest tag in the interest degree comprehensive set; For each knowledge point node, the interest level of all related interest tags is aggregated, and difficulty adjustment is introduced to calculate the knowledge attraction value of that node for students. The specific calculation formula is as follows: ; in, The summation symbol Σ represents the knowledge attraction value of knowledge node j to students, and the summation symbol Σ indicates that all interest tags i associated with node j are traversed. This represents the current level of interest in the interest tag i. α represents the semantic association strength between interest tag i and knowledge point node j, and α represents the association sensitivity index. j β represents the inherent difficulty attribute of knowledge point node j, and β represents the difficulty adjustment factor. Summarize all knowledge point nodes and their knowledge gravity values to generate a knowledge gravity set.
[0038] Specifically, after completing the real-time interest score updates for all interest tags, the knowledge gravity value calculation for knowledge points begins. The goal of this step is to map interest tags to knowledge point nodes in the knowledge graph space and calculate the aggregated interest score for each knowledge point node.
[0039] By traversing every knowledge point node in the humanities knowledge graph, a many-to-many semantic relevance calculation process is initiated for the currently computed node. This process aims to identify all interest tags semantically related to the given knowledge point node within the comprehensive interest set, and to precisely quantify the semantic relevance between them. The semantic relevance calculation is not based on simple string matching, but rather relies on the rich semantic relationships and vectorized representations pre-stored in the humanities knowledge graph. It comprehensively determines the semantic relevance by comparing the cosine similarity of the semantic vectors of interest tags and knowledge point nodes in the shared vector space, thereby ensuring that the deep semantic connections between interest tags and knowledge point nodes can be captured.
[0040] After obtaining the semantic relevance between all knowledge point nodes and interest tags, a weighted aggregation operation is performed on the current knowledge point node. The specific calculation formula is as follows: ; in, The summation symbol Σ represents the knowledge attraction value of knowledge node j to students, and the summation symbol Σ indicates that all interest tags i associated with node j are traversed. This represents the current level of interest in interest tag i. This level of interest may be the initial level of interest that has not been reactivated or updated, or it may be the real-time level of interest that has been updated subsequently. α represents the semantic association strength between interest tag i and knowledge point node j, and α represents the association sensitivity index. This parameter controls the influence pattern of association strength on attraction. When α>1, the system has an extreme preference for strong associations and will prioritize recommending content that highly matches the core interest. When α<1, the system maintains a certain degree of openness to weak associations, allowing for exploration and divergence of interests. Here, it is set to 0.8. j The inherent difficulty attribute of knowledge point node j is typically a scalar greater than 0. β represents the difficulty adjustment factor, an adjustable parameter within the range of [0.5, 3.0]. This factor controls the system to favor students' comfort zones during recommendations. When the β value is large, the system tends to recommend content that is interesting and relatively easy, keeping students within their comfort zone. When the β value is small, the system is more willing to recommend content that is interesting but challenging, promoting cognitive development. Here, it can be set to 1.5. Through this series of calculations, the system ultimately assigns a quantified, personalized, real-time knowledge gravity value to each knowledge point in the humanities knowledge graph.
[0041] Specifically, this calculation formula multiplies the interest level of each associated interest tag by its semantic relevance to the current knowledge point to obtain a contribution value. Then, the contribution values of all associated interest tags are summed. To incorporate the control of the teaching dimension, this summation result is usually divided by the inherent difficulty of the current knowledge point to the power of β.
[0042] Finally, the identifiers of all knowledge point nodes are paired with their calculated knowledge gravity values to form a large set of key-value pairs, generating a knowledge gravity set. This dataset quantitatively represents the relative attractiveness of all available knowledge points to students across the entire knowledge domain at a specific moment.
[0043] S5: Generate a push sequence based on the knowledge gravity set to push personalized content. The push sequence includes main knowledge points and exploratory knowledge points. Furthermore, step S5 also includes: Based on knowledge gravity sets, sorting and clustering analysis are performed to identify the top high-gravity knowledge point clusters and their core themes, forming structured knowledge modules; Introduce a balance mechanism between exploration and utilization to screen exploratory content that has high potential but has not yet reached the top; The system generates a sequence of push notifications that includes key knowledge points and exploratory knowledge points, and adapts the content format and complexity to students' historical preferences to achieve personalized content delivery.
[0044] Specifically, after generating the knowledge gravity set, content needs to be pushed based on the knowledge gravity value of each knowledge point node. This process is not simply about selecting the single knowledge point with the highest gravity value for push. Instead, a more sophisticated and strategic multi-objective optimization push process is executed.
[0045] First, the knowledge gravity set is traversed, and all knowledge point nodes are sorted in descending order and clustered. All knowledge point nodes are arranged from highest to lowest gravity value, forming a knowledge gravity spectrum. Next, a clustering algorithm is used to identify the top high-gravity clusters in the knowledge gravity spectrum; these clusters represent knowledge point groups with gravity values significantly higher than the average. Then, the distribution of knowledge point nodes within these top high-gravity clusters in the humanities knowledge graph is analyzed to identify their core themes or categories, such as "Tang Dynasty poetry" or "Song Dynasty Neo-Confucianism," ultimately forming structured knowledge modules to ensure that the pushed content organically unfolds around core areas of interest.
[0046] Next, to avoid the recommended content becoming trapped in an information cocoon, resulting in an overemphasis on students' existing interests and limiting their knowledge horizons, a small amount of high-potential exploratory content needs to be proactively introduced into the recommendation list. This selection must meet one of the following two conditions: First, while the knowledge attraction value of this knowledge point node has not yet entered the top high-attraction cluster, it is on a rapid upward trajectory; second, this knowledge point node has a close graph connection with the current high-attraction knowledge point nodes, but differs slightly in theme, possessing value for knowledge transfer.
[0047] Finally, an optimal push sequence is generated based on the above processing. This sequence includes key knowledge points and exploratory knowledge points. When pushing content, the system will call the content resource library associated with the knowledge point nodes and dynamically determine the presentation format and complexity of the content based on the student's historical preferences.
[0048] S6: Track the interaction results between students and the pushed sequence, adjust the parameters in the interest update model based on the students' feedback behavior, and optimize the interest calculation.
[0049] Specifically, the content push is not the end of the process, but rather the starting point for a new round of data collection and model optimization. After the content is pushed, the system captures students' interactive behavior again. However, this time, the multimodal interactions acquired are no longer the unstructured text dialogues, continuous voice streams, discrete image snapshots, and implicit behavior logs required when constructing interest tag sets. Instead, they include multi-dimensional interaction metrics encompassing behavioral indicators, sentiment indicators, and explicit feedback. Specifically, behavioral indicators include content completion rate, dwell time on the page, and number of rereads; sentiment indicators include the emotional state reflected in the voice or text data; and explicit feedback includes students' direct responses to the pushed content. The resulting multi-dimensional interaction metrics can track students' feedback behavior towards the pushed content.
[0050] Multi-dimensional interactive metrics are transmitted in real time to the adaptive engine to drive dynamic tuning of parameters in the formula. The core principle is to set the student's long-term learning gains as the optimization goal, using feedback data to calibrate key parameters in the interest update model and the knowledge gravity value calculation formula, making them more aligned with the student's individual cognitive patterns. Specific parameters optimized include the decay coefficient, nonlinear adjustment factor, difficulty sensitivity index, and correlation sensitivity index. The following explains how each parameter is adjusted.
[0051] Personalized calibration of the decay coefficient λ: If the system observes that a student consistently shows high engagement with content under a certain interest tag after multiple pushes, such as long dwell time and high completion rate, it indicates that the student's interest has strong persistence, and the system will appropriately lower the corresponding λ value to slow down the rate of interest decay. Conversely, if the interest fades rapidly, the λ value will be appropriately increased.
[0052] Contextual adaptation of the nonlinear adjustment factor γ: If the system finds that students only provide positive feedback when the pushed content is highly relevant to their interest tags, while the feedback is bland or even negative under weakly relevant content, the system will increase the value of γ to focus the model on core interests. Conversely, if students also show acceptance of weakly relevant but high-quality exploratory content, the system will decrease the value of γ to encourage bolder exploration of interests.
[0053] Adaptive adjustment of difficulty adjustment factor β: The system continuously monitors the difficulty D of the pushed content. j The relationship between beta and student frustration / affectiveness metrics. If students frequently provide correct feedback and demonstrate a sense of accomplishment on slightly more challenging content, it indicates their potential to tackle higher levels of difficulty, and the system will lower the beta value, gradually increasing the average difficulty of the recommended content. Conversely, if students frequently experience frustration with challenging content, the system will raise the beta value, providing more conservative, comfort-zone-friendly content.
[0054] Strategic optimization of the relevance sensitivity index α: The system optimizes α by analyzing students' feedback on different types of relevance (strong and weak relevance) content. If exploratory content with weak relevance generally receives good feedback, the α value is lowered to broaden the scope of interest-based relevance; if exploratory content receives poor feedback, the α value is increased to tighten the recommendation strategy and focus on deep relevance.
[0055] Through continuous, data-driven parameter fine-tuning, the system achieves a complete intelligent closed loop of perception-decision-execution-optimization. Each interaction, each push notification, and each feedback deepens the system's understanding of the student's cognitive model, leading to more accurate, efficient, and personalized learning support decisions. This is not merely a static algorithm, but a self-evolving learning support system that grows alongside students.
[0056] Example 2: Based on the same inventive concept as the multimodal adaptive learning companion method for humanities learning based on a large language model in the previous examples, this application also provides a multimodal adaptive learning companion system for humanities learning based on a large language model. Please refer to the appendix. Figure 2 The system includes: The knowledge graph construction module 11 is used to input primary and secondary school humanities texts into a large language model, perform deep semantic analysis and structured processing, and generate a humanities knowledge graph. Interest tag construction module 12 is used to acquire student interaction behavior, analyze student interaction behavior through multimodal technology, extract student interest tags, calculate initial interest degree, and construct dynamic interest tag set; Interest score comprehensive calculation module 13 is used to calculate the interest score of relevant knowledge points according to the interest score update formula and combined with dynamic interest tags to obtain a comprehensive set of interest scores. The knowledge gravity calculation and integration module 14 is used to traverse the knowledge graph of the humanities field, calculate the knowledge gravity value of each knowledge point in combination with the interest degree comprehensive set, integrate all knowledge gravity values, and construct a knowledge gravity set. Personalized content push sequence generation module 15 is used to generate a push sequence based on a knowledge gravity set for personalized content push. The push sequence includes main knowledge points and exploratory knowledge points. The parameter adaptive optimization module 16 is used to track the interaction results between students and the pushed sequence, and adjust the parameters in the interest update model according to the students' feedback behavior to optimize the interest calculation.
[0057] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0058] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multimodal adaptive learning companion method for humanities learning based on a large language model, characterized in that, The method includes: Input humanities texts from primary and secondary schools into a large language model, perform in-depth semantic analysis and structured processing, and generate a knowledge graph for the humanities field; Acquire student interaction behavior, analyze it using multimodal technology, extract student interest tags, calculate initial interest levels, and construct a dynamic interest tag set, including: Collect students' multimodal interaction data, construct a multimodal dataset, preprocess the multimodal dataset to obtain a structured feature set, the multimodal dataset includes unstructured text dialogues, continuous speech streams, discrete image snapshots, and implicit behavior logs within the system; Construct an interest update model, combine dynamic interest tag set to calculate the interest of relevant knowledge points, and obtain a comprehensive interest set; Traverse the knowledge graph of the humanities field, calculate the knowledge gravity value of each knowledge point by combining it with the comprehensive set of interest levels, integrate all knowledge gravity values, and construct a knowledge gravity set, which includes: Traverse the knowledge point nodes in the knowledge graph of the humanities domain and calculate the semantic association strength between them and each interest tag in the interest degree comprehensive set; For each knowledge point node, the interest level of all related interest tags is aggregated, and difficulty adjustment is introduced to calculate the knowledge attraction value of that node for students. The specific calculation formula is as follows: ; in, The summation symbol Σ represents the knowledge attraction value of knowledge node j to students, and the summation symbol Σ indicates that all interest tags i associated with knowledge node j are traversed. This represents the real-time interest level of interest tag i. α represents the semantic association strength between interest tag i and knowledge point node j, and α represents the association sensitivity index. j β represents the inherent difficulty attribute of knowledge point node j, and β represents the difficulty adjustment factor. Summarize all knowledge point nodes and their knowledge gravity values to generate a knowledge gravity set; A push sequence is generated based on the knowledge gravity set to deliver personalized content. The push sequence includes key knowledge points and exploratory knowledge points. By tracking the interaction results between students and the pushed sequences, and adjusting the parameters in the interest update model based on the students' feedback behavior, the interest calculation is optimized.
2. The multimodal adaptive learning companion method for humanities learning based on a large language model as described in claim 1, characterized in that, Generate knowledge graphs in the humanities field, including: Texts in the humanities field from primary and secondary schools are collected to construct an unstructured initial corpus. The initial corpus is then preprocessed to obtain a standardized corpus. The preprocessing includes format standardization and primary structure parsing. Select a large language model and perform domain-adaptive fine-tuning on it; The large language model, after domain-adaptive fine-tuning, is used to process the standardized corpus and generate a set of knowledge point relationships. The text processing includes knowledge element extraction, semantic relationship mining, entity alignment and semantic disambiguation. Generate a knowledge graph for the humanities field based on the set of relationships between knowledge points.
3. The multimodal adaptive learning companion method for humanities learning based on a large language model as described in claim 1, characterized in that, Constructing dynamic interest tag sets, including: Based on the structured feature set, cross-modal semantic alignment is performed using a domain-adaptive fine-tuned large language model to generate a semantic representation of student intent. Based on the semantic representation of student intent, the large language model with domain-adaptive fine-tuning is used to perform interest tag recognition and initial interest degree calculation, and a dynamic interest tag set is constructed.
4. The multimodal adaptive learning companion method for humanities learning based on a large language model as described in claim 1, characterized in that, Obtain a comprehensive set of interest levels, including: Traverse the dynamic interest tag set, and for interest tags that have not been repeatedly activated, calculate their real-time interest level at the current moment according to the time decay model. The specific calculation formula is as follows: ; Where i represents the interest tag, and t represents the current time. This represents the real-time interest level of interest tag i. This represents the interest level of the interest tag at the time of the last record. If there is no previous record, this value is the initial interest level in the tuple. e is the exponential function, λ represents the decay coefficient, and Δt is the precise interval between the current time and the last intensity update timestamp. When the interest tag is identified again, the existing interest level is enhanced and updated by combining the initial interest level and semantic relevance of this interaction. The specific calculation formula is as follows: ; in, This represents the real-time interest level after the interest tag i is dynamically updated. A represents the real-time interest score of interest tag i, A is the initial interest score given by the large language model when parsing this interaction, and S is a key semantic relevance coefficient. Represents a nonlinear adjustment factor. Represents the inhibition factor based on the current level of interest, where It controls based on the current level of interest. The hyperparameter of compressive strength; Integrate all updated interest tags and their interest scores to construct a comprehensive set of interest scores.
5. The multimodal adaptive learning companion method for humanities learning based on a large language model as described in claim 1, characterized in that, Personalized content delivery is achieved by generating push sequences based on knowledge gravity sets, including: Based on knowledge gravity sets, sorting and clustering analysis are performed to identify the top high-gravity knowledge point clusters and their core themes, forming structured knowledge modules; Introduce a balance mechanism between exploration and utilization to screen exploratory content that has high potential but has not yet reached the top; The system generates a sequence of push notifications that includes key knowledge points and exploratory knowledge points, and adapts the content format and complexity to students' historical preferences to achieve personalized content delivery.
6. A multimodal adaptive learning companion system for humanities learning based on a large language model, characterized in that: The system is used to implement the multimodal adaptive learning companion method for humanities learning based on a large language model as described in any one of claims 1 to 5, and the system comprises: The knowledge graph construction module is used to input humanities texts from primary and secondary schools into a large language model, perform deep semantic analysis and structured processing, and generate a humanities knowledge graph. An interest tag construction module is used to acquire student interaction behavior, analyze student interaction behavior through multimodal technology, extract student interest tags, calculate initial interest level, and construct dynamic interest tag set. An interest score calculation module is used to construct an interest score update model, combine a dynamic interest tag set to calculate the interest score of relevant knowledge points, and obtain an interest score comprehensive set. The knowledge gravity calculation and integration module is used to traverse the knowledge graph in the humanities field, calculate the knowledge gravity value of each knowledge point by combining the interest degree comprehensive set, integrate all knowledge gravity values, and construct a knowledge gravity set. A personalized content push sequence generation module is used to generate a push sequence based on a knowledge gravity set for personalized content push. The push sequence includes main knowledge points and exploratory knowledge points. The parameter adaptive optimization module is used to track the interaction results between students and the pushed sequence, and adjust the parameters in the interest update model according to the students' feedback behavior to optimize the interest calculation.