An intelligent analysis system for students' self-regulated learning strategies

By using an intelligent analysis system to collect and analyze classroom discussions in a full, real-time, and objective manner, and to generate intuitive and visual feedback, the system solves the problems of insufficient timeliness of data collection and insufficient depth of content analysis in existing technologies, thereby improving the data-driven capabilities of teaching and students' self-regulating learning abilities.

CN121503932BActive Publication Date: 2026-03-13TIANJIN NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-14
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies cannot effectively support the real-time and objective analysis of students' self-regulated learning strategies in classroom discussions, resulting in insufficient timeliness of data collection and depth of content analysis. They also lack text processing capabilities tailored to educational scenarios, have limited feedback formats, and fail to form a closed loop with teaching.

Method used

An intelligent analysis system for students' self-regulating learning strategies was designed, including a data acquisition and preprocessing layer, a text processing and enhancement layer, a core analysis layer, and an application display layer. Through audio acquisition, speech transcription, text cleaning, semantic normalization, and concept graph analysis, it generates intuitive visual feedback to help teachers conduct teaching interventions.

Benefits of technology

It enables the full, real-time, and objective collection and in-depth analysis of classroom discussion data, provides intuitive and interactive visual feedback, enhances the data-driven capability of teaching, and promotes the improvement of students' self-regulation learning ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121503932B_ABST
    Figure CN121503932B_ABST
Patent Text Reader

Abstract

This invention provides an intelligent analysis system for students' self-regulated learning strategies. The system includes a data acquisition and preprocessing layer, a text processing and enhancement layer, a core analysis layer, and an application display layer. The data acquisition and preprocessing layer includes an audio acquisition module and a speech-to-text and structuring module. The text processing and enhancement layer includes a text cleaning module and a semantic normalization module. The core analysis layer includes a concept graph analysis module and a learning strategy analysis module. The application display layer includes a visualization rendering module and a dashboard integration module. This invention achieves full, objective, and real-time collection of classroom discussion data, eliminating the subjectivity and lag of manual recording. Through deep semantic analysis, it reveals the knowledge structure of the discussion content, providing core technical support for empirical research on "cultivating students' self-regulated learning abilities based on intelligent learning analysis tools."
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, and in particular relates to an intelligent analysis system for students' self-regulating learning strategies. Background Technology

[0002] Currently, the focus of education has shifted from knowledge transmission to competency development, with particular emphasis on enhancing core competencies such as self-regulated learning. Zimmerman (1986) defines self-regulated learning as the process by which learners systematically activate and maintain their cognitive activities and behaviors in order to achieve their learning goals. From a social cognitive perspective, he proposed a self-regulated learning model comprising three cyclical stages: the pre-thinking stage (task analysis and self-motivational beliefs), the performance stage (self-control and self-observation), and the self-reflection stage (self-judgment and self-response). The pre-thinking stage occurs before actual learning behavior, where learners prepare for learning and set direction; the performance stage occurs during the learning behavior, where learners execute their plans and monitor their progress; and the self-reflection stage occurs after the learning behavior, where learners reflect on the learning process and its results. He emphasizes that learners actively set goals, utilize strategies, and adjust based on feedback throughout this process, making it a dynamic and cyclical regulatory process.

[0003] Pintrich (2000) defines self-regulated learning (or self-regulation) as: "an active, constructive process in which learners set goals for their learning and then, guided and constrained by the characteristics of their goals and the environment, attempt to monitor, regulate, and control their cognition, motivation, and behavior." He further expands the dimensions of self-regulation, proposing a comprehensive model that includes four regulatory stages (pre-thinking and planning activation, monitoring, control, response and reflection) and four regulatory domains (cognition, motivation, behavior, and context). His model emphasizes the role of motivational and affective regulation in the learning process.

[0004] Regarding learning strategies, both scholars pointed out that self-regulated learners actively employ various strategies to enhance learning outcomes. Zimmerman et al. summarized ten commonly used strategies through interviews, including self-evaluation, organization and transformation, institutional goals and planning, information seeking, record keeping and monitoring, environment construction, self-reward and punishment, repetition and memorization, seeking social help, and reviewing records. Pintrich, on the other hand, systematically divided learning strategies into three categories: cognitive strategies, metacognitive strategies, and resource management strategies, emphasizing students' ability to select, monitor, and regulate strategies. These strategies are not only concrete manifestations of self-regulated learning but also key to improving academic achievement and learning autonomy. Classroom discussions are a crucial scenario for cultivating students' self-regulated learning abilities, but traditional analytical techniques cannot effectively support this process.

[0005] Previously, while students' participation in classroom discussions or reflections contained rich information about self-regulation learning strategies, they lacked real-time, objective analytical tools to aid their self-awareness. Students struggled to systematically identify whether they actively used strategies such as planning, monitoring, reflection, and resource management during discussions, and were unable to understand the patterns and effects of their strategy usage. Due to the fast pace and fragmented content of classroom discussions, students found it difficult to immediately perceive and evaluate their strategy usage in dynamic exchanges. This resulted in self-regulation processes remaining largely subconscious, failing to translate into conscious, optimizable learning behaviors, thus limiting the effective development and improvement of their self-regulation learning abilities.

[0006] Currently, the analysis techniques for classroom discussion content mainly take the following forms:

[0007] Basic electronic questionnaires and voting tools: Some interactive platforms (such as Rain Classroom and Seewo Easy Classroom) allow students to vote on multiple-choice questions and send short text comments via their devices. This method can quickly collect scattered feedback, but the information is fragmented, lacks depth, and cannot reflect students' mastery of learning strategies. The text content is mostly pre-set answers or brief reflections, making it difficult to systematically present the core framework and knowledge focus of group discussions.

[0008] Common word cloud generation tools: Some teachers try using existing online word cloud tools (such as WordArt, MicroWordCloud, etc.). The usual practice is for teachers to manually input keywords from classroom discussions based on their recollections and notes after class, and then generate a word cloud. While this method can present visual results, it has drawbacks: the data source is subjective, incomplete, and not an objective record; the word cloud reflects more the teacher's "opinions" than the students' actual "expressions."

[0009] Based on the above situation, the following key shortcomings of existing technologies can be summarized:

[0010] 1. Insufficient timeliness and objectivity in data collection: It relies heavily on manual recording, which makes the data prone to omissions, subject to the subjective selection of the recorder, and cannot achieve real-time processing, resulting in delayed feedback.

[0011] 2. Insufficient depth of content analysis: Existing analytical methods only focus on superficial and fragmented language, while ignoring the value of the discussion content itself regarding self-regulation learning strategies, and cannot answer deeper questions such as "What learning strategies did the students actually discuss?" and "What are the core concepts related to self-regulation learning?"

[0012] 3. Lack of dedicated text processing capabilities for educational scenarios: General word cloud tools are not optimized for classroom speech characteristics (such as colloquialism, repetition, pauses, and alternating teacher-student dialogues) and subject-specific terminology (such as theories related to self-regulated learning). The generated results are crude, noisy, and have limited educational significance.

[0013] 4. The feedback format is too simplistic and fails to form a closed loop with teaching improvement: The analysis results are often presented in the form of reports, which are disconnected from the improvement of classroom teaching behavior and fail to provide teachers with intuitive and actionable insights to support their teaching reflection and adjustment.

[0014] In conclusion, existing technologies have significant shortcomings in supporting the cultivation of self-regulated learning abilities through discussion and analysis. Therefore, there is an urgent need for an intelligent tool that can deeply analyze discussion content, interpret the current state of students' self-regulated learning strategies, provide teachers with accurate learning information, and offer precise evidence for instructional interventions to meet the practical needs of quality education and precision teaching in the new era. Summary of the Invention

[0015] In view of this, the present invention aims to overcome the shortcomings of the prior art and propose an intelligent analysis system for students' self-regulated learning strategies. This system achieves full, objective, and real-time collection of classroom discussion data, eliminating the subjectivity and lag of manual recording. Through deep semantic analysis, it reveals the knowledge structure of the discussion content, providing core technical support for empirical research on "cultivating students' self-regulated learning ability based on intelligent learning analysis tools." By providing intuitive and interactive visual feedback, it helps teachers quickly understand the dynamics of the discussion, conduct precise teaching interventions, and promote the transformation of classroom teaching from "experience-driven" to "data-driven," ultimately achieving the goal of improving students' learning effectiveness and core competencies.

[0016] To achieve the above objectives, the technical solution of the present invention is implemented as follows:

[0017] An intelligent analysis system for students' self-regulating learning strategies, the system comprising a data acquisition and preprocessing layer, a text processing and enhancement layer, a core analysis layer, and an application display layer;

[0018] The data acquisition and preprocessing layer includes an audio acquisition module and a speech transcription and structuring module. The audio acquisition module is used for hardware control and raw acquisition of audio streams, and the speech transcription and structuring module is used to convert audio streams into text with speaker tags.

[0019] The text processing and enhancement layer includes a text cleaning module and a semantic normalization module. The text cleaning module is used to clean the text, remove noise, and retain valuable educational features. The semantic normalization module is used to improve the standardization and semantic clarity of the text.

[0020] The core analysis layer includes a concept graph analysis module and a learning strategy analysis module. The concept graph analysis module is used to construct the knowledge network of the current classroom discussion, and the learning strategy analysis module is used to analyze the discussion content from the perspective of teaching objectives.

[0021] The application presentation layer includes a visualization rendering module and a dashboard integration module. The visualization rendering module is used to transform data into intuitive graphics, and the dashboard integration module is used to uniformly manage and display all visualization components.

[0022] The concept mapping analysis module is implemented as follows:

[0023] The discussion text is segmented and filtered by part of speech, retaining only content words as candidate keywords to form a set of nodes in the graph. A sliding window of size 7 is defined, and the co-occurrence relationship between any two candidate words in the window is counted. If two words co-occur in the window, an edge is established between them, and the number of co-occurrences is used as the initial weight of the edge.

[0024] Introducing word frequency factors into the edge weight calculation gives higher weight to the connections between high-frequency core words, while assigning higher initial scores to words that appear in key positions.

[0025] The importance score of each word is obtained by iterative calculation. After multiple rounds of iteration until the score converges, all candidate words are sorted in descending order according to the final score, and the K words with the highest ranking are selected as the concept set that best represents the core of the discussion.

[0026] The trained concept entity recognition model is used to identify entities, dependency parsing is used to obtain dependency trees, and the identified concept entities are used as central nodes. The path-first search algorithm is used to find grammatical paths connecting them, the relationship between core concepts in the context is analyzed, and a classroom discussion concept graph is constructed with concepts as nodes and relationships as edges.

[0027] Furthermore, the audio acquisition module includes:

[0028] Device driver unit: controls the start / stop and parameter settings of the microphone array;

[0029] Streaming data receiving unit: continuously receives raw streaming data from the audio acquisition device and performs buffering, packetization, and timestamp marking;

[0030] The speech transcription and structuring module includes:

[0031] Automatic Speech Recognition Unit: Calls the trained and optimized ASR engine API to convert the audio stream into raw text with timestamps;

[0032] Speaker Separation and Recognition Unit: Analyzes the audio, distinguishes different speakers, and aligns and binds the identified speaker IDs with the text generated by ASR;

[0033] Text formatting unit: Generates the final result in a standard structured log format.

[0034] Furthermore, the text cleaning module includes:

[0035] Noise filtering unit: Based on preset rules and a vocabulary database, it filters out verbal tics and repetitive words;

[0036] Educational Feature Identification and Retention Unit: Identifies and ensures that key features are not removed through keyword matching or simple rules;

[0037] The semantic normalization module includes:

[0038] Thesaurus Management Unit: Maintains a thesaurus for the education field, providing query and merging services;

[0039] Concept Entity Recognition Unit: Identifies key entities in the text and assigns type labels to these key entities.

[0040] Furthermore, the concept mapping analysis module includes:

[0041] Core concept extraction unit: Extracts core keywords from the text as concept nodes;

[0042] Relation extraction unit: Utilizes a pre-trained model to analyze the grammatical relations of concepts in a sentence and extract the semantic relations between concepts;

[0043] Graph Construction and Storage Unit: Constructs nodes and relationships into a graph data structure for storage;

[0044] The learning strategy analysis module is based on a self-regulating learning model. It identifies self-regulating learning strategies that students directly indicate or imply in their language. It scans cleaned text, matches keywords and contextual expressions related to the learning strategies, and statistically analyzes the frequency, distribution patterns, and contextual relationships of different strategy words. This reveals the specific self-regulating learning strategies that students use in classroom discussions.

[0045] Furthermore, the visualization rendering module includes:

[0046] Intelligent word cloud generation unit: Receives data from the concept graph analysis module and the learning strategy analysis module, and generates a multi-dimensional word cloud;

[0047] Timeline rendering unit: Based on the initial structured text log, it generates an interactive timeline, which can be clicked to view the discussion content at a specific moment;

[0048] The instrument panel integration module includes:

[0049] Layout Management Unit: Responsible for arranging visual components on the webpage;

[0050] Data Interface Unit: Retrieves data from the core analytics layer and database, and distributes it to various visualization components.

[0051] Furthermore, the text cleaning module employs a BERT fine-tuning method based on sequence labeling to filter text noise and preserve educational features, specifically including:

[0052] Based on the BERT text model, we fine-tuned the noise filtering as a three-class classification task: O - Keep / B - DELETE - Delete the beginning / I - DELETE - Delete the middle.

[0053] Using the Focal Loss loss function, a weight factor α is introduced, and α is preset for each class k. k Weight;

[0054] By using rule-based annotation strategies, a clear "whitelist" of terms can be established to avoid misjudgments.

[0055] Furthermore, the semantic normalization module is implemented as follows:

[0056] A multi-task learning framework is constructed. For synonym merging, a fine-tuned BERT model is used to classify sentence pairs, and manual annotation is used to determine whether the expressions are synonymous. A thesaurus is built to merge different words that express the same concept into standard terms.

[0057] Based on the traditional BIO annotation method, new labels for educational entities are added, and the model is fine-tuned into a sequence labeling system, enabling it to identify and label lexical segments in the text that represent educational entities.

[0058] Furthermore, the learning strategy analysis module is implemented as follows:

[0059] A classification system for learning strategies was constructed, and a corresponding library of keywords and expression patterns was established for each type of strategy.

[0060] Based on the concept graph obtained by the concept graph analysis module, the identified text fragments are classified according to the learning strategy classification system and projected onto the nodes of the concept graph. The word frequency, distribution pattern and contextual relationship of different strategy words are statistically analyzed, thereby showing the specific self-regulating learning strategies adopted by students in classroom discussions.

[0061] Furthermore, the system operates as follows:

[0062] The audio acquisition module was used to collect multimodal audio data of student discussions in class.

[0063] The collected audio data is converted into speech and processed into structure using the speech-to-text and structuring modules to obtain the processed text data.

[0064] The text cleaning module is used to filter noise and retain educational features from the obtained text data, resulting in cleaned text data.

[0065] The semantic normalization module is used to perform semantic normalization and concept annotation on the cleaned text data;

[0066] The concept graph analysis module is used to construct a concept graph of classroom discussions with concepts as nodes and relationships as edges, so as to deeply interpret classroom discussions from multiple dimensions.

[0067] The learning strategy analysis module is used to perform semantic analysis on the identified text in order to identify the self-regulated learning strategies demonstrated by students in the discussion.

[0068] Use the visualization rendering module to generate intuitive and visual word clouds;

[0069] The dashboard integration module is used to generate a dashboard that displays a smart word cloud and a timeline of the discussion process.

[0070] Compared with existing technologies, the intelligent analysis system for student self-regulation learning strategies described in this invention has the following advantages:

[0071] The innovative method of this invention lowers the threshold for data interpretation, transforms data into action guidelines, and provides clear and concise advice, greatly improving the ease of use and accessibility of the technology.

[0072] This invention upgrades the analysis system from a "diagnostic tool" to a "decision support system," enabling it to provide suggested solutions and drive changes in teaching practices.

[0073] The interactive and interconnected visualization dashboard of this invention showcases an innovative output format. On the one hand, it provides exploratory analysis capabilities, transforming "passive viewing" into "active exploration." On the other hand, it fosters a holistic understanding, integrating fragmented analytical dimensions (content, interaction, process) into a single interface. This helps teachers quickly build a comprehensive and systematic understanding of the entire lesson's discussion, avoiding bias and ensuring that the system's final output is intuitive, operable, and capable of driving teaching improvement.

[0074] The innovative technical approach of this invention ensures that the system processing is efficient, automated, and accurate. Attached Figure Description

[0075] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0076] Figure 1 This is a schematic diagram of the structure of an intelligent analysis system for student self-regulation learning strategies according to the present invention;

[0077] Figure 2 This is a schematic diagram of the working process of the system of the present invention;

[0078] Figure 3 This is an example of the original text used in the data collection process of this invention;

[0079] Figure 4 This is an example of an intermediate result of data cleaning in this invention;

[0080] Figure 5 This is an example of the final cleaned and enhanced text of the present invention. Detailed Implementation

[0081] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0082] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.

[0083] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0084] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0085] like Figure 1 As shown, the present invention provides an intelligent analysis system for students' self-regulation of learning strategies. The system includes a data acquisition and preprocessing layer, a text processing and enhancement layer, a core analysis layer, and an application display layer.

[0086] The data acquisition and preprocessing layer includes an audio acquisition module and a speech transcription and structuring module. The audio acquisition module is used for hardware control and raw acquisition of audio streams, and the speech transcription and structuring module is used to convert audio streams into text with speaker tags.

[0087] The text processing and enhancement layer includes a text cleaning module and a semantic normalization module. The text cleaning module is used to clean the text, remove noise, and retain valuable educational features. The semantic normalization module is used to improve the standardization and semantic clarity of the text.

[0088] The core analysis layer includes a concept graph analysis module and a learning strategy analysis module. The concept graph analysis module is used to construct the knowledge network of the current classroom discussion, and the learning strategy analysis module is used to analyze the discussion content from the perspective of teaching objectives.

[0089] The application presentation layer includes a visualization rendering module and a dashboard integration module. The visualization rendering module is used to transform data into intuitive graphics, and the dashboard integration module is used to uniformly manage and display all visualization components.

[0090] In this embodiment, the audio acquisition module is responsible for hardware control and raw audio stream acquisition, specifically including:

[0091] Device driver unit: controls the start / stop and parameter settings of the microphone array to ensure reliable startup and optimize recording quality during specific discussion sessions of the self-regulating learning ability intervention experiment.

[0092] Streaming data receiving unit: continuously receives raw streaming data from audio acquisition devices, and performs buffering, packetization, and timestamp marking to prepare for subsequent real-time speech transcription.

[0093] In this embodiment, the speech transcription and structuring module intelligently converts the audio stream into text with speaker tags, specifically including:

[0094] Automatic Speech Recognition Unit: Calls the trained and optimized ASR engine API to convert the audio stream into raw text with timestamps.

[0095] Speaker Separation and Recognition Unit: Analyzes the audio, distinguishes different speakers, and aligns and binds the identified speaker IDs with the text generated by ASR.

[0096] Text formatting unit: Generates the final result in a standard structured log format [timestamp] [speaker ID]: [conversation content].

[0097] In this embodiment, the text cleaning module is used to clean the text, remove noise, and retain valuable educational features. It employs a parameter-sharing architecture—a shared BERT encoder is used to construct different units that provide services for multiple tasks such as noise filtering, concept recognition, and synonym detection. Specifically, this includes:

[0098] Noise filtering unit: Based on preset rules and a vocabulary database, it quickly filters out verbal tics, repetitive words, etc.

[0099] Educational Feature Identification and Retention Unit: Through keyword matching or simple rules, identify and ensure that key features such as logical words and interrogative words are not removed.

[0100] In this embodiment, the semantic normalization module is used to improve the standardization and semantic clarity of the text, specifically including:

[0101] Thesaurus Management Unit: Maintains a thesaurus for the education field, providing query and merging services.

[0102] Concept Entity Recognition Unit: Uses NLP techniques (such as named entity recognition) to identify key entities in the text (such as "Pythagorean theorem" and "Xinhai Revolution") and assigns them type labels (such as "[historical event]").

[0103] In this embodiment, the concept mapping analysis module is used to construct the knowledge network for this classroom discussion, specifically including...

[0104] Core concept extraction unit: Using graph algorithms or TF-IDF-based variants, core keywords are extracted from the text as concept nodes.

[0105] Relation extraction unit: Using a pre-trained model, analyze the grammatical relations of concepts in a sentence (such as subject, verb, and object) and extract the semantic relations between concepts (such as "belongs to" or "leads to").

[0106] Graph construction and storage unit: Constructs nodes and relationships into a graph data structure and stores it in a graph database or memory for querying and visualization.

[0107] In this embodiment, the learning strategy analysis module is used to analyze the discussion content from the perspective of teaching objectives, specifically including:

[0108] The learning strategy analysis unit, based on a self-regulated learning model (comprising three stages: planning, execution, and self-reflection), aims to identify self-regulated learning strategies that students directly indicate or imply in their language. This involves scanning cleaned texts, matching keywords and contextual expressions related to these strategies, and statistically analyzing the frequency, distribution patterns, and contextual relationships of vocabulary related to different strategies. This objectively and quantitatively reveals the specific self-regulated learning strategies students employ in classroom discussions.

[0109] In this embodiment, the visualization rendering module is used to transform data into intuitive graphics, specifically including:

[0110] Intelligent word cloud generation unit: It receives "concept coreness" data from the concept graph analysis module to determine the word size; and can also receive "learning strategy" data from the teaching analysis module to determine the word color, generating a multi-dimensional word cloud.

[0111] Timeline rendering unit: Based on the initial structured text log, it generates an interactive timeline, which can be clicked to view the discussion content at a specific moment.

[0112] In this embodiment, the dashboard integration module serves as the system's user interface, uniformly managing and displaying all visual components, specifically including...

[0113] Layout Management Unit: Responsible for arranging components such as word clouds, timelines, and indicator cards on the webpage.

[0114] Data Interface Unit: Retrieves data from the core analytics layer and database, and distributes it to various visualization components.

[0115] The working process of the intelligent analysis system of the present invention is as follows: Figure 2 As shown, specifically including

[0116] Step 1: Acquire multimodal audio data;

[0117] When students engage in reflective strategy discussions in class, microphone arrays deployed in the classroom or student terminal devices are used to record the entire process to ensure that clear and complete audio data is collected, providing raw materials for subsequent analysis.

[0118] Step 2: Perform speech-to-text and structured processing on the collected audio data;

[0119] The acquired audio stream is transmitted to the ASR engine in real time. This engine, optimized for educational scenarios, possesses the ability to resist noise, recognize colloquial expressions and subject-specific terminology, and convert speech into timestamped raw text.

[0120] The transcribed text is processed using traditional Mel-frequency cepstral coefficients (MFCCs) as acoustic features to automatically distinguish and label the speaker's identity for each segment of text (e.g., "Student 1", "Student 2", etc.). The final result is a structured text log in the format: [Timestamp] [Speaker ID]: [Conversation Content].

[0121] Step 3: Noise filtering and educational feature preservation of the text data;

[0122] This invention employs a BERT fine-tuning scheme based on sequence labeling to achieve intelligent filtering of text noise and accurate preservation of educational features. Specifically, this invention uses the BERT text model as a foundation for fine-tuning. By using a model architecture specifically designed for classifying each single character in the text (BERTForTokenClassification) based on the BERT architecture, noise filtering is created as a three-classification task (O - retain / B - DELETE - delete the beginning / I - DELETE - delete the middle).

[0123] Meanwhile, in the noise filtering task, the vast majority of tokens need to be retained (class O), and only a very small number of tokens are noise that needs to be deleted (classes B-DELETE, I-DELETE). This causes the loss function to be dominated by a large number of "simple" negative samples (class O), resulting in extremely low recognition rate (recall) for the minority classes. This invention uses the Focal Loss loss function, introducing a weight factor α, and pre-setting α for each class k. k Weights. The resulting loss function is given by formula (1), where α k p represents the weight of category k. k Let γ be the probability that the model predicts a class k, FL be the focusing parameter, and γ be the loss function.

[0124] (1)

[0125] This minimalist labeling design significantly reduces model complexity. Combined with mixed precision training (FP16), it increases training speed by 40%–60%. While removing meaningless filler words such as “um” and “ah”, it establishes a clear “whitelist” vocabulary through rule-based labeling strategies. Words in this vocabulary will not be filtered under any circumstances and can be enriched and improved through continuous application. This ensures that logical words such as “because” and “therefore” and interrogative words such as “what” and “how” are retained even if they are misjudged as noise by the model, achieving a balance between purification and fidelity.

[0126] Step 4: Perform semantic normalization and concept labeling on the purified text data;

[0127] A multi-task learning framework is constructed. For synonym merging, a locally fine-tuned BERT model is used for sentence pair classification, supplemented by manual annotation to determine whether expressions are synonymous, thereby building a thesaurus. Different words expressing the same concept are merged into standard terms (e.g., unifying "PPT" and "slides" into "slides").

[0128] By adding new labels for educational entities to the traditional BIO annotation method, the model is fine-tuned into a sequence labeling system, which can accurately identify and label the lexical segments representing educational entities in the text. For example, "Li Bai" in "Li Bai wrote "Quiet Night Thoughts" is labeled as B-PERSON, and "Quiet Night Thoughts" is labeled as B-WORKS.

[0129] Step 5: Construct a concept map of classroom discussions with concepts as nodes and relationships as edges, and deeply interpret classroom discussions from multiple dimensions;

[0130] A keyword extraction algorithm based on a graph ranking model is adopted, which models text units and their interrelationships as a graph structure, and identifies key concepts by iteratively calculating the importance of nodes in the graph.

[0131] First, the discussion text is segmented and filtered by part of speech, retaining only content words such as nouns and verbs as candidate keywords, forming the node set of the graph. Then, a sliding window of size 7 is defined, and the co-occurrence relationship between any two candidate words within the window is counted. If two words co-occur within the window, an edge is established between them, with the co-occurrence count used as the initial weight of the edge. This graph construction method effectively captures the semantic relationship network in the text.

[0132] Building upon the basic graph structure, word frequency and positional information are further incorporated to enhance model performance. By introducing word frequency into the edge weight calculation, connections between high-frequency core words receive higher weight. Simultaneously, based on the linguistic principle that "important concepts often appear in specific positions," words appearing at the beginning of paragraphs and other key positions are assigned higher initial scores, thus providing them with an importance bias from the outset of the algorithm.

[0133] The core calculation process employs the TextRank algorithm framework, using a damping coefficient to control the probability of random jumps. The final importance score for each word is calculated iteratively, and its value depends on the importance of all its connected neighboring words, as well as the weight ratio of the connecting edges. This process can be understood as an importance vote in a word association network; the more important neighboring words a word is connected to, the higher its importance.

[0134] The core formula for TextRank is shown in (2), where the calculation result is WS(V i ) is the word node V iThe TextRank importance score; d is the damping coefficient representing the probability of continuing the walk along the edge. As a random jump probability, it ensures algorithm convergence; In(V i W represents the set of all points to that node in the constructed graph; ji Represents word node V j To word node V i The edge weights reflect the strength of the association between two words in the text; Out(V j ) represents the word node V in the constructed graph. j The set of all external nodes pointed to; WS(V j ) is the word node V j Importance score Represents a set Any out-neighbor node in (i.e. The (one external node); Represents a node Its out-neighboring nodes The edge weights between words are used to measure the strength of the association between them. For word node V j The sum of the weights of all outer edges is used to... Normalization makes Indicates from node To the node The normalized proportion of propagation weight (transition probability).

[0135] (2)

[0136] The improved weighting formula is shown in (3), W ji Represents word node V j With words and nodes V i Edge weights between; count co (i,j) represents the number of times word i and word j co-occur within a specified sliding window; freq(i) and freq(j) represent the word frequencies of word i and word j in the entire document, respectively. By introducing a word frequency factor into the edge weights, connections between high-frequency core words can be given higher weights, thereby strengthening the influence of core concepts in the graph structure.

[0137] (3)

[0138] After multiple iterations until the scores converge, all candidate words are sorted in descending order according to their final scores, and the K highest-ranking words are selected as the set of concepts that best represent the core of the discussion. Manual labeling and calibration are used between iterations. The entire algorithm captures semantic associations through co-occurrence relationships, strengthens core concepts through word frequency, and introduces linguistic priors through location information, ultimately achieving accurate extraction of the core concepts of the discussion.

[0139] Using the concept entity recognition model trained in the previous stage, entities such as B-CONCEPT and B-WORKS are identified. Based on the dependency tree obtained from dependency parsing, and with the identified concept entities as central nodes, a path-first search algorithm is used to find grammatical paths connecting them. The relationships between core concepts in the context (such as "is part of," "causes," "has attributes") are analyzed to construct a classroom discussion concept graph with concepts as nodes and relationships as edges. This graph intuitively reveals the knowledge structure and logical flow of the discussion.

[0140] Step Six: Perform semantic analysis on the identified text to identify the self-regulated learning strategies demonstrated by students in the discussion;

[0141] By deeply integrating educational psychology and semantic analysis technology, semantic analysis is performed on the cleaned text to identify the self-regulating learning strategies exhibited by students in the discussion.

[0142] This analytical framework integrates the classic learning strategy classification system in educational psychology—cognitive strategies, metacognitive strategies, and resource management strategies—and establishes a corresponding keyword and expression pattern library for each type of strategy. For example, cognitive strategies correspond to specific learning methods such as "summarizing," "drawing charts," "frequently explaining problems to classmates," and "timely review"; metacognitive strategies are reflected in the planning and regulation of one's own cognitive processes, such as "monitoring attention" and "reflecting on the effectiveness of learning methods"; and resource management strategies include behaviors that optimize the learning environment and resources, such as "rationally allocating time," "organizing learning materials," and "actively seeking help from teachers."

[0143] Based on the concept map completed in the previous stage, the identified text fragments are categorized using a learning strategy classification system and projected onto nodes in the concept map. Because the graph structure provides richer contextual information, it can be used to verify, refine, and discover more complex semantic behaviors, upgrading from "bag-of-words analysis" which only targets individual words and word frequencies to "graph contextual analysis" that incorporates the entire text context. From the "content skeleton" of student discussions, the word frequencies and distribution patterns of different strategy vocabulary, as well as their contextual relationships, are statistically analyzed, thus objectively and quantitatively revealing the specific self-regulated learning strategies adopted by students in classroom discussions.

[0144] Step 7: Generate an intuitive and visual word cloud;

[0145] Word size is determined by its coreness in the concept graph (rather than simple word frequency); word color can be mapped to different dimensions.

[0146] Step 8: Display the intelligent word cloud and the timeline of the discussion process;

[0147] Provide teachers with a unified dashboard that displays: a smart word cloud (reflecting core content) and a discussion process timeline (supporting backtracking).

[0148] The following examples illustrate the solution of the present invention.

[0149] like Figure 3 The image shows the original text. After preprocessing, the text is regularized and then broken down into individual characters or tokens using the Chinese-RoBERTa token segmenter. At the same time, the fine-tuned BERTForTokenClassification model performs a three-class classification prediction for each token, thereby removing redundant repetitions and colloquialisms, eliminating a large number of repetitive and filler words without actual semantic meaning, correcting unfluent punctuation and grammatical errors, and completing incomplete sentences in spoken language; thus simplifying and optimizing the expression.

[0150] For example: Redundant repetitions and filler words like, "If you can't memorize ancient poems, I'll give you a method, how about it? It's not about rote memorization," have been removed. The phrase "I'll give you a method, how about it?" used to start conversations and buy time has been replaced with "Don't memorize by rote." Sentence breaks and grammatical errors have been corrected: "For example, you can...you can read it first," has been corrected by merging the pauses in "can...you can" into a fluent "can," etc. Expressions have been simplified: "Wait until you've read it thoroughly before you memorize it, that will be much easier," has had colloquial elements like "wait," "it," and "you" removed. Figure 4 As shown.

[0151] After semantic normalization and concept labeling modules, the model scans the text word by word, identifies words or phrases belonging to predefined educational entity categories, and labels them in the format of [entity type]. Based on the classification model results, it determines whether different expressions are synonymous and uses the constructed thesaurus for replacement and merging.

[0152] For example, "reciting ancient poems" is labeled as "[Learning Problem] Reciting Ancient Poems", and "gradual progress" is labeled as "[Core Principle] Gradual Progress", etc. Figure 5 As shown.

[0153] Traditional methods rely on coding systems such as FIAS to manually classify and count classroom speech (e.g., "teacher questions" and "student answers"). This is a behaviorist analysis method that only cares about "what type of words were said by whom" and ignores "what the specific content of the words is".

[0154] The present invention completely shifts the analysis focus from "speech act types" to "the semantics of the discussion content itself". By constructing a concept map, it can not only answer "how many times did the students speak", but also answer "which core concepts did the students discuss? What are the logical relationships between these concepts?", thus revealing the knowledge structure of classroom discussions.

[0155] Based on the framework of self-regulated learning theory, the present invention maps students' language expressions to the specific self-regulated learning strategies adopted. This is an innovative attempt to deeply integrate educational psychology and natural language processing technology, enabling the system to go beyond simple word frequency statistics and providing a new and quantifiable analysis dimension for insight into students' self-regulated learning ability and teaching effectiveness.

[0156] Traditional methods use general ASR and word cloud tools, which have poor processing capabilities for classroom spoken and subject-term-rich texts and generate noisy results. The present invention designs a complete automated pipeline from audio acquisition → voiceprint recognition → speech transcription → text cleaning → semantic analysis. In particular, it deeply combines the speaker separation technology with ASR to automatically generate structured texts with identity markers, providing a solid foundation for subsequent interactive network analysis and solving the problems of subjectivity and lag in manual recording.

[0157] Instead of directly using general NLP tools, the present invention specifically designs preprocessing and analysis modules for classroom texts: while selectively filtering noise, the present invention retains logical words that reflect the thinking process; through synonym merging and concept entity recognition, the present invention injects subject knowledge into the system to make the analysis results more accurate and educational.

[0158] When traditional methods output data, the analysis results are usually a static report or an isolated word cloud map, which is disconnected from the classroom teaching process and difficult to directly guide teaching actions. The present invention integrates various visualization forms such as intelligent word clouds, concept maps, interactive network diagrams, and discussion timelines into one interface through a publish-subscribe mode. All visualization components subscribe to a common data status center. When a user operates on a component (such as clicking), it does not directly modify other components but "publishes" an event to the status center. After the status center updates the data, all components that have subscribed to this data will automatically re-render, thus showing a linkage effect. Through component linkage, this provides teachers with an all-round and explorable "digital twin" of classroom discussions, greatly enhancing the intuitiveness and depth of analysis.

[0159] In response to the characteristics of primary school students in the upper grades (fourth and fifth grades) who speak quickly, use colloquial language, engage in a lot of repetition and inversion, and include specific subject-specific terminology in classroom discussions, this system adopts a deep transfer learning strategy based on the pre-trained language model (BERT).

[0160] I. Dataset Construction and Preprocessing

[0161] Data Sources and Feature Analysis: The training data comes from real fourth and fifth grade elementary school Chinese language classroom group discussions and interview recordings. The dataset has the following significant features, which are the focus of model fine-tuning:

[0162] The speech is rich in noise: it contains a lot of meaningless repetitions (such as “I…I…I think”), filler words (such as “that”, “just”), and informal expressions (such as “I’m not happy”, “It’s closed now”).

[0163] It is highly subject-specific: it contains a large number of specialized terms for Chinese language teaching, such as "daily accumulation", "writing words based on pinyin", "similar-looking characters", "sentence reduction", "sentence expansion", and "reading and writing integration".

[0164] Self-regulated learning (SRL) has clear semantics: it contains rich reflective statements (such as "inadequate review", "careless", "set small goals"), making it suitable for training policy analysis models.

[0165] Data annotation strategy: In order to train the various modules described in the patent, the original corpus was manually annotated in multiple dimensions.

[0166] Sequence Labeling for Text Cleaning: The BIO labeling system is adopted. Valid components in the transcribed text are marked as O (retain); stuttering, repetition, and meaningless interjections (such as "um..." or "that...") are marked as B-DELETE (delete the beginning) and I-DELETE (delete the middle).

[0167] Entity and Concept Labeling (NER): Mark "mistake notebook", "mind map", "preview", "dictation" and other terms in the text as B-EDU_TOOL (learning tool) or B-STRATEGY (learning behavior).

[0168] Synonym pair annotation: Extract semantically similar expressions from the corpus to construct positive sample pairs. For example, construct synonym or related concept pairs such as "write words based on pinyin" and "basic questions", and "recognize characters" and "new characters" to train the semantic normalization module.

[0169] II. Fine-tuning of the text cleaning module

[0170] This module aims to remove redundant noise from elementary school students' spoken language while preserving educational features. BERT-base-chinese is chosen as the pre-training foundation, with a fully connected (dense) layer at the top for classification. Text cleaning is defined as a character-level three-class classification task (Token Classification): categories 0 (retain), 1 (delete the beginning), and 2 (delete the continuation). Since the number of valid characters to be retained (category 0) in the corpus far exceeds the number of noisy characters to be deleted (categories 1 and 2), there is a severe sample imbalance. This embodiment uses Focal Loss instead of the traditional Cross Entropy Loss, as shown in the following formula:

[0171]

[0172] in, It is the model's predicted probability of the true class. The focusing parameter is used to reduce the contribution of easily classified samples to the loss and emphasize difficult-to-classify samples; The class weight factor is used to alleviate the class imbalance problem. By assigning greater weight to the minority class and less weight to the majority class, the model can reduce the bias towards a large number of simple samples during training and improve its ability to identify minority and difficult samples.

[0173] set up To focus on noisy samples that are difficult to classify; to set weighting factors This reduces the weight of a large number of simple samples (reserved words) and increases the model's sensitivity to stuttering and redundant words.

[0174] III. Training of the Semantic Normalization and Entity Recognition Module

[0175] Synonym Determination: Input [CLS] Sentence A [SEP] Sentence B [SEP]. Construct positive samples such as ("Review vocabulary at the end of the book", "Recite the vocabulary list") and negative samples such as ("Review ancient poems", "Preview the text"), and determine whether the two refer to the same type of learning behavior.

[0176] Educational Entity Recognition: In addition to traditional BIO annotation, the EDU_ENTITY tag is added. The model is trained to identify specific entities that frequently appear in the audio, such as: “Daily Accumulation” (textbook section), “Resource Bag”, “53” (workbook), and “Error Notebook”.

[0177] Fine-tuning strategy: Freeze the underlying parameters of BERT (the first 6 layers) and only fine-tune the high-level parameters to retain general language understanding capabilities while adapting to the specific terminology distribution in the field of primary education.

[0178] To verify the effectiveness of the "Intelligent Analysis System for Student Self-Regulated Learning Strategies" described in this invention in real educational scenarios, particularly the technical advantages of the data acquisition and preprocessing layer and the core analysis layer in handling high classroom noise, colloquial language, and subject terminology recognition, this embodiment constructs a rigorous comparative experiment.

[0179] IV. Experiment Setup

[0180] Experimental Dataset: The experiment selected real-world audio recordings of group discussions in Chinese language classrooms as test samples, aiming to simulate the most challenging real-world application environment. Three audio recordings were selected, each approximately 4 minutes long. The audio environment exhibited typical classroom characteristics, including multiple speakers taking turns speaking, background noise interference, and frequently used colloquial expressions by students (such as repetition and pauses).

[0181] Compared to the baseline model: Our model is specifically optimized for educational scenarios, integrating text cleaning and semantic normalization modules, and is able to handle specific subject-specific terms and logical connectors.

[0182] Comparison Model 1 (Vosk): A lightweight speech recognition model developed based on the Kaldi toolkit. It is ultra-lightweight and supports streaming processing on embedded devices, and is often used as a benchmark for offline recognition.

[0183] Comparison Model 2 (Wav2Vec 2.0): An end-to-end speech recognition model based on self-supervised pre-training and CTC fine-tuning architecture, representing the mainstream technical level in low-resource languages ​​in general domains.

[0184] Speech-to-text evaluation metrics:

[0185] Word Error Rate (WER):

[0186] Word error rate measures the edit distance between the recognized text and the standard reference text, and is a core indicator for evaluating basic transcription accuracy.

[0187]

[0188] Where N is the number of words in the reference text, and S, I, and D represent the number of words replaced, inserted, and deleted, respectively.

[0189] Sentence Error Rate (SER):

[0190] Sentence error rate measures semantic integrity at the sentence level. A sentence is considered to have an error if it contains any core semantic deviation or error in key information.

[0191]

[0192] This metric reflects whether the text generated by the system is sufficient to support subsequent semantic understanding and strategy analysis.

[0193] Keyword Extraction F1 Score:

[0194] The keyword extraction F1 score is specifically used to assess the system's ability to capture keywords related to "self-regulated learning strategies" (such as "review", "memorize", "plan").

[0195]

[0196] in, This indicates the number of correctly extracted target items (i.e., the number of prediction results that match the manual annotations, also known as the number of true positives). This represents the total number of target items extracted by the system (the number predicted to be positive). This represents the total number of target items (the number of truly positive ones) in the manually annotated (reference answer) entries. Precision rate represents the proportion of correct items in the extracted results; Recall rate represents the proportion of reference answers that are successfully extracted. It is the harmonic mean of precision and recall, used to comprehensively evaluate extraction performance.

[0197] V. Analysis of Experimental Results

[0198] Speech-to-text performance comparison:

[0199] The transcription performance comparisons of the three models on the audio test set are shown in Tables 1, 2, and 3. Table 1 shows the comparison results of the speech transcription performance of audio 1, Table 2 shows the comparison results of the speech transcription performance of audio 2, and Table 3 shows the comparison results of the speech transcription performance of audio 3.

[0200] Table 1

[0201]

[0202] Table 2

[0203]

[0204] Table 3

[0205]

[0206] The performance comparison of speech transcription from Audio 1 to Audio 3 shows that the model proposed in this invention exhibits extremely high robustness and stability in classroom speech environments of varying difficulty. In all three test sets, the word accuracy of this system consistently remained at a high level, ranging from 94.8% to 96.3%. In contrast, the word accuracy of the Vosk model fluctuated between 68.4% and 73.4%, while the Wav2Vec 2.0 model even dropped to 45.2% in Audio 3. This indicates that existing general-purpose models struggle to adapt to the complexity of classroom environments, while this system, through domain-specific optimization, effectively overcomes noise and spoken language interference. Sentence accuracy is a key indicator for evaluating whether text can be used for subsequent semantic analysis. The sentence accuracy of this system remained between 84.0% and 88.0%, meaning that the vast majority of sentences retained complete semantic logic. In contrast, the highest sentence accuracy of the Vosk model was only 32.0%, and Wav2Vec 2.0 even fell to 4.0% in Audio 3. The data fully demonstrates that the text generated by existing technologies suffers from severe "semantic fragmentation" and cannot support in-depth pedagogical analysis, while the text generated by this system has extremely high usability.

[0207] Learning strategy keyword extraction performance:

[0208] Statistical analysis was conducted on the core strategy vocabulary involved in the classroom discussion. The results are shown in Tables 4, 5, and 6. Table 4 shows the comparison results of keyword extraction performance for Audio 1, Table 5 shows the comparison results of keyword extraction performance for Audio 2, and Table 6 shows the comparison results of keyword extraction performance for Audio 3.

[0209] Table 4

[0210]

[0211] Table 5

[0212]

[0213] Table 6

[0214]

[0215] The F1 score is a comprehensive indicator for measuring the extraction effect. The F1 scores of this system in Audio 1 and Audio 2 reached 83.9% and 98.2% respectively, far exceeding those of Vosk (73.8% - 85.1%) and Wav2Vec 2.0 (47.6% - 53.8%). The recall rate directly reflects whether the system "missed" the students' strategies. The recall rate of Wav2Vec 2.0 is extremely low (34.5% - 39.7%), meaning that more than 60% of the learning strategy information is lost. The recall rate of this system can reach up to 96.5% at most, proving that the system can keenly capture almost all strategy-related signals in students' discussions, laying a foundation for subsequent work.

[0216] Analysis of typical error cases:

[0217] To deeply explore the specific advantages of this system in dealing with complex classroom speech environments, this embodiment conducts a deep semantic comparison on the experimental data. The results show that compared with the frequent "semantic hallucinations" and "logical collapses" in the prior art, the model of this invention (Our) demonstrates extremely high accuracy and robustness when dealing with subject terms, core strategy words, and long and difficult sentences.

[0218] Accurate recognition of teaching-specific terms:

[0219] Accurately recognizing Chinese teaching concepts such as "imitative writing" and "four-character phrases" is the bottom line for the system to be usable. The prior art often shows serious semantic distortions in such terms, and even generates misleading negative words.

[0220] Standard text: "I want to consolidate the four-character phrase 'turning complexity into simplicity' in the book."

[0221] The Vosk model had extremely unreasonable semantic errors and misrecognized it as "drawing plants into simplicity". This kind of error of misrecognizing a standard idiom in a teaching scenario as a non-standard phrase will directly interfere with the system's judgment of students' mastery of idioms and expression accuracy, resulting in deviations in subsequent teaching diagnoses. This system accurately recognizes it as "turning complexity into simplicity" and at the same time retains the teaching category term "four-character phrase", with the characters and meanings being completely correct, ensuring the reliability and zero ambiguity of the subsequent analysis of students' knowledge point mastery.

[0222] Standard text: "And there is also imitative writing of sentences."

[0223] The Vosk model completely lost the core verb "imitation writing," incorrectly identifying it as "industry data collection," causing the context to change from "Chinese language practice" to "industrial data collection," resulting in a 100% semantic deviation. This system identifies it as "anti-writing sentence." Although homophones exist (anti- / imitation), this system fully preserves the initial and final vowel structure. Combined with the "semantic normalization module" and "educational whitelist" described in this invention, it is easily corrected into standard terminology.

[0224] Accurate capture of core learning strategy words

[0225] The core task of this system is to analyze students' "self-regulated learning strategies." Experiments have shown that existing technologies are prone to missing key strategic intentions, while this system can accurately pinpoint this high-value information.

[0226] Standard text: "All stem from strategies accumulated over time."

[0227] The Vosk model failed to recognize the idiom, incorrectly outputting "originating from the strategy of the sun never setting." It mistakenly replaced "accumulation" with "the sun never sets," preventing the system from determining whether the student had grasped the cognitive strategy of "long-term persistence." Our system, however, accurately output "originating from the strategy of daily accumulation," word for word. This demonstrates our system's significant advantage in capturing long idioms and complex strategic concepts.

[0228] Standard text: "A strategy for seeking challenges."

[0229] The Vosk model misidentified it as "Planet Challenge," while the Wav2Vec model identified it as "New Ball Battle." These errors completely masked the student's positive learning attitude of "actively seeking challenging tasks." Our system accurately identified it as "a strategy for seeking challenges," not only with accurate verbs but also by preserving interjections, providing strong support for subsequent sentiment analysis.

[0230] Semantic robustness in complex spoken contexts:

[0231] In classroom discussions, students often use inversions, obscure words, or ellipses. Existing technologies often create an "illusion" as a result, while this system can accurately reproduce the true context.

[0232] Standard text: "Sometimes the unit introduction will also be tested."

[0233] The Vosk model recognized it as "big fainting also good", and the Wav2Vec model recognized it as "bring someone back good". Both of these comparison models produced completely meaningless gibberish, resulting in the complete loss of information in this sentence. This system recognized it as "single-person lead". Although "yuan" was recognized as the homophonous "person", the core concept of "lead" was accurately captured. Compared with the "fainting" and "bringing someone" of the comparison models, the result of this system is highly semantically close to the true value and has extremely high comprehensibility and reparability.

[0234] Standard text: "It's just in the corners."

[0235] The Vosk model couldn't handle such rhotacized sounds and rare words and misrecognized it as "subway things", forcefully associating the semantics with means of transportation. This system recognized it as "ridge feet drooping". This system successfully captured the phonetic features of the four syllables, avoiding the irrelevant "semantic reconstruction" like in the existing technology, proving that the acoustic model has extremely strong stability when dealing with non-standard classroom spoken language.

[0236] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent analysis system for students' self-regulating learning strategies, characterized in that: The system includes a data acquisition and preprocessing layer, a text processing and enhancement layer, a core analysis layer, and an application display layer. The data acquisition and preprocessing layer includes an audio acquisition module and a speech transcription and structuring module; the audio acquisition module is used for hardware control and raw acquisition of audio streams, and the speech transcription and structuring module is used to convert audio streams into text with speaker tags; The text processing and enhancement layer includes a text cleaning module and a semantic normalization module. The text cleaning module is used to clean the text, remove noise, and retain valuable educational features. The semantic normalization module is used to improve the standardization and semantic clarity of the text. The core analysis layer includes a concept graph analysis module and a learning strategy analysis module; the concept graph analysis module is used to construct the knowledge network of the current classroom discussion, and the learning strategy analysis module is used to analyze the discussion content from the perspective of teaching objectives. The application presentation layer includes a visualization rendering module and a dashboard integration module; the visualization rendering module is used to transform data into intuitive graphics, and the dashboard integration module is used to uniformly manage and display all visualization components. The concept mapping analysis module is implemented as follows: The discussion text is segmented and filtered by part of speech, retaining only content words as candidate keywords to form a set of nodes in the graph. A sliding window of size 7 is defined, and the co-occurrence relationship between any two candidate words in the window is counted. If two words co-occur in the window, an edge is established between them, and the number of co-occurrences is used as the initial weight of the edge. Introducing word frequency factors into the edge weight calculation gives higher weight to the connections between high-frequency core words, while assigning higher initial scores to words that appear in key positions. The importance score of each word is obtained by iterative calculation. After multiple rounds of iteration until the score converges, all candidate words are sorted in descending order according to the final score, and the K words with the highest ranking are selected as the concept set that best represents the core of the discussion. The trained concept entity recognition model is used to identify entities, dependency parsing is used to obtain dependency trees, and the identified concept entities are used as central nodes. The path-first search algorithm is used to find grammatical paths connecting them, the relationship between core concepts in the context is analyzed, and a classroom discussion concept graph is constructed with concepts as nodes and relationships as edges.

2. The intelligent analysis system for student self-regulation learning strategies according to claim 1, characterized in that: The audio acquisition module includes: Device driver unit: controls the start / stop and parameter settings of the microphone array; Streaming data receiving unit: continuously receives raw streaming data from the audio acquisition device and performs buffering, packetization, and timestamp marking; The speech transcription and structuring module includes: Automatic Speech Recognition Unit: Calls the trained and optimized ASR engine API to convert the audio stream into raw text with timestamps; Speaker Separation and Recognition Unit: Analyzes the audio, distinguishes different speakers, and aligns and binds the identified speaker IDs with the text generated by ASR; Text formatting unit: Generates the final result in a standard structured log format.

3. The intelligent analysis system for student self-regulation learning strategies according to claim 1, characterized in that: The text cleaning module includes: Noise filtering unit: Based on preset rules and a vocabulary database, it filters out verbal tics and repetitive words; Educational Feature Identification and Retention Unit: Identifies and ensures that key features are not removed through keyword matching or simple rules; The semantic normalization module includes: Thesaurus Management Unit: Maintains a thesaurus for the education field, providing query and merging services; Concept Entity Recognition Unit: Identifies key entities in the text and assigns type labels to these key entities.

4. The intelligent analysis system for student self-regulation learning strategies according to claim 1, characterized in that: The concept mapping analysis module includes: Core concept extraction unit: Extracts core keywords from the text as concept nodes; Relation extraction unit: Utilizes a pre-trained model to analyze the grammatical relations of concepts in a sentence and extract the semantic relations between concepts; Graph Construction and Storage Unit: Constructs nodes and relationships into a graph data structure for storage; The learning strategy analysis module is based on a self-regulating learning model. It identifies self-regulating learning strategies that students directly indicate or imply in their language. It scans cleaned text, matches keywords and contextual expressions related to the learning strategies, and statistically analyzes the frequency, distribution patterns, and contextual relationships of different strategy words. This reveals the specific self-regulating learning strategies that students use in classroom discussions.

5. The intelligent analysis system for student self-regulation learning strategies according to claim 1, characterized in that: The visualization rendering module includes: Intelligent word cloud generation unit: Receives data from the concept graph analysis module and the learning strategy analysis module, and generates a multi-dimensional word cloud; Timeline rendering unit: Based on the initial structured text log, it generates an interactive timeline, which can be clicked to view the discussion content at a specific moment; The instrument panel integration module includes: Layout Management Unit: Responsible for arranging visual components on the webpage; Data Interface Unit: Retrieves data from the core analytics layer and database, and distributes it to various visualization components.

6. The intelligent analysis system for student self-regulation learning strategies according to claim 1, characterized in that: The text cleaning module employs a BERT fine-tuning method based on sequence labeling to filter text noise and preserve educational features. The implementation process is as follows: Based on the BERT text model, we fine-tuned the noise filtering as a three-class classification task: O - Keep / B - DELETE - Delete the beginning / I - DELETE - Delete the middle. Using the Focal Loss loss function, a weight factor α is introduced, and α is preset for each class k. k Weight; By using rule-based annotation strategies, a clear "whitelist" of terms can be established to avoid misjudgments.

7. The intelligent analysis system for student self-regulation learning strategies according to claim 6, characterized in that: The semantic normalization module is implemented as follows: A multi-task learning framework is constructed. For synonym merging, a fine-tuned BERT model is used to classify sentence pairs, and manual annotation is used to determine whether the expressions are synonymous. A thesaurus is built to merge different words that express the same concept into standard terms. Based on the traditional BIO annotation method, new labels for educational entities are added, and the model is fine-tuned into a sequence labeling system, enabling it to identify and label lexical segments in the text that represent educational entities.

8. The intelligent analysis system for student self-regulation learning strategies according to claim 1, characterized in that: The learning strategy analysis module is implemented as follows: A classification system for learning strategies was constructed, and a corresponding library of keywords and expression patterns was established for each type of strategy. Based on the concept graph obtained by the concept graph analysis module, the identified text fragments are classified according to the learning strategy classification system and projected onto the nodes of the concept graph. The word frequency and distribution patterns of different strategy words and their contextual relationships are statistically analyzed, thereby showing the specific self-regulating learning strategies adopted by students in classroom discussions.

9. The intelligent analysis system for student self-regulation learning strategies according to claim 1, characterized in that: The system operates as follows: The audio acquisition module was used to collect multimodal audio data of student discussions in class. The collected audio data is converted into speech and processed into structure using the speech-to-text and structuring modules to obtain the processed text data. The text cleaning module is used to filter noise and retain educational features from the obtained text data, resulting in cleaned text data. The semantic normalization module is used to perform semantic normalization and concept annotation on the cleaned text data; The concept graph analysis module is used to construct a concept graph of classroom discussions with concepts as nodes and relationships as edges, so as to deeply interpret classroom discussions from multiple dimensions. The learning strategy analysis module is used to perform semantic analysis on the identified text in order to identify the self-regulated learning strategies demonstrated by students in the discussion. Use the visualization rendering module to generate intuitive and visual word clouds; The dashboard integration module is used to generate a dashboard that displays a smart word cloud and a timeline of the discussion process.

Citation Information

Patent Citations

  • Knowledge content weight analysis system and method based on knowledge graph

    CN107644062A

  • Method for intelligently generating exhaustion report of financial unfavorable assets

    CN120198232A