A Method and System for Extracting and Analyzing Knowledge Points in First-Person Classroom Based on AI Glasses

CN122569747APending Publication Date: 2026-08-14SHENZHEN SMART CLOUD TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

因此,本发明提供一种基于AI眼镜的第一视角课堂知识点提炼分析方法与系统,旨在解决现有技术在课堂教学场景下知识提取不准确、结构化程度低、知识点溯源困难等技术问题

Benefits of technology

[0015]本发明的有益效果:本申请通过AI眼镜获取课堂第一视角音视频流,并结合视觉锚点与语音锚点建立跨模态对齐关系,实现教师讲解内容与板书、课件内容之间的对应关联,降低仅依赖语音转写或单纯视觉识别造成的知识点偏移与语义断裂问题;通过内容坐标与时间范围构建对齐记录,实现课堂知识点在不同讲解阶段下的稳定定位、关联回溯与证据追踪;通过断言模板生成候选知识断言,并结合证据覆盖、跨模态一致性、时序一致性与反事实核验,实现课堂知识点的结构化提炼与一致性校验,提高第一视角课堂知识提炼结果的准确性、可归并性与可解释性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122569747A_ABST
    Figure CN122569747A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for extracting and analyzing classroom knowledge points from a first-person perspective using AI glasses, relating to the field of artificial intelligence technology. The method includes: acquiring, segmenting, and recognizing first-person audio and video streams to obtain a sequence of event segments; extracting visual and audio anchor point sets based on the sequences and associating them to obtain an anchor point enhancement sequence; performing spatiotemporal alignment on the two anchors to generate an aligned record sequence; constructing candidate knowledge assertions and performing evidence package construction, consistency verification, and merging to obtain a set of knowledge assertions; constructing an argument graph to generate a minimum evidence playback set and a playback set sequence corresponding to the backtracking sequence. This invention achieves stable cross-perspective positioning through content coordinates and accurately binds verbal explanations with visual content using cross-modal temporal adjacency pairing, effectively solving the problems of inaccurate knowledge extraction, low structure, and difficulty in tracing in existing technologies in classroom teaching scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for extracting and analyzing classroom knowledge points based on first-person perspective using AI glasses. Background Technology

[0002] With the rapid development of artificial intelligence technology, AI glasses, as a new generation of wearable smart devices, are gradually showing broad application prospects in education, healthcare, and industry. In the education field, how to use AI glasses to intelligently extract and record classroom knowledge has become a hot research topic. However, existing AI glasses technology still has significant shortcomings in knowledge extraction in classroom teaching scenarios, making it difficult to meet the actual needs of classroom learning.

[0003] Currently, Chinese invention patent application CN121364777A discloses AI smart glasses, including a hardware system (main control chip, camera module, microphone array, etc.) and a software system (image processing module, visual analysis module, semantic generation module, etc.). It possesses functions such as image recognition, knowledge reasoning, and language generation, and claims to be applicable to educational scenarios. However, this technology lacks the ability to handle specific classroom scenarios, failing to distinguish between teacher lectures and student speeches, filtering background noise and student chatter, and failing to achieve spatiotemporal alignment and association between visual content (blackboard writing, PPT) and audio explanations. Furthermore, this solution lacks a mechanism for prioritizing the extraction of knowledge points required by curriculum standards, failing to meet the actual needs of classroom teaching for extracting core knowledge points. Summary of the Invention

[0004] The technical problem addressed by this invention is that while existing AI glasses technology possesses certain audio and video processing and multimodal interaction capabilities, it still suffers from the following major shortcomings in the specific application scenario of classroom knowledge extraction: First, it lacks dedicated large-scale model fine-tuning technology for classroom scenarios, making it unable to effectively distinguish between teacher lectures and student speeches, or filter non-teaching audio; second, it lacks spatiotemporal alignment technology to accurately bind visual content (formulas, charts, blackboard writing) and audio explanations to the same timeline; third, it lacks a system function that dynamically generates mind maps based on classroom progress and supports knowledge point tracing. Therefore, this invention provides a first-person perspective classroom knowledge point extraction and analysis method and system based on AI glasses, aiming to solve the technical problems of inaccurate knowledge extraction, low degree of structure, and difficulty in knowledge point tracing in existing technologies in classroom teaching scenarios.

[0005] To address the aforementioned technical problems, this invention provides the following technical solution: a method for extracting and analyzing classroom knowledge points based on the first-person perspective of AI glasses, comprising the following steps: Step S1: Obtain the first-person audio and video stream, and perform event segmentation, speaker category identification, and semantic pattern recognition on the first-person audio and video stream to obtain an event segment sequence; Step S2: Extract the visual anchor set and the speech anchor set based on the event segment sequence, and associate the visual anchor set and the speech anchor set with the corresponding event segments to obtain the anchor-enhanced event segment sequence; Step S3: Perform dual-anchor alignment on the visual anchor set and the speech anchor set to generate an aligned record sequence; Step S4: Construct candidate knowledge assertions based on the aligned record sequence, and perform evidence package construction, consistency verification and assertion merging processing on the candidate knowledge assertions to obtain a set of knowledge assertions; Step S5: Construct an argument graph based on the knowledge assertion set, and generate a minimum evidence replay set and a minimum evidence replay set sequence corresponding to the backtracking sequence based on the argument graph.

[0006] As a preferred embodiment of the first-person perspective classroom knowledge point extraction and analysis method based on AI glasses described in this invention, step S1 specifically includes: Step S11: Obtain the first-view audio and video stream, and determine the event boundaries in the first-view audio and video stream. The event boundaries include audio content change event boundaries, video content change event boundaries, and audio and video content change event boundaries. Step S12: Divide the first-view audio and video stream into a sequence of event segments arranged in the order of events and without overlap, based on the event boundaries, and record the corresponding time range for each event segment. Step S13: Segment the audio portion of each event segment to obtain a sequence of speech segments; Step S14: Determine the speaker category for each speech segment in the speech segment sequence, and determine the role label for the event segment based on the speaker category corresponding to each speech segment in the speech segment sequence. The role label includes teacher lecturing, student asking questions, classroom management, and idle chatter. Step S15: Perform semantic pattern recognition on the semantic analysis text corresponding to the speech fragment sequence of each event fragment, and determine the explanation status label for the event fragment based on the distribution results of each semantic pattern in the semantic analysis text. The explanation status label includes background definition, derivation and proof, example explanation, summary and emphasis, and question and answer clarification. Step S16: Associate and store the time range, role tag, and explanation status tag of the event segment to obtain the event segment sequence.

[0007] As a preferred embodiment of the first-person perspective classroom knowledge point extraction and analysis method based on AI glasses described in this invention, step S2 specifically includes: Step S21: For each event segment in the event segment sequence, determine the candidate content region from the video frame sequence corresponding to the event segment. The candidate content region includes text region, symbol region and chart region. Step S22: Perform text and symbol parsing on the candidate content area to obtain visual content units, and determine the content type corresponding to each visual content unit; Step S23: Determine the visual content unit as a visual anchor point according to the content type of the visual content unit, and form a visual anchor point set by combining all visual anchor points within the same event segment. The set of visual anchor points includes title anchor points, number anchor points, formula anchor points, icon anchor points, and whiteboard block anchor points; Step S24: For each visual anchor point in the visual anchor point set, record the corresponding video frame position and the time position of the video frame within the event segment time range, and store the video frame position, time position and the visual anchor point in association. Step S25: For each event segment in the event segment sequence, perform speech-to-text transcription on the corresponding audio part of the event segment to obtain anchor-transcribed text; Step S26: Identify semantic anchor trigger expressions in the anchor-transcribed text, determine the anchor-transcribed text segments corresponding to the semantic anchor trigger expressions as speech anchors, and form a speech anchor set from all speech anchors; The voice anchors include definition anchors, introduction anchors, emphasis anchors, error-prone anchors, and question anchors; Step S27: Record the time position of each voice anchor point in the set of voice anchor points within the time range of the event segment, and store the time position associated with the voice anchor point. Step S28: Associate and store the visual anchor point set and voice anchor point set corresponding to each event segment with the event segment to obtain the anchor point enhanced event segment sequence.

[0008] As a preferred embodiment of the first-person perspective classroom knowledge point extraction and analysis method based on AI glasses described in this invention, step S3 specifically includes: Step S31: For each anchor enhancement event segment in the anchor enhancement event segment sequence, read the time range of the anchor enhancement event segment, the set of visual anchors corresponding to the anchor enhancement event segment, and the set of speech anchors corresponding to the anchor enhancement event segment. Step S32: Determine the content coordinates for each visual anchor point in the set of visual anchor points. The content coordinates include the carrier identifier, page identifier, and region identifier. Step S33: Determine the speech time position for each speech anchor in the speech anchor set, and determine the visual time position for each visual anchor in the visual anchor set. Step S34: Within the same anchor point enhancement event segment, for each voice anchor point, select visual anchor points in the visual anchor point set that satisfy the temporal adjacency relationship, and establish an anchor point pairing relationship between the voice anchor point and the selected visual anchor points.

[0009] As a preferred embodiment of the first-person perspective classroom knowledge point extraction and analysis method based on AI glasses described in this invention, step S3 further includes: Step S35: Generate an alignment record for each anchor point pairing relationship. The alignment record includes the event segment identifier of the corresponding anchor point enhanced event segment, the time range determined by the voice time position of the voice anchor point and the visual time position of the visual anchor point, the content coordinates of the visual anchor point, the voice anchor point identifier, and the visual anchor point identifier. Step S36: All alignment records generated within the same anchor point enhancement event segment are arranged in chronological order to form an alignment record sequence, and the alignment record sequence is associated with and stored with the corresponding anchor point enhancement event segment to obtain the alignment record sequence.

[0010] As a preferred embodiment of the first-person perspective classroom knowledge point extraction and analysis method based on AI glasses described in this invention, step S4 specifically includes: Step S41: For each anchor point enhancement event segment in the anchor point enhancement event segment sequence, read the explanation status label corresponding to the anchor point enhancement event segment and the alignment record sequence corresponding to the anchor point enhancement event segment; Step S42: Based on the explanation status label, select the assertion template corresponding to the explanation status label from the preset assertion template set; The pre-set set of assertion templates includes templates for object properties and conditions, templates for steps, conclusions, templates for question types, templates for common mistakes, and templates for summarizing and merging. Step S43: For each alignment record in the alignment record sequence, read the time range, content coordinates, voice anchor identifier and visual anchor identifier in the alignment record, and combine the anchor transcribed text corresponding to the voice anchor identifier with the visual content unit corresponding to the visual anchor identifier based on the assertion template to generate candidate knowledge assertions. Step S44: Assign an assertion type to each candidate knowledge assertion, wherein the assertion type is determined by the assertion template, and establish an association between the candidate knowledge assertion and the alignment record that generated the candidate knowledge assertion. Step S45: All candidate knowledge assertions generated within the same anchor point augmentation event segment are arranged in chronological order to form a candidate knowledge assertion set, and the candidate knowledge assertion set is associated with the corresponding anchor point augmentation event segment. Step S46: For each candidate knowledge assertion in the candidate knowledge assertion set, construct an evidence package. The processing logic is as follows: Read the time range and content coordinates of the aligned record associated with the candidate knowledge assertion; Select video frames from the video frame sequence corresponding to the visual anchor point identifier that correspond to the visual content unit as the keyframe set; Extract the anchor point transcribed text segments corresponding to the speech anchor point identifiers from the anchor point transcribed text corresponding to the time range, and form a set of transcribed text segments from the anchor point transcribed texts; The time range, content coordinates, keyframe set, and transcribed fragment set are combined into an evidence package, and the evidence package is associated with the candidate knowledge assertion.

[0011] As a preferred embodiment of the first-person perspective classroom knowledge point extraction and analysis method based on AI glasses described in this invention, step S4 further includes: Step S47: Perform consistency verification for each candidate knowledge assertion to obtain consistency verification results. The consistency verification results include evidence coverage results, cross-modal consistency verification results, temporal consistency verification results, and counterfactual verification results. Step S48: Combine each candidate knowledge assertion with its corresponding evidence package and consistency verification result into a knowledge assertion record, generate a confidence level label for the knowledge assertion record, merge all knowledge assertion records, group the knowledge assertion records by content coordinates as the merging key, merge knowledge assertion records with the same assertion atom set into the same knowledge assertion within each group, and merge the evidence packages corresponding to the merged knowledge assertion records into an evidence package set. Step S49: Output the merged knowledge assertions as a knowledge assertion set, and establish an association between the knowledge assertion set and the corresponding content coordinates for storage.

[0012] As a preferred embodiment of the first-person perspective classroom knowledge point extraction and analysis method based on AI glasses described in this invention, step S5 specifically includes: Step S51: Construct an argument graph, read the knowledge assertion set, assign an argument graph node identifier to each knowledge assertion in the knowledge assertion set, and map each knowledge assertion to an argument graph node. Each node in the argument graph is associated with and stores the assertion type, assertion text, content coordinates, evidence package set, confidence level label, and assertion atom set of the knowledge assertion. Step S52: Establish dependency edges, derivation edges, applicable edges, and misconception edges between nodes in the argument graph, and establish association storage between each argument graph edge and the corresponding evidence package set.

[0013] As a preferred embodiment of the first-person perspective classroom knowledge point extraction and analysis method based on AI glasses described in this invention, step S5 further includes: Step S53: For each node or edge of the argument diagram, read the associated set of evidence packages, group the evidence package sets according to content coordinates, and perform redundancy reduction processing on evidence packages with the same content coordinates and overlapping time ranges. Within each group, select keyframes from the keyframe set from the evidence package corresponding to that group, and select transcribed segments from the transcribed segment set. Combine the keyframes, transcribed segments, and their corresponding time ranges and content coordinates into the playback entries of that group. Merge the playback entries of each group into the minimum evidence playback set and output it. Step S54: According to the backtracking protocol, determine the backtracking path corresponding to the node or edge of the argument graph in the argument graph, and form a backtracking sequence by combining the argument graph nodes and edges in the backtracking path in the traversal order. Read the evidence package set corresponding to the backtracking sequence, and output the minimum evidence replay set sequence corresponding to the backtracking sequence based on the evidence package set.

[0014] Secondly, the first-person classroom knowledge point extraction and analysis system based on AI glasses includes an event fragment construction module, an anchor point enhancement module, a double anchor alignment module, a knowledge assertion generation module, and an argument playback module. The event segment construction module is used to acquire first-person audio and video streams, and to perform event segmentation, speaker category recognition, and semantic pattern recognition on the first-person audio and video streams to obtain an event segment sequence; The anchor point enhancement module is used to extract a set of visual anchor points and a set of voice anchor points based on an event segment sequence, and associate the set of visual anchor points and the set of voice anchor points with the corresponding event segments to obtain an anchor point enhanced event segment sequence. The dual-anchor alignment module is used to perform dual-anchor alignment on the visual anchor point set and the voice anchor point set to generate an alignment record sequence. The knowledge assertion generation module is used to construct candidate knowledge assertions based on the aligned record sequence, and to perform evidence package construction, consistency verification and assertion merging processing on the candidate knowledge assertions to obtain a set of knowledge assertions; The argument replay module is used to construct an argument graph based on the knowledge assertion set, and generate a minimum evidence replay set and a minimum evidence replay set sequence corresponding to the backtracking sequence according to the argument graph.

[0015] The beneficial effects of this invention are as follows: This application acquires first-person perspective audio and video streams in the classroom through AI glasses, and establishes cross-modal alignment relationships by combining visual anchor points and voice anchor points, realizing the correspondence between the teacher's explanation content and the content of the blackboard and courseware, reducing the problems of knowledge point offset and semantic breakage caused by relying solely on speech transcription or simple visual recognition; by constructing alignment records through content coordinates and time ranges, it realizes stable positioning, correlation backtracking and evidence tracking of classroom knowledge points at different explanation stages; by generating candidate knowledge assertions through assertion templates, and combining evidence coverage, cross-modal consistency, temporal consistency and counterfactual verification, it realizes the structured extraction and consistency verification of classroom knowledge points, improving the accuracy, merging and interpretability of the first-person perspective classroom knowledge extraction results. Attached Figure Description

[0016] Figure 1 A flowchart illustrating the steps of a first-person perspective classroom knowledge point extraction and analysis method based on AI glasses, as provided in one embodiment of the present invention; Figure 2 A schematic diagram of the basic process of a first-person classroom knowledge point extraction and analysis system based on AI glasses, provided as an embodiment of the present invention; Figure 3 A flowchart for cross-modal pairing of voice and visual anchors; Figure 4 A flowchart for consistency verification and merging of knowledge assertions. Detailed Implementation

[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0018] Example, refer to Figure 1 , 3 4. A method for extracting and analyzing knowledge points from a first-person classroom perspective using AI glasses, including the following steps: Step S1: Obtain the first-person audio and video stream, and perform event segmentation, speaker category identification, and semantic pattern recognition on the first-person audio and video stream to obtain an event segment sequence; Step S2: Extract the visual anchor set and the speech anchor set based on the event segment sequence, and associate the visual anchor set and the speech anchor set with the corresponding event segments to obtain the anchor-enhanced event segment sequence; Step S3: Perform dual-anchor alignment on the visual anchor set and the speech anchor set to generate an aligned record sequence; Step S4: Construct candidate knowledge assertions based on the aligned record sequence, and perform evidence package construction, consistency verification and assertion merging processing on the candidate knowledge assertions to obtain a set of knowledge assertions; Step S5: Construct an argument graph based on the knowledge assertion set, and generate a minimum evidence replay set and a minimum evidence replay set sequence corresponding to the backtracking sequence based on the argument graph.

[0019] In specific implementation, step S1 includes: Step S11: Obtain the first-view audio and video stream, and determine the event boundaries in the first-view audio and video stream. The event boundaries include the audio content change event boundary, the video content change event boundary, and the audio and video content change event boundary. Step S12: Divide the first-view audio and video stream into a sequence of event segments arranged in the order of events and without overlap, based on the event boundaries, and record the corresponding time range for each event segment. Step S13: Segment the audio portion of each event segment to obtain a sequence of speech segments; Step S14: Determine the speaker category for each speech segment in the speech segment sequence, and determine the role label for the event segment based on the speaker category corresponding to each speech segment in the speech segment sequence. The role labels include teacher lecturing, student asking questions, classroom management, and idle chatter. Step S15: Perform semantic pattern recognition on the semantic analysis text corresponding to the speech fragment sequence of each event fragment, and determine the explanation status label for the event fragment based on the distribution results of each semantic pattern in the semantic analysis text. The explanation status label includes defining background, derivation and proof, example explanation, summary and emphasis, and question and answer clarification. Step S16: Associate and store the time range, role tag, and explanation status tag of the event segment to obtain the event segment sequence.

[0020] Specifically, first-person perspective audio and video streams are typically acquired simultaneously by the front-facing camera and microphone of AI glasses. During acquisition, video frames and audio sampling blocks are uniformly bound to the same time base timestamp. The time base can be the device's monotonic clock or an absolute time after network synchronization, thus ensuring that the audio and video are aligned on the same timeline. The acquired data can be recorded locally and written into a packaged file containing video and audio tracks, retaining the display timestamp of each frame or audio block. Alternatively, it can be sent to a mobile phone, edge computing device, or server in real-time, carrying the acquisition end timestamp for synchronization and playback positioning at the receiving end. This completes the acquisition of the first-person perspective audio and video stream, and subsequently, the corresponding video frame sequence and audio segment can be accurately read within any time range. Based on this time-axis audio and video stream, when determining event boundaries and dividing event segments on the timeline, we can first detect audio content change boundaries, video content change boundaries, and audio-video content change boundaries: Audio content change boundaries can be triggered by the switching point between long silences and speaking states detected by speech activity detection, or by changes in the speaker, or by calculating audio feature vectors through a sliding window and detecting sudden increases in differences between adjacent windows; Video content change boundaries can be triggered by a significant decrease in the structural similarity between adjacent frames or a significant increase in the difference in color histograms, or by a sudden drop in the similarity of text recognition results in content areas such as blackboards, projections, or books within adjacent time windows, or by drastic changes in perspective; the above boundaries can then be considered as potential boundaries. After selecting points and sorting them by time, fusion and deduplication are performed within a preset time tolerance window. This merges multiple candidate points falling within the same time tolerance window into a final boundary. When audio and video are triggered simultaneously within the same time tolerance window, they can be marked as audio-video joint change boundaries to improve confidence. Finally, the start and end times of two adjacent final boundaries constitute the time range of an event segment, forming a sequence of event segments arranged in order of events and without overlap. The start and end timestamps and duration of each event segment are recorded. In engineering, minimum and maximum segment constraints can also be added. Excessively short segments are merged with adjacent segments, and excessively long segments are further divided internally according to secondary boundaries to improve the integrity and consistency of each segment.In this embodiment, the preset time tolerance window is set to 1 second, which can absorb the inherent time jitter and delay difference between audio and video detection. For example, classroom behaviors such as teachers speaking before writing on the blackboard, pointing to the screen before explaining, and teachers turning around to write after students ask questions usually cause misalignment between audio and video changes for hundreds of milliseconds to about one second. The one-second tolerance window can merge two types of changes that are essentially the same classroom event, avoiding cutting the explanation process of the same knowledge point into two or more fragments, significantly reducing the chain reaction problems caused by excessive segmentation. After the number of event fragments is reduced, the computational load and redundancy of subsequent steps such as speech segmentation, semantic pattern recognition, anchor point extraction, double anchor alignment, and evidence package construction will decrease. At the same time, the event fragments are more complete, the explanation status labels are more stable, and the candidate knowledge assertions are less likely to have breaks or missing evidence, improving the usability of double anchor alignment and evidence playback. Because the boundary is closer to the rhythm of the real classroom, it is easier to contain the corresponding voice anchor and visual anchor within the same event fragment, so the time range of the alignment record is more concentrated, the key frames of the evidence package are more consistent with the transcribed fragments, and the final generated minimum evidence playback set is shorter and more accurate, and it is also more in line with the human understanding order when reviewing.

[0021] When segmenting the audio portion of each event segment to obtain a sequence of speech segments, speech activity detection is generally performed on the audio of the event segment first to separate continuous speech intervals from silence and noise. Speech intervals with very short intervals are merged, and extremely short noise segments are filtered to obtain a set of speech intervals. In order to make the speech segments closer to the real speaking rounds, speaker separation and speaker switching detection can also be performed within the speech intervals. For example, the speaker embedding of each short window is extracted and clustered and change point detection is performed. When a speaker switching is detected, the speech interval is cut at the switching point, and finally a sequence of speech segments with start and end times is obtained. When determining the speaker category for each speech segment in a speech segment sequence and further determining the role label for the event segment, the speaker category can first be defined as a set of categories such as teacher, student, and others. If pre-registration of teacher voiceprints is allowed, the speaker embedding is extracted for each speech segment and the similarity is calculated with the teacher voiceprint. If the similarity is higher than the threshold, the speaker is identified as a teacher; otherwise, the speaker is identified as a non-teacher. The student is further distinguished from others by combining the volume distance feature and semantic intent feature. If pre-registration of teacher voiceprints is not allowed, multiple speaker clusters can be obtained by speaker separation first. Then, teacher candidate clusters can be selected according to rules such as the proportion of total speaking time in the classroom, speaking continuity, and the proportion of explanatory language. A light calibration is performed when the program is first used. The role labels of event segments are obtained by aggregating the speaker categories and semantic intent of the speech segments within the event segment. For example, when the teacher's speaking time exceeds the threshold and the intention to ask questions is weak, it is labeled as teacher-led; when students show obvious intention to ask questions and their turn to speak is prominent, it is labeled as student questioning; when expressions related to classroom organization appear, such as quiet, turning to a page, taking attendance, or collecting homework, it is labeled as classroom management; and when the confidence of the identified text is consistently low, multiple people overlap, and the content is weakly related to the classroom topic, it is labeled as idle chatter or noise.When performing semantic pattern recognition and determining the explanation state label for the semantic analysis text corresponding to the speech fragment sequence of each event segment, the speech fragments within the event segment can first be transcribed into speech and concatenated into paragraph text according to time. Then, the paragraph text can be segmented into sentences and feature extracted. Semantic pattern recognition can adopt a method of trigger expression and sentence pattern rules plus distribution statistics. For example, defining the background often involves expressions such as definition, called, we let, denoted as, concept, property, etc.; derivation and proof often involve expressions such as proof, derivation, therefore, so, thus obtained, deduced, thus, equivalent, etc.; example explanation often involves examples, known, find, solve, ask questions, substitute, steps; summarizing and emphasizing often involves summarizing, pay attention, emphasize, conclusion, key, common mistakes, etc.; question and answer clarification often involves your question, further explanation, in other words, is it, why, etc. Then, the number of hits, coverage ratio, and time concentration of each type of trigger expression in the event segment text are statistically analyzed to form a distribution result. The one with the largest distribution result is used as the explanation state label of the event segment. Alternatively, a trained text classification model can be used to directly output the probability distribution of each explanation state and determine the label according to the highest probability. If necessary, thresholds and priorities can be added to avoid misjudgment. When storing the event segment sequence by associating the time range, role tags, and explanation status tags of event segments, a unique event segment identifier can be assigned to each event segment. The start timestamp, end timestamp, role tag, explanation status tag, and necessary reference information are written into the same event segment record. The reference information includes the list of speaking segments corresponding to the event segment, the start and end times of each speaking segment, the speaker category, and the time location information of the transcribed segment. It also includes the location information of the original audio and video stream, such as the file location and timeline mapping or the stream address and timestamp range. Then, the output is sorted by the start timestamp and persistently saved as an event segment sequence, so that subsequent steps can stably read back the corresponding audio and video evidence and text data by event segment identifier or by time range.

[0022] In specific implementation, step S2 includes: Step S21: For each event segment in the event segment sequence, determine the candidate content region from the video frame sequence corresponding to the event segment. The candidate content region includes text region, symbol region and chart region. Step S22: Perform text and symbol parsing on the candidate content area to obtain visual content units, and determine the content type corresponding to each visual content unit; Step S23: Determine the visual content unit as a visual anchor point according to the content type of the visual content unit, and form a visual anchor point set by combining all visual anchor points within the same event segment. The set of visual anchors includes title anchors, number anchors, formula anchors, icon anchors, and whiteboard block anchors; Step S24: For each visual anchor point in the visual anchor point set, record the corresponding video frame position and the time position of the video frame within the event segment time range, and store the video frame position, time position and the visual anchor point in association. Step S25: For each event segment in the event segment sequence, perform speech-to-text transcription on the corresponding audio part of the event segment to obtain anchor-transcribed text; Step S26: Identify semantic anchor trigger expressions in the anchor-transcribed text, determine the anchor-transcribed text segments corresponding to the semantic anchor trigger expressions as speech anchors, and form a speech anchor set from all speech anchors; Voice anchors include definition anchors, introduction anchors, emphasis anchors, error-prone anchors, and question anchors; Step S27: Record the time position of each voice anchor point in the set of voice anchor points within the time range of the event segment, and store the time position associated with the voice anchor point. Step S28: Associate and store the visual anchor point set and voice anchor point set corresponding to each event segment with the event segment to obtain the anchor point enhanced event segment sequence.

[0023] Specifically, when determining candidate content regions in the video frame sequence corresponding to an event segment, several representative keyframes are first selected from the video of that event segment as detection carriers. Keyframes can be extracted at fixed intervals, such as one frame per second, or frames can be extracted during periods when the image is stable, thereby reducing false detections caused by head shaking. Subsequently, the classroom content carrier plane is preferentially determined on the keyframes, namely the blackboard area, projection screen area, lecture book area, or tablet screen area. The carrier plane can be determined through object detection or semantic segmentation, and the feature criteria include large rectangular planes, edge regularity, texture consistency, and similarity to other elements. The co-occurrence relationship of chalk or projected text is analyzed. Layout analysis and region proposal are then performed within the carrying plane. Text candidate boxes are output through a text detection network, symbol candidate boxes are output through symbol detection or high-contrast connected component analysis, and chart candidate boxes are output through chart detection models or rule recognition of structures such as coordinate axes, table lines, and flowcharts. Finally, overlap elimination and temporal consistency tracking are performed on all candidate boxes. Candidate boxes that are close in position and similar in content in adjacent keyframes are merged into stable candidate content regions, thus obtaining three types of candidate content regions: text regions, symbol regions, and chart regions. Each region has corresponding intra-frame position coordinates.

[0024] When parsing text and symbols in candidate content regions, the processing logic typically involves image enhancement followed by recognition: For text regions, denoising, perspective correction, tilt correction, and contrast enhancement are performed, then a character recognition model is invoked to output the text sequence and confidence score, while preserving the local bounding box positions line by line or word by word. For symbol regions, a mathematical or physical / chemical symbol recognition process is used to segment character blocks within the candidate region into basic strokes or connected components. A symbol classification model is then used to recognize plus signs, minus signs, equal signs, arrows, Greek letters, integral signs, and square roots. If necessary, further structural parsing is performed to assemble the symbols into formulas based on spatial relationships. For chart regions, their structural elements are parsed, extracting coordinate axes, scales, table headers, flow lines, geometric shapes, illustrative arrows, and legend text, and outputting these elements as referable content items. After the above recognition and structuring, visual content units are generated. Each visual content unit includes at least the recognition result, region location coordinates, frame identifier, and recognition confidence score, facilitating subsequent anchor references and evidence playback.

[0025] The content type of a visual content unit is used to categorize visual content units into different anchor categories. In this embodiment, the content type can be defined as title, number, formula, icon, and blackboard block. Titles are usually located at the top of the display plane or at the beginning of a paragraph and have a relatively larger font or present chapter words such as chapter number, section number, knowledge point, or example. Numbers usually conform to a serial number pattern such as I, II, III, or I-II-III, or I-II-III-IV, or I-II-III-IV-V, or numbers with decimal points or parentheses. Formulas usually contain mathematical structures or physical and chemical expressions such as equal signs, inequalities, fractions, radicals, superscripts and subscripts, and nested parentheses. Icons are usually represented as geometric figures, coordinate graphs, function graphs, force diagrams, circuit diagrams, chemical structure diagrams, or flowcharts. Blackboard blocks are used to represent a collection of content units that are spatially clustered within the same time period, forming a blackboard area that can be referenced as a whole. Content type determination can be made by combining positional features, size features, text pattern features, and structural features. For example, content conforming to a numbering pattern and located near the beginning of a line is classified as numbered; content containing formula structures is classified as formulas; content located at the top and containing chapter terms is classified as titles; content with coordinate axes or closed graphic outlines is classified as icons; and the rest, appearing densely in the same spatial block and increasing in number over time, can be classified as whiteboard blocks. After determining the content type, visual content units can be identified as visual anchors according to their type and aggregated into a set of visual anchors within the same event segment. Simultaneously, the corresponding video frame position and the time position of that frame relative to the event segment's time range are recorded for each visual anchor, which is the time offset obtained by subtracting the start timestamp of the event segment from the frame's timestamp. This provides a calculable temporal basis for subsequent alignment.

[0026] When transcribing the audio portion corresponding to an event segment, a streaming or offline automatic speech recognition engine can be used. First, the audio of the event segment is enhanced and denoised. Then, speech activity detection is used to remove long silent intervals to improve recognition stability. The text is then input into the speech recognition model to output anchor-transcribed text. To support subsequent anchor time localization, the transcribing result should ideally include word-level or phrase-level timestamps, that is, the start and end times of each word or phrase. The time offset is calculated based on the start time of the event segment. In addition, sentence segmentation and punctuation restoration, as well as the injection of hot words from the domain vocabulary, such as the pronunciation of common terms and symbols in mathematics, physics, and chemistry, can be performed on the anchor-transcribed text to reduce misrecognition of proper nouns.

[0027] When identifying semantic anchor trigger expressions in anchor-transcribed text and determining the corresponding anchor-transcribed text segments as speech anchors, a combination of rule dictionary and classification model can be used: First, establish a trigger expression library, setting the trigger expressions for defining anchors as defined as, called, we let, denoted as, concept is, property is; setting the trigger expressions for derivation anchors as so, therefore, thus, derived, deduced, thereby, can be obtained, equivalent to; setting the trigger expressions for emphasis anchors as pay attention, key point, must, especially important, remember; setting the trigger expressions for error-prone anchors as easy to make mistakes, common errors, do not confuse, error-prone points; setting the trigger expressions for question anchors as why, is it, can you explain, I want to ask, where do I not understand; scan these trigger expressions in anchor-transcribed text and combine them with syntactic clues to determine their true pragmatic function. At the same time, an intent classification model can be used to verify sentence-level text to avoid misjudging verbal tics or irrelevant terms as anchors. After determining the location of the triggering expression, it needs to be expanded into anchor-transcribed text fragments that can be used for knowledge extraction. The expansion strategy can be set to extract text from the word containing the triggering expression to the end of the sentence or the next long pause. The long pause can be determined by the silence interval or punctuation boundary in the speech recognition output, and an upper limit should be set on the fragment length to avoid mixing multiple knowledge points in the same anchor. The temporal position of the fragment is determined by the start time of the first word and the end time of the last word in the fragment, or by taking the time of the triggering expression word as the anchor center time and saving its time offset relative to the event fragment. In this way, definition anchors, inference anchors, emphasis anchors, error-prone anchors, and question anchors in the anchor-transcribed text can be identified, and the temporal position of each speech anchor within the event fragment time range can be recorded. Finally, it is associated with the event fragment along with the visual anchors to form an anchor-enhanced event fragment sequence, providing input for subsequent double-anchor alignment and knowledge assertion construction.

[0028] In specific implementation, step S3 includes: Step S31: For each anchor enhancement event segment in the anchor enhancement event segment sequence, read the time range of the anchor enhancement event segment, the set of visual anchors corresponding to the anchor enhancement event segment, and the set of speech anchors corresponding to the anchor enhancement event segment. Step S32: Determine the content coordinates for each visual anchor point in the set of visual anchor points. The content coordinates include the carrier identifier, page identifier, and region identifier. Step S33: Determine the speech time position for each speech anchor in the speech anchor set, and determine the visual time position for each visual anchor in the visual anchor set. Step S34: Within the same anchor point enhancement event segment, for each voice anchor point, select visual anchor points in the visual anchor point set that satisfy the temporal adjacency relationship, and establish an anchor point pairing relationship between the voice anchor point and the selected visual anchor points.

[0029] Specifically, determining the content coordinates can begin by identifying the teaching medium where the visual anchor point is located. The teaching medium refers to a medium capable of supporting a stable layout structure and reusing positioning rules, such as projected courseware, electronic whiteboard courseware, textbook pages, lecture notes pages, or workbook pages. When an external teaching medium is detected, the medium type can be identified and a layout baseline established. The medium identifier will no longer use abstract camera image coordinates, but rather the page identifier or paragraph identifier within the external medium to indicate the location source. The page identifier can come from the page number of the courseware or the page number area text obtained through visual recognition, or it can be obtained by matching the overall page feature fingerprint with a known courseware page database. The paragraph identifier can be obtained by parsing the heading hierarchy or numbering structure, such as the section number, knowledge point number, or example question number. Subsequently, region identifiers are determined within page or paragraph identifiers. Region identifiers can be defined as element identifiers within that page. Element identifiers are generated from the set of elements obtained through layout analysis, such as heading elements, numbered elements, formula elements, graphic elements, table elements, and body text elements. Each element within the page has a unique number or unique hash. The generation method can be sequential numbering from top to bottom and left to right, or a stable identifier can be generated using the element's detection box position and content fingerprint. The resulting content coordinates are composed of carrier identifiers, page identifiers, and region identifiers, indicating which external carrier, page, and element within that page the anchor point belongs to, thus achieving stable positioning across shooting perspectives and time periods.

[0030] When no external teaching medium is available, content coordinates can be constructed by placing them on the whiteboard plane. First, the whiteboard plane (the planar area of ​​the blackboard or whiteboard) is detected in the video frame, and perspective correction is used to obtain a unified standard whiteboard plane coordinate system. The whiteboard plane identifier then serves as the medium identifier. To ensure consistent area numbering for the same piece of whiteboard material across different frames, the standard plane can be divided into grids according to fixed rules, such as dividing it into several grid units by rows and columns. Alternatively, the whiteboard content can be clustered to obtain several whiteboard blocks, which are then numbered. The area identifier is then determined as the grid unit identifier or whiteboard block identifier. During generation, it can be assigned based on which grid unit the center point of the visual anchor point falls into, or based on the maximum intersection-union ratio (IUU) between the anchor point detection box and which whiteboard block. In a whiteboard scenario, the page identifier can be fixed as the current whiteboard plane's session page, or it can be incremented using the whiteboard plane content fingerprint to generate new page identifiers when there are significant temporal updates, thus enabling pagination management when the whiteboard expands across time. By using a dual-path design of external carrier coordinate system and blackboard coordinate system, it can be ensured that stable content coordinates that can be merged, retrieved and replayed can be generated for each visual anchor point, whether viewing projected courseware or blackboard writing.

[0031] The establishment of anchor pairing relationships takes temporal adjacency as the primary constraint. The implementation method is to first determine the speech time position for speech anchors and the visual time position for visual anchors. The speech time position can be the start time of the word that triggers the expression or the center time of the speech anchor segment. The visual time position can be the timestamp of the frame where the visual anchor first appears, the timestamp of the clearest frame, or the timestamp of the frame with the highest confidence in multi-frame tracking. Then, for each audio anchor within the same enhanced event segment, visual anchors that satisfy the temporal adjacency relationship are selected from the set of visual anchors. The temporal adjacency relationship can be defined as the visual time position falling within a preset time window before or after the audio time position. The preset time window can be 1 second or 1.5 seconds based on classroom behavior experience. After obtaining the candidate set, if there are multiple visual anchors that satisfy the temporal adjacency relationship, the visual anchor closest to the audio anchor in time sequence is selected to establish a pair. The closest determination can be achieved by minimizing the absolute time difference between the audio time position and the visual time position. When there are still ties, disambiguation can be performed using the content type priority of visual anchors. For example, anchors are defined to be paired with title anchors or formula anchors first, which leads to anchors being paired with formula anchors first, emphasis anchors being paired with blackboard block anchors or numbered anchors first, and question anchors being paired with the most recently updated blackboard block anchors first, so that the pairing is more in line with the classroom explanation pattern.

[0032] This application significantly improves the accuracy and interpretability of knowledge point extraction by further transforming the visual and audio information in anchor-enhanced event fragments into a computable, traceable, and reusable alignment basis. Through the introduction of content coordinates, visual anchors are no longer just detection boxes within a single frame, but are mapped to pages and elements of external teaching materials, or to grids or blocks on the whiteboard. This allows for reliable merging of the same knowledge point when it appears at different times or is repeatedly explained, and enables subsequent evidence playback to directly locate the corresponding page and area, rather than relying on unstable pixel coordinates. Through temporally adjacent cross-modal pairing, the system can establish a correspondence between the definitions, derivations, emphasis, and questions presented by the teacher and the titles, numbers, formulas, or diagrams currently being displayed. This reduces drift caused by relying solely on audio summaries and minimizes the inability to understand semantic key points through visual recognition alone. Ultimately, this results in a more concentrated time range for alignment records, more consistent keyframes and transcribed fragments in the evidence package, more reliable generated knowledge assertions, and easier understanding of which sentence, whiteboard, or slide in the classroom the knowledge assertion originated from during playback and replay.

[0033] In specific implementation, step S3 also includes: Step S35: Generate an alignment record for each anchor point pairing relationship. The alignment record includes the event segment identifier of the corresponding anchor point enhanced event segment, the time range determined by the speech time position of the speech anchor point and the visual time position of the visual anchor point, the content coordinates of the visual anchor point, the speech anchor point identifier, and the visual anchor point identifier. Step S36: All alignment records generated within the same anchor point enhancement event segment are arranged in chronological order to form an alignment record sequence, and the alignment record sequence is associated with and stored with the corresponding anchor point enhancement event segment to obtain the alignment record sequence.

[0034] Specifically, when generating an alignment record for each anchor pairing, the anchor pairing can be regarded as a predetermined cross-modal binding, that is, a certain voice anchor corresponds to a certain visual anchor. After reading the pairing, the system first obtains the event segment identifier of the enhanced event segment of the anchor to which it belongs, then reads the voice anchor identifier and the visual anchor identifier respectively, and extracts the voice time position from the voice anchor, the visual time position from the visual anchor, and the already determined content coordinates of the visual anchor. Then, these fields are written into the same structured record to form an alignment record. The time range is determined by jointly defining the audio and visual time positions. Specifically, the audio and visual time positions can be unified to a timestamp under the same time base or a time offset relative to the start of the event segment. When an audio anchor point corresponds to a single audio time position, the earlier of the audio and visual time positions is used as the start time of the time range, and the later of the audio and visual time positions is used as the end time of the time range. A preset buffer duration is extended before the start time and after the end time of the time range. When an audio anchor point corresponds to an anchor point transcribed text segment and has an audio segment start time and an audio segment end time, the smaller of the audio segment start time and the visual time position is used as the start time of the time range, and the larger of the audio segment end time and the visual time position is used as the end time of the time range. A preset buffer duration is extended before the start time and after the end time of the time range. The extended time range is then cropped to the time range of the corresponding anchor point enhanced event segment. After the above fields are filled in, the alignment record contains event fragment identifiers, time ranges, content coordinates of visual anchors, voice anchor identifiers, and visual anchor identifiers. Alignment record identifiers can also be generated for subsequent indexing. In this way, each alignment record can directly locate a key phrase in a certain period of time in the classroom and a specific area in a certain blackboard or a certain page of courseware, providing a readable and deterministic basis for subsequent candidate knowledge assertion generation and evidence package construction.

[0035] The time range satisfies: ; ; in, This represents the start time of the speech segment corresponding to the speech anchor point. This represents the end time of the speech segment corresponding to the speech anchor point. For the visual temporal position of the visual anchor paired with the voice anchor, For the preset buffer duration, Add the start time of the event fragment corresponding to the current anchor point. Add the end time of the event fragment corresponding to the current anchor point. The start time of the time range. This is the end time of the time range.

[0036] In specific implementation, step S4 includes: Step S41: For each anchor point enhancement event segment in the anchor point enhancement event segment sequence, read the explanation status label corresponding to the anchor point enhancement event segment and the alignment record sequence corresponding to the anchor point enhancement event segment; Step S42: Based on the explanation status label, select the assertion template corresponding to the explanation status label from the preset assertion template set; The pre-set set of assertion templates includes templates for object properties and conditions, templates for steps, conclusions, templates for question types, templates for common mistakes, and templates for summarizing and merging. Step S43: For each alignment record in the alignment record sequence, read the time range, content coordinates, voice anchor identifier and visual anchor identifier in the alignment record, and combine the anchor text fragment corresponding to the voice anchor identifier with the visual content unit corresponding to the visual anchor identifier based on the assertion template to generate candidate knowledge assertions. Step S44: Assign an assertion type to each candidate knowledge assertion. The assertion type is determined by the assertion template. Also, associate the candidate knowledge assertion with the alignment record that generated the candidate knowledge assertion. Step S45: All candidate knowledge assertions generated within the same anchor point augmentation event segment are arranged in chronological order to form a candidate knowledge assertion set, and the candidate knowledge assertion set is associated with the corresponding anchor point augmentation event segment. Step S46: For each candidate knowledge assertion in the candidate knowledge assertion set, construct an evidence package. The processing logic is as follows: Read the time range and content coordinates of the aligned record associated with the candidate knowledge assertion; Select video frames from the video frame sequence corresponding to the visual anchor point identifier that correspond to the visual content unit as the keyframe set; Extract the anchor point transcribed text segments corresponding to the speech anchor point identifiers from the anchor point transcribed text corresponding to the time range, and form a set of transcribed text segments from the anchor point transcribed texts; The time range, content coordinates, keyframe set, and transcribed fragment set are combined into an evidence package, and the evidence package is associated with the candidate knowledge assertion.

[0037] Specifically, the pre-defined assertion template set is a text generation skeleton used to organize oral explanations and visual content such as blackboard projections in the classroom into standardized knowledge expressions. Each template specifies the slot fields to be extracted, the source of the slots, and the organization order of the final assertion text. This embodiment includes object property condition templates, step basis conclusion templates, question type routine and common mistake templates, and summary merging templates. Among them, the object property condition template is used to express concept definitions, object attributes, applicable conditions, and limiting conditions. The template semantics can be set to the object field satisfying the property field when the condition field is true. The condition field can be empty or extracted by definition anchors and emphasis anchors. The object field often comes from keywords and symbols in title anchors, number anchors, or blackboard block anchors. The property field often comes from formula anchors or property descriptions in anchor transcribing text. The step basis conclusion template is used to express the derivation and proof chain. The template semantics can be set to obtain the conclusion field based on the basis field and through the step field. The basis field usually comes from the previous established assertion or from the formulas and rules that have appeared in the same content coordinate. The templates for questions, steps, and common mistakes are primarily derived from explanatory statements near the anchor points, while conclusions are primarily derived from formulas or blackboard result lines corresponding to visual anchor points. The template semantics can be set to target question type fields, use routine fields, and pay attention to error-prone and correction fields. Question type fields are often extracted from numbered or title anchor points in example explanations, routine fields are extracted from verbal step descriptions by the teacher, and error-prone fields are extracted from text fragments triggered by error-prone anchor points. The summary and merging templates are used to express periodic summaries and generalizations. The template semantics can be set to summarize topic fields to obtain a set of key points and applicable scope fields. Topic fields typically come from title anchor points or blackboard block anchor points, and the set of key points is obtained by aggregating multiple sentences from emphasis anchor points and summary emphasis states. When explaining the use of status labels for template selection, deterministic mapping rules can be adopted. For example, defining the background prioritizes the object property condition template, derivation and proof prioritizes the step basis conclusion template, example explanation prioritizes the step basis conclusion template and question type routine and error-prone template, summary and emphasis prioritize the summary and merging template, and question and answer clarification prioritizes the reuse of the object property condition template or the step basis conclusion template and adds clarification fields or limiting condition fields in the assertion.

[0038] When combining the anchor-transcribed text fragments corresponding to voice anchors with the visual content units corresponding to visual anchors based on assertion templates, the process can be implemented by first aligning, then filling slots, and finally generating. Since the alignment record already provides the time range and content coordinates, the system first uses the voice anchors to locate the anchor-transcribed text fragments related to the triggering expression within that time range, and then uses the visual anchors to locate the visual content units within the same time range, such as a formula, a conclusion line, a number, or a whiteboard block. Subsequently, keyword and structure extraction is performed on the anchor-transcribed text fragments, extracting object names, conditional phrases, causal connectors, derivation words, emphasis words, and negation correction words, etc. Then, the visual content units are read in a structured manner. For example, formula anchors are parsed into left-hand and right-hand expressions, title anchors are read into chapter topics, number anchors are read into question numbers or step numbers, and whiteboard block anchors are read into a collection of several lines of content within that block. After extraction, this information is written into the slots of the assertion template to generate candidate knowledge assertion text and structured fields. For example, under the object property condition template, core terms from the title or blackboard block are filled into the object field, the applicable scope appearing in the audio text is filled into the condition field, and visual formulas or audio property descriptions are filled into the property field, thus generating assertions such as "an object satisfies a certain property under certain conditions." Under the step-based conclusion template, the basis and derivation sentences from the audio segment are filled into the basis field and step field, and the circled or final result from the visual content unit is filled into the conclusion field, thus generating assertions such as "a conclusion is derived from a certain basis." The key point of this approach is that audio provides semantic relationships and explanatory intent, while visuals provide symbolic results and precise writing. After the two are bound together by the same content coordinates and the same time range, the combined assertions can express meaning and also be applied to specific blackboard or courseware elements.

[0039] When assigning an assertion type to each candidate knowledge assertion and establishing an association with the alignment record that generated the candidate knowledge assertion, a template-driven deterministic type assignment method can be used. Since the template identifier used by the candidate assertion is already obtained when selecting the template, the assertion type can be directly taken as the type code or type name that corresponds one-to-one with the template. For example, a candidate assertion generated using the object property condition template is labeled as object property condition; a candidate assertion generated using the step-based conclusion template is labeled as step-based conclusion; a candidate assertion generated using the question type, pattern, and error-prone template is labeled as question type, pattern, and error-prone; and a candidate assertion generated using the summary and merging template is labeled as summary and merging. Meanwhile, to ensure the construction and replayability of subsequent evidence packages, each candidate knowledge assertion is written with the identifier of its source alignment record as the association key during generation, and the event fragment identifier, time range, and content coordinates are simultaneously written as redundant index fields, thus forming a traceable link: the candidate knowledge assertion can return to the alignment record through the association key, and then return to the specific voice anchor point and visual anchor point, as well as its time range and content coordinates through the alignment record. This ensures that the key frame set and transcription fragment set can be stably extracted and the evidence package can be constructed. It also ensures that when merging assertions, the content coordinates can be used as the merging key to merge repeated explanations on the same whiteboard area or the same courseware element into the same knowledge assertion.

[0040] In specific implementation, step S4 also includes: Step S47: Perform consistency verification for each candidate knowledge assertion to obtain consistency verification results. The consistency verification results include evidence coverage results, cross-modal consistency verification results, temporal consistency verification results, and counterfactual verification results. Step S48: Combine each candidate knowledge assertion with its corresponding evidence package and consistency verification result into a knowledge assertion record, generate a confidence level label for the knowledge assertion record, merge all knowledge assertion records, group the knowledge assertion records by content coordinates as the merging key, merge knowledge assertion records with the same assertion atom set into the same knowledge assertion within each group, and merge the evidence packages corresponding to the merged knowledge assertion records into an evidence package set. Step S49: Output the merged knowledge assertions as a knowledge assertion set, and establish an association between the knowledge assertion set and the corresponding content coordinates for storage.

[0041] Specifically, when performing consistency verification for each candidate knowledge assertion, the candidate knowledge assertion can be regarded as composed of several smallest semantic units that can be independently supported or refuted by evidence. Then, evidence coverage, cross-modal mutual verification, temporal constraint checks, and counterfactual distinguishability checks are performed around these smallest semantic units. Finally, a consistency verification result is output and input is provided for subsequent confidence level scoring. In practical implementation, the system first performs assertion atomization on candidate knowledge assertions, that is, breaks down the assertions into sets of assertion atoms. The generation method of assertion atoms is strongly related to the assertion template, so template-driven parsing rules can be adopted to break down object property condition type assertions into object atoms, property atoms, and condition atoms; step-based conclusion type assertions into basis atoms, derivation step atoms, and conclusion atoms; question type, routine, and error-prone type assertions into question type atoms, routine step atoms, error-prone point atoms, and correction atoms; and summary and merging type assertions into theme atoms and key point atom sets. If the assertion contains a formula, the formula is parsed into a structured expression and treated as an independent formula atom or equation atom, while the key symbols and variable names in the formula are written into the atom attributes as matching elements. After atomization, the system performs evidence coverage calculation based on the assertion atom set and evidence package. The evidence package contains time range, content coordinates, keyframe set, and transcribed fragment set. Therefore, the coverage calculation can determine whether each atom is supported by keyframes or transcribed fragments: For atoms with strong visualization, such as formula atoms, number atoms, and chart element atoms, the system prioritizes confirming their presence and consistency in the keyframe set through text recognition and formula structure matching. For atoms with strong semantic relationships, such as definition relationships, derivation relationships, emphasis relationships, and error-prone warning relationships, the system prioritizes confirming their presence and consistency in the transcribed fragment set through keyword matching, dependency syntax relationships, or intent classification results. Each atom receives a binary coverage tag or a continuous coverage score. The assertion-level evidence coverage result can be obtained by dividing the number of covered atoms by the total number of atoms to get the coverage rate, or by using a weighted coverage rate. The weight is determined by the atom type, with conclusion atoms and formula atoms typically having higher weights.Next, cross-modal consistency verification results are generated. The key to this step is to ensure that the same assertion atom corroborates each other on both the visual and linguistic evidence chains. The system extracts a set of visual candidate facts from the keyframe set, such as identifying equations, inequalities, and variable names from the formula area, chapter numbers and keywords from the title number area, and axis labels and legend text from the chart area. At the same time, it extracts a set of linguistic candidate facts from the transcribed fragment set, such as concept names, conditional phrases, derivation conjunctions, conclusion phrases, and error-prone prompt phrases. Then, the same atom is matched on both sides and the matching results are merged. The matching method can be string similarity plus synonym normalization, or the formula can be converted into a standardized expression and then structural consistency judgment is performed. When an atom can be matched with consistent content in both visual and linguistic aspects, the cross-modal consistency of the atom is passed. When it appears only on one side, it is partially passed or weakly passed. When it appears on both sides but the content conflicts, it is failed. Next, a temporal consistency check is performed. This step uses the order of event fragments and the reference relationships between candidate knowledge assertions to check whether the order of assertions is reasonable. The reference relationship can come from template fields, such as the "based on" field in the step-based conclusion assertion referencing the formula on the previous page or the previous numbered step, or it can come from content coordinate references, such as the later assertion in the same content coordinate referencing the conclusion atom of the previous assertion. Based on this, the system determines the order of appearance of candidate knowledge assertions and checks whether the referenced assertion actually appears first in time, checks whether there are reverse steps in the derivation chain within the same content coordinate, and checks whether there are abnormal situations within the same assertion where the conclusion is given first and the basis is given later without a summary. The check results are output as the temporal consistency check result. Finally, counterfactual verification is performed to assess the distinguishability of the assertions. The system constructs a counterfactual assertion atom set based on the assertion atom set and generates candidate counterfactual knowledge assertions. The construction of counterfactuals follows the principle of minimum perturbation, meaning that only one key atom is changed as much as possible to form an assertion that is similar to the original assertion but has the opposite conclusion: for example, replacing the sign on the right side of the equation with a common confusing term, changing the greater than sign to a less than sign, changing a necessary and sufficient condition to a necessary condition, inverting the range of conditions or deleting key conditions, and changing the negative relation of a common point of error to an affirmative relation. After generating candidate counterfactual knowledge assertions, counterfactual evidence is constructed for them. The counterfactual evidence package is used to calculate the counterfactual evidence coverage and perform cross-modal consistency verification. The counterfactual evidence package usually prioritizes reusing keyframes and transcribed segments corresponding to the time range and content coordinates of the original assertion, because the goal is to verify whether the same classroom evidence can also support counterfactual evidence. If the counterfactual evidence still achieves high coverage and high cross-modal consistency under the same evidence, it indicates that the original assertion lacks distinguishability, and the counterfactual verification is judged as failing or weakly passing. If the counterfactual evidence coverage is significantly reduced under the evidence or there is obvious cross-modal consistency conflict, it indicates that the original assertion and the evidence have good distinguishability, and the counterfactual verification is judged as passing.At this point, the consistency test results can be output in a structured manner, including evidence coverage results, evidence coverage results, cross-modal consistency verification results, temporal consistency verification results, and counterfactual verification results. Each item can be refined to the atomic level, which facilitates subsequent interpretation and playback.

[0042] When generating confidence level labels for knowledge assertion records, the aforementioned verification results can be converted into a comparable comprehensive confidence score, which is then mapped to discrete levels. The comprehensive confidence score consists of four parts: evidence coverage score, cross-modal consistency score, temporal consistency score, and counterfactual distinguishability score. The evidence coverage score can be a weighted coverage rate; the cross-modal consistency score can be the percentage of passed atoms with penalties for conflicting atoms; the temporal consistency score can be full marks for passing and deducted points according to severity for failing; and the counterfactual distinguishability score can be obtained by normalizing the difference between the evidence coverage score of the original assertion and the evidence coverage score of the counterfactual assertion under the same evidence package. A larger difference indicates a higher degree of distinguishability between the original assertion and the counterfactual assertion. The comprehensive confidence score is obtained by weighted summation of the four components, using the following formula: ; in, To calculate the overall confidence score, For evidence coverage weight, For cross-modal consistency weights, For time-series consistency weights, For counterfactual distinguishability weight, For evidence coverage score, For cross-modal consistency score, For time-series consistency score, The counterfactual distinguishability score; In this embodiment, the evidence coverage weight is set to 0.4, the cross-modal consistency weight to 0.3, the temporal consistency weight to 0.1, and the counterfactual distinguishability weight to 0.2. Several hard downgrade rules are also set, such as the maximum weight not exceeding "medium" if any cross-modal conflicting atom exists, and the weight being directly judged as "low" if any conclusion atom is not covered. Finally, the overall confidence score is mapped to a confidence level label. For example, a score not lower than 0.8 is labeled as high confidence, a score between 0.6 and 0.8 is labeled as medium-high confidence, a score between 0.4 and 0.6 is labeled as medium-low confidence, and a score lower than 0.4 is labeled as low confidence. The confidence levels generated in this way reflect not only the sufficiency of evidence but also the degree of mutual corroboration between voice and vision and the counterfactual distinguishability, directly supporting subsequent assertion merging and ranking, argument graph construction, and priority selection for minimum evidence playback.

[0043] In specific implementation, step S5 includes: Step S51: Construct an argument graph, read the knowledge assertion set, assign an argument graph node identifier to each knowledge assertion in the knowledge assertion set, and map each knowledge assertion to an argument graph node. Each node in the argument graph is associated with and stores the assertion type, assertion text, content coordinates, evidence package set, confidence level label, and assertion atom set of the knowledge assertion. Step S52: Establish dependency edges, derivation edges, applicable edges, and misconception edges between nodes in the argument graph, and establish association storage between each argument graph edge and the corresponding evidence package set.

[0044] Specifically, after reading the knowledge assertion set, the system first traverses the knowledge assertions according to their generation order or chronological order within the set, generating a unique argument graph node identifier for each knowledge assertion. The argument graph node identifier can be generated by combining an event fragment identifier, a content coordinate identifier, and a node sequence number. The event fragment identifier represents the classroom event fragment from which the knowledge assertion originates; the content coordinate identifier represents the corresponding blackboard or courseware area; and the node sequence number distinguishes different knowledge assertions under the same content coordinate. After generating the argument graph node identifier, the system maps each knowledge assertion to a corresponding argument graph node and creates a node record for that node. The node record stores the assertion type, assertion text, content coordinates, evidence package set, confidence level label, and assertion atom set for that knowledge assertion. The assertion atom set describes the object atom, condition atom, formula atom, step atom, conclusion atom, or fallible atom corresponding to the knowledge assertion. Subsequently, the node identifiers of the argument graph are written into the corresponding node records, and a node index table is established according to the node identifiers of the argument graph. This enables the corresponding knowledge assertions and evidence packages to be quickly located through the node identifiers of the argument graph during subsequent graph edge creation, path traversal, and evidence playback.

[0045] When establishing dependency edges between nodes in the argument graph, for any first argument graph node and any second argument graph node, the assertion atom sets corresponding to the first and second argument graph nodes are read. It is then checked whether the assertion atom set corresponding to the first argument graph node contains references to object atoms, formula atoms, or symbol atoms corresponding to the second argument graph node. These references include identical variable names, identical formula symbols, identical object names, or identical conclusion field references. When the check is successful, a dependency edge is established between the first and second argument graph nodes, and the argument graph edge type of the dependency edge is marked as `requires`. Simultaneously, the corresponding source and target argument graph node identifiers are recorded. Dependency edges are primarily used to indicate that a knowledge assertion requires a prior knowledge assertion as a basis for understanding; for example, a formula derivation depends on a prior definition, or a conclusion depends on a prior theorem or conditional assertion.

[0046] When establishing imputation edges between nodes in the argument graph, a subset of argument graph nodes with the same content coordinates is first determined from the knowledge assertion set. These nodes are then sorted according to the start time of the time range corresponding to the evidence package sets associated with each node, generating a node order. Subsequently, for adjacent preceding and subsequent argument graph nodes in the node order, it is checked whether the assertion atom set corresponding to the subsequent argument graph node contains references to the conclusion atom, formula atom, or symbol atom corresponding to the preceding argument graph node. When the check is successful, an imputation edge is established between the preceding and subsequent argument graph nodes, and the argument graph edge type of the imputation edge is marked as "implies." The imputation edge indicates that the subsequent knowledge assertion is derived from the preceding knowledge assertion. Since the same content coordinates usually correspond to the same blackboard area, the same courseware page, or the same derivation area, a stable knowledge derivation chain can be constructed in the classroom by combining the time order with atomic reference relationships.

[0047] When establishing applicable edges between nodes in the argument graph, first, a subset of rule-type argument graph nodes whose assertion types are object property conditions, step-based conclusions, or summaries are determined from the knowledge assertion set. Then, a subset of question-type argument graph nodes whose assertion types are prone to errors are also determined. Subsequently, for any rule-type and any question-type argument graph node, it is checked whether their corresponding assertion atom sets contain shared atoms. Shared atoms include atoms with the same object, formula, condition, or symbol. When the check is successful, an applicable edge is established between the rule-type and question-type argument graph nodes, and the edge type of the applicable edge is marked as "applies-to." The applicable edge indicates that the corresponding rule applies to the corresponding question type or problem-solving scenario. For example, when a rule-type knowledge assertion and a question-type knowledge assertion simultaneously contain the same formula and condition atoms, it can be determined that the rule applies to that question type.

[0048] When establishing misconception edges between nodes in an argument graph, first, a subset of argument graph nodes in the knowledge assertion set that contain assertion types that are prone to errors, misconceptions, incorrect examples, or confusing explanations is identified. Then, for any misconception-type argument graph node and any non-misconception-type argument graph node, it is checked whether their corresponding assertion atom sets contain shared or conflicting atoms. Conflicting atoms include those with opposite sign directions, opposite condition ranges, inconsistent formula results, or inconsistent object relationships. When the detection is successful, a misconception edge is established between the misconception-type and non-misconception-type argument graph nodes, and the argument graph edge type of the misconception edge is marked as "misconception-of." The misconception edge represents the error-correction correspondence between misconception knowledge assertions and correct knowledge assertions. For example, a misconception edge can be established when the formula atoms in a misconception-type knowledge assertion and the formula atoms in a correct knowledge assertion differ only in sign direction or condition range.

[0049] When establishing an association between each argument graph edge and its corresponding evidence package set, for each argument graph edge, the evidence package sets associated with the source argument graph node and the target argument graph node connected to that edge are read, and a merging process is performed on the two evidence package sets. The merging process includes time range deduplication, content coordinate deduplication, and keyframe duplication reduction, thereby generating an edge evidence package set corresponding to that argument graph edge. Subsequently, the edge evidence package set is associated with the corresponding argument graph edge, and the argument graph edge type, source argument graph node identifier, target argument graph node identifier, and edge evidence package set identifier are recorded in the corresponding argument graph edge. Among them, the evidence packages in the edge evidence package set are used to represent the source of the knowledge relationship corresponding to the argument graph edge, including the corresponding keyframe set, transcription fragment set, time range, and content coordinate. In this way, each argument graph edge can not only represent the logical relationship between knowledge assertions, but also directly associate with the classroom evidence that forms this logical relationship, so that subsequent backtracking path generation, minimum evidence playback set generation, and argument chain interpretation can all be located and played back based on the evidence package set corresponding to the argument graph edge.

[0050] In specific implementation, step S5 also includes: Step S53: For each node or edge of the argument diagram, read the associated set of evidence packages, group the evidence package sets according to content coordinates, and perform redundancy reduction processing on evidence packages with the same content coordinates and overlapping time ranges. Within each group, select keyframes from the keyframe set from the evidence package corresponding to that group, and select transcribed segments from the transcribed segment set. Combine the keyframes, transcribed segments, and their corresponding time ranges and content coordinates into the playback entries of that group. Merge the playback entries of each group into the minimum evidence playback set and output it. Step S54: According to the backtracking protocol, determine the backtracking path corresponding to the node or edge of the argument graph in the argument graph, and form a backtracking sequence by combining the argument graph nodes and edges in the backtracking path in the traversal order. Read the evidence package set corresponding to the backtracking sequence, and output the minimum evidence replay set sequence corresponding to the backtracking sequence based on the evidence package set.

[0051] Specifically, after reading the set of evidence packages associated with a node or edge of the argument diagram, the system first reads the content coordinates and time range corresponding to each evidence package. The content coordinates include the carrier identifier, page identifier, and region identifier, and the time range includes the start time and end time. Subsequently, the evidence package set is grouped using the content coordinates as the grouping key, so that evidence packages with the same content coordinates are grouped into the same evidence group. Since the same content coordinates correspond to the same whiteboard area, the same courseware page area, or the same knowledge content area, the evidence packages in the same group correspond to the repeated or continuous explanation process of the same knowledge point in different time periods.

[0052] After grouping, the system performs redundancy reduction processing on evidence packages within the same evidence group. During redundancy reduction, it first checks whether the time ranges of any two evidence packages overlap. If the overlap ratio of the time ranges of two evidence packages exceeds a preset threshold, they are determined to be duplicate evidence packages. Subsequently, keyframe redundancy detection and transcription fragment redundancy detection are performed on the duplicate evidence packages. Keyframe redundancy detection includes calculating image similarity, text recognition result similarity, or formula structure similarity between keyframes. Transcription fragment redundancy detection includes calculating text similarity, keyword overlap, or assertion atom overlap between anchor point transcribed texts. When the similarity exceeds a preset threshold, only evidence packages with shorter time ranges, higher confidence levels, or higher keyframe clarity are retained, and the remaining duplicate evidence packages are deleted. This method reduces the large amount of repetitive classroom evidence generated during repeated explanations of the same knowledge point, thereby reducing redundancy in subsequent replays.

[0053] After redundancy reduction, within each evidence group, keyframes are selected from the set of keyframes corresponding to the remaining evidence packages, and transcribed fragments are selected from the set of transcribed fragments. Keyframe selection can be based on blackboard completeness, text recognition confidence, formula clarity, or image stability; transcribed fragment selection can be based on semantic completeness, keyword coverage, or assertion atom coverage. Subsequently, the selected keyframes, transcribed fragments, and their corresponding time ranges and content coordinates are combined to form the playback entries for that group. The playback entries from each group are then merged in chronological order to generate a minimal evidence playback set. This minimal evidence playback set is used to reduce repetitive playback content while retaining core evidence of the knowledge assertion formation process, thereby improving the efficiency of classroom knowledge playback.

[0054] The backtracking protocol is used to determine the upstream derivation path corresponding to the knowledge assertion in the argument graph and generate the corresponding minimal evidence replay set sequence. Its processing logic includes: Determine the starting object for backtracking. The starting object for backtracking is either the node in the argument graph corresponding to the selection instruction or the target argument graph node pointed to by the edge in the argument graph corresponding to the selection instruction. The selection instruction includes a knowledge assertion click instruction, a knowledge relation edge click instruction, or a knowledge replay trigger instruction.

[0055] Starting from the backtracking starting object, determine the set of upstream argument graph nodes along the edges of the requires type argument graph and the impies type argument graph that enter the argument graph node; where the edges of the requires type argument graph are used to represent the prerequisite dependency relationship, and the edges of the impies type argument graph are used to represent the derivation relationship. Therefore, by performing backtracking along the above two types of argument graph edges, the source knowledge assertion and derivation process corresponding to the current knowledge assertion can be gradually located.

[0056] The upstream argument graph node is added to the backtracking queue, and the current argument graph node is retrieved sequentially from the backtracking queue. For the current argument graph node, the set of evidence packets associated with it is read, and the minimum evidence replay set generation process is performed to output the minimum evidence replay set corresponding to the current argument graph node. The minimum evidence replay set generation process includes content coordinate grouping processing, time range overlap detection processing, redundancy reduction processing, keyframe filtering processing, and transcription fragment filtering processing.

[0057] When the current argument graph node has no required argument graph edge or implies argument graph edge that leads to it, or when the current argument graph node has already appeared in the backtracking sequence, the upstream expansion of the current argument graph node is terminated, and the next argument graph node in the backtracking queue is processed. The step of stopping the expansion of the already appeared node is to prevent circular references in the argument graph from causing infinite backtracking.

[0058] When the backtracking queue is empty, the backtracking is terminated, and the minimum evidence replay sets corresponding to each current argument graph node obtained during the backtracking process are combined into a minimum evidence replay set sequence according to the traversal order and output. The minimum evidence replay set sequence is used to display the knowledge formation process, derivation process and source of evidence in the classroom in the order of knowledge derivation.

[0059] Specifically, when a user clicks on a knowledge assertion node, the system can use that node as the starting point for backtracking and automatically search for the corresponding antecedent and derived knowledge assertions along the edges of the upstream require and implies type argument graphs. During each expansion, upstream expansion is only performed on argument graph nodes that have not yet appeared in the backtracking sequence, thus avoiding circular backtracking. Once all upstream nodes have been traversed, the system generates a minimal evidence playback set sequence arranged in the order of knowledge derivation, allowing users to review the formation, derivation, and sources of evidence for that knowledge point in the order it was presented in class.

[0060] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0061] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the protection scope of the present invention. All data acquisition actions in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located and with the authorization granted by the owner of the corresponding device.

Claims

1. A method for extracting and analyzing classroom knowledge points based on first-person perspective using AI glasses, characterized in that: Includes the following steps: Step S1: Obtain the first-person audio and video stream, and perform event segmentation, speaker category identification, and semantic pattern recognition on the first-person audio and video stream to obtain an event segment sequence; Step S2: Extract the visual anchor set and the speech anchor set based on the event segment sequence, and associate the visual anchor set and the speech anchor set with the corresponding event segments to obtain the anchor-enhanced event segment sequence; Step S3: Perform dual-anchor alignment on the visual anchor set and the speech anchor set to generate an aligned record sequence; Step S4: Construct candidate knowledge assertions based on the aligned record sequence, and perform evidence package construction, consistency verification and assertion merging processing on the candidate knowledge assertions to obtain a set of knowledge assertions; Step S5: Construct an argument graph based on the knowledge assertion set, and generate a minimum evidence replay set and a minimum evidence replay set sequence corresponding to the backtracking sequence based on the argument graph.

2. The method for extracting and analyzing classroom knowledge points based on first-person perspective using AI glasses as described in claim 1, characterized in that, Step S1 specifically includes: Step S11: Obtain the first-view audio and video stream, and determine the event boundaries in the first-view audio and video stream. The event boundaries include audio content change event boundaries, video content change event boundaries, and audio and video content change event boundaries. Step S12: Divide the first-view audio and video stream into a sequence of event segments arranged in the order of events and without overlap, based on the event boundaries, and record the corresponding time range for each event segment. Step S13: Segment the audio portion of each event segment to obtain a sequence of speech segments; Step S14: Determine the speaker category for each speech segment in the speech segment sequence, and determine the role label for the event segment based on the speaker category corresponding to each speech segment in the speech segment sequence. The role label includes teacher lecturing, student asking questions, classroom management, and idle chatter. Step S15: Perform semantic pattern recognition on the semantic analysis text corresponding to the speech fragment sequence of each event fragment, and determine the explanation status label for the event fragment based on the distribution results of each semantic pattern in the semantic analysis text. The explanation status label includes background definition, derivation and proof, example explanation, summary and emphasis, and question and answer clarification. Step S16: Associate and store the time range, role tag, and explanation status tag of the event segment to obtain the event segment sequence.

3. The method for extracting and analyzing classroom knowledge points based on first-person perspective using AI glasses as described in claim 2, characterized in that, Step S2 specifically includes: Step S21: For each event segment in the event segment sequence, determine the candidate content region from the video frame sequence corresponding to the event segment. The candidate content region includes text region, symbol region and chart region. Step S22: Perform text and symbol parsing on the candidate content area to obtain visual content units, and determine the content type corresponding to each visual content unit; Step S23: Determine the visual content unit as a visual anchor point according to the content type of the visual content unit, and form a visual anchor point set by combining all visual anchor points within the same event segment. The set of visual anchor points includes title anchor points, number anchor points, formula anchor points, icon anchor points, and whiteboard block anchor points; Step S24: For each visual anchor point in the visual anchor point set, record the corresponding video frame position and the time position of the video frame within the event segment time range, and store the video frame position, time position and the visual anchor point in association. Step S25: For each event segment in the event segment sequence, perform speech-to-text transcription on the audio part corresponding to the event segment to obtain anchor-transcribed text; Step S26: Identify semantic anchor trigger expressions in the anchor-transcribed text, determine the anchor-transcribed text segments corresponding to the semantic anchor trigger expressions as speech anchors, and form a speech anchor set from all speech anchors; The voice anchors include definition anchors, introduction anchors, emphasis anchors, error-prone anchors, and question anchors; Step S27: Record the time position of each voice anchor point in the set of voice anchor points within the time range of the event segment, and store the time position associated with the voice anchor point. Step S28: Associate and store the visual anchor point set and voice anchor point set corresponding to each event segment with the event segment to obtain the anchor point enhanced event segment sequence.

4. The method for extracting and analyzing classroom knowledge points based on first-person perspective using AI glasses as described in claim 3, characterized in that, Step S3 specifically includes: Step S31: For each anchor enhancement event segment in the anchor enhancement event segment sequence, read the time range of the anchor enhancement event segment, the set of visual anchors corresponding to the anchor enhancement event segment, and the set of speech anchors corresponding to the anchor enhancement event segment. Step S32: Determine the content coordinates for each visual anchor point in the set of visual anchor points. The content coordinates include the carrier identifier, page identifier, and region identifier. Step S33: Determine the speech time position for each speech anchor in the speech anchor set, and determine the visual time position for each visual anchor in the visual anchor set. Step S34: Within the same anchor point enhancement event segment, for each voice anchor point, select visual anchor points in the visual anchor point set that satisfy the temporal adjacency relationship, and establish an anchor point pairing relationship between the voice anchor point and the selected visual anchor points.

5. The method for extracting and analyzing classroom knowledge points based on first-person perspective using AI glasses as described in claim 4, characterized in that, Step S3 also includes: Step S35: Generate an alignment record for each anchor point pairing relationship. The alignment record includes the event segment identifier of the corresponding anchor point enhanced event segment, the time range determined by the voice time position of the voice anchor point and the visual time position of the visual anchor point, the content coordinates of the visual anchor point, the voice anchor point identifier, and the visual anchor point identifier. Step S36: All alignment records generated within the same anchor point enhancement event segment are arranged in chronological order to form an alignment record sequence, and the alignment record sequence is associated with and stored with the corresponding anchor point enhancement event segment to obtain the alignment record sequence.

6. The method for extracting and analyzing classroom knowledge points based on first-person perspective using AI glasses as described in claim 5, characterized in that, Step S4 specifically includes: Step S41: For each anchor point enhancement event segment in the anchor point enhancement event segment sequence, read the explanation status label corresponding to the anchor point enhancement event segment and the alignment record sequence corresponding to the anchor point enhancement event segment; Step S42: Based on the explanation status label, select the assertion template corresponding to the explanation status label from the preset assertion template set; The pre-set set of assertion templates includes templates for object properties and conditions, templates for steps, conclusions, templates for question types, templates for common mistakes, and templates for summarizing and merging. Step S43: For each alignment record in the alignment record sequence, read the time range, content coordinates, voice anchor identifier and visual anchor identifier in the alignment record, and combine the anchor transcribed text corresponding to the voice anchor identifier with the visual content unit corresponding to the visual anchor identifier based on the assertion template to generate candidate knowledge assertions. Step S44: Assign an assertion type to each candidate knowledge assertion, wherein the assertion type is determined by the assertion template, and establish an association between the candidate knowledge assertion and the alignment record that generated the candidate knowledge assertion. Step S45: All candidate knowledge assertions generated within the same anchor point augmentation event segment are arranged in chronological order to form a candidate knowledge assertion set, and the candidate knowledge assertion set is associated with the corresponding anchor point augmentation event segment. Step S46: For each candidate knowledge assertion in the candidate knowledge assertion set, construct an evidence package. The processing logic is as follows: Read the time range and content coordinates of the aligned record associated with the candidate knowledge assertion; Select video frames from the video frame sequence corresponding to the visual anchor point identifier that correspond to the visual content unit as the keyframe set; Extract the anchor point transcribed text segments corresponding to the speech anchor point identifiers from the anchor point transcribed text corresponding to the time range, and form a set of transcribed text segments from the anchor point transcribed texts; The time range, content coordinates, keyframe set, and transcribed fragment set are combined into an evidence package, and the evidence package is associated with the candidate knowledge assertion.

7. The method for extracting and analyzing classroom knowledge points based on first-person perspective using AI glasses as described in claim 6, characterized in that, Step S4 also includes: Step S47: Perform consistency verification for each candidate knowledge assertion to obtain consistency verification results. The consistency verification results include evidence coverage results, cross-modal consistency verification results, temporal consistency verification results, and counterfactual verification results. Step S48: Combine each candidate knowledge assertion with its corresponding evidence package and consistency verification result into a knowledge assertion record, generate a confidence level label for the knowledge assertion record, merge all knowledge assertion records, group the knowledge assertion records by content coordinates as the merging key, merge knowledge assertion records with the same assertion atom set into the same knowledge assertion within each group, and merge the evidence packages corresponding to the merged knowledge assertion records into an evidence package set. Step S49: Output the merged knowledge assertions as a knowledge assertion set, and establish an association between the knowledge assertion set and the corresponding content coordinates for storage.

8. The method for extracting and analyzing classroom knowledge points based on first-person perspective using AI glasses as described in claim 7, characterized in that, Step S5 specifically includes: Step S51: Construct an argument graph, read the knowledge assertion set, assign an argument graph node identifier to each knowledge assertion in the knowledge assertion set, and map each knowledge assertion to an argument graph node. Each node in the argument graph is associated with and stores the assertion type, assertion text, content coordinates, evidence package set, confidence level label, and assertion atom set of the knowledge assertion. Step S52: Establish dependency edges, derivation edges, applicable edges, and misconception edges between nodes in the argument graph, and establish association storage between each argument graph edge and the corresponding evidence package set.

9. The method for extracting and analyzing classroom knowledge points based on first-person perspective using AI glasses as described in claim 8, characterized in that, Step S5 also includes: Step S53: For each node or edge of the argument diagram, read the associated set of evidence packages, group the evidence package sets according to content coordinates, and perform redundancy reduction processing on evidence packages with the same content coordinates and overlapping time ranges. Within each group, select keyframes from the keyframe set from the evidence package corresponding to that group, and select transcribed segments from the transcribed segment set. Combine the keyframes, transcribed segments, and their corresponding time ranges and content coordinates into the playback entries of that group. Merge the playback entries of each group into the minimum evidence playback set and output it. Step S54: According to the backtracking protocol, determine the backtracking path corresponding to the node or edge of the argument graph in the argument graph, and form a backtracking sequence by combining the argument graph nodes and edges in the backtracking path in the traversal order. Read the evidence package set corresponding to the backtracking sequence, and output the minimum evidence replay set sequence corresponding to the backtracking sequence based on the evidence package set.

10. A first-person perspective classroom knowledge point extraction and analysis system based on AI glasses, which is applied in the first-person perspective classroom knowledge point extraction and analysis method based on AI glasses as described in any one of claims 1-9, characterized in that, It includes an event fragment construction module, an anchor point enhancement module, a double anchor alignment module, a knowledge assertion generation module, and an argument replay module; The event segment construction module is used to acquire first-person audio and video streams, and to perform event segmentation, speaker category recognition, and semantic pattern recognition on the first-person audio and video streams to obtain an event segment sequence; The anchor point enhancement module is used to extract a set of visual anchor points and a set of voice anchor points based on an event segment sequence, and associate the set of visual anchor points and the set of voice anchor points with the corresponding event segments to obtain an anchor point enhanced event segment sequence. The dual-anchor alignment module is used to perform dual-anchor alignment on the visual anchor point set and the voice anchor point set to generate an alignment record sequence. The knowledge assertion generation module is used to construct candidate knowledge assertions based on the aligned record sequence, and to perform evidence package construction, consistency verification and assertion merging processing on the candidate knowledge assertions to obtain a set of knowledge assertions; The argument replay module is used to construct an argument graph based on the knowledge assertion set, and generate a minimum evidence replay set and a minimum evidence replay set sequence corresponding to the backtracking sequence based on the argument graph.

Citation Information

Patent Citations

  • AI intelligent glasses

    CN121364777A