A method for aligning timing information data of a lecture framework

CN122451281BActive Publication Date: 2026-08-21SHENZHEN THINKING MUSIC CULTURE EDUCATION TECH DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610914418.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-08-21
Estimated Expiration
2046-06-24

AI Technical Summary

Technical Problem

[0004]本发明旨在至少在一定程度上解决现有技术中的技术问题之一,通过提出一种授课框架的时序信息数据对齐方法,用于解决现有的授课框架的时序信息数据对齐方法中,在授课时序对齐判定方面,缺少依托授课环节的文本特征做语义匹配的方法,造成仅依靠固定钟表时间划分数据区间,时段无法自适应授课时的节奏变化,导致多类时序数据环节错配且对齐可靠性下降的问题

Benefits of technology

[0015]本发明的有益效果:本申请首先基于对齐参考框架获取授课框架对应的所有授课环节以及每个授课环节对应的时序信息类型;然后基于对齐参考框架对所有授课环节进行分析,并基于分析结果获取每个授课环节的文本特征数据、行为特征数据以及时序限定区间,这样的好处在于,通过获取授课环节,并基于授课环节的时序信息类型获取文本特征数据、行为特征数据以及时序限定区间,能够得到授课框架中的每个授课环节对应的多类时序数据的特征,以便于在数据对齐时,基于授课环节的文本特征数据、行为特征数据以及时序限定区间,在时序信息数据中匹配对应的数据,从而实现更加灵活且符合授课节奏变化的时序信息对齐,避免因多类时序数据环节错配导致对齐后的数据不准确的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122451281B_ABST
    Figure CN122451281B_ABST
Patent Text Reader

Abstract

The application discloses a kind of alignment methods of timing information data of teaching framework, it is related to timing signal processing technical field, including: based on alignment reference framework obtains teaching link and timing information type;Text feature data, behavior feature data and timing limit interval of the teaching link are obtained;When link time domain is obtained based on data extraction;Timing information data alignment is carried out in teaching framework based on link time domain;The present application is used to solve the alignment method of timing information data of existing teaching framework, in teaching timing alignment determination aspect, lack the method of doing semantic matching relying on the text feature of teaching link, cause only rely on fixed clock time division data interval, time period cannot adapt to the rhythm change of teaching time, lead to the problems, such as multiple timing data link mismatch and alignment reliability decline.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of timing signal processing technology, specifically to a method for aligning timing information data in a teaching framework. Background Technology

[0002] The teaching framework is a structured logical system that supports a single lesson. It is a systematic design of teaching objectives, content, process, and organization. Its core function is to make the classroom logic clear and highlight the key points, avoiding confusion and inefficiency. The temporal information data alignment of the teaching framework is essentially a data processing operation that unifies and organizes course segments and teaching data according to the time dimension, ensuring the consistency between the classroom temporal logic and the data. It belongs to the data alignment application in the context of digital teaching.

[0003] Existing methods for aligning temporal information data in teaching frameworks typically involve acquiring the audio and video data to be processed and then matching it with preset timestamps to align the temporal information data within the teaching framework. While this improved method can effectively match data based on audio and video timestamps, it lacks a method for semantic matching based on the textual features of the teaching segments in determining the timing of the teaching process. This results in data intervals being divided solely by fixed clock times, which cannot adapt to changes in the rhythm of the teaching sessions, leading to mismatches in multiple types of temporal data segments and decreased alignment reliability. For example, patent application CN112423075A discloses audio and video... The solution involves processing timestamps, devices, electronic equipment, and storage media. This solution corrects the first or second timestamp based on the acquisition delay, enabling the acquisition of time-synchronized audio and video data. Improvements to other methods for aligning time-series information data in teaching frameworks typically focus on audio-visual synchronization and latency reduction. However, in determining the timing alignment of teaching sessions, there is still a lack of methods for semantic matching based on the text features of the teaching segments. This results in data intervals being divided solely by fixed clock times, which cannot adapt to changes in the rhythm of the teaching sessions, leading to mismatches in various timing data segments and decreased alignment reliability. Therefore, it is necessary to improve the existing methods for aligning time-series information data in teaching frameworks. Summary of the Invention

[0004] This invention aims to at least partially solve one of the technical problems in the prior art by proposing a method for aligning time-series information data in a teaching framework. This method addresses the issue that existing methods for aligning time-series information data in a teaching framework lack a semantic matching method based on the text features of the teaching segments. This results in the data intervals being divided solely by fixed clock times, which cannot adapt to changes in the rhythm of the teaching sessions, leading to mismatches in multiple types of time-series data segments and a decrease in alignment reliability.

[0005] To achieve the above objectives, this application provides a method for aligning time-series information data in a teaching framework, comprising the following steps: Obtain multiple teaching frameworks that have undergone time-series data alignment, and denote them as alignment reference frameworks; based on the alignment reference frameworks, obtain all teaching segments corresponding to the teaching frameworks and the time-series information type corresponding to each teaching segment; All teaching segments are analyzed based on the alignment reference framework, and text feature data, behavioral feature data, and time-series limited intervals are obtained based on the analysis results. The time-series limited intervals include short-domain intervals and long-domain intervals. When aligning time-series information data within the teaching framework, the time-series information data is extracted based on the text feature data, behavioral feature data, and time-series defined intervals of each teaching segment, and the segment time domain of all teaching segments is obtained based on the data extraction results. Based on the time domain of all teaching segments, time sequence information data is aligned within the teaching framework.

[0006] Furthermore, based on the alignment reference frame, the acquisition of all teaching segments corresponding to the teaching framework and the timing information types corresponding to each teaching segment includes: For any obtained alignment reference frame: the stage in the alignment reference frame where time sequence information data is inserted is recorded as a teaching stage; based on time from first to last, all teaching stages in the alignment reference frame are obtained sequentially. For any teaching segment: the data type contained in the time-series information data corresponding to the teaching segment is recorded as the time-series information type of the teaching segment. The data types in the time-series information data include continuous data types and discrete data types. Continuous data types include video time-series data and audio time-series data. Discrete data types include courseware interaction data, classroom interaction data, blackboard writing data, and teaching operation data. Sequentially obtain all teaching segments in all aligned reference frames, as well as the timing information type corresponding to each teaching segment.

[0007] Furthermore, based on the alignment reference framework, all teaching segments are analyzed, and based on the analysis results, text feature data, behavioral feature data, and time-series defined intervals for each teaching segment are obtained, including: For any teaching segment α obtained from the alignment reference frame: all video time-series data corresponding to teaching segment α in all alignment reference frames are recorded as video data to be analyzed; all audio time-series data corresponding to teaching segment α in all alignment reference frames are recorded as audio data to be analyzed; all data of discrete data type corresponding to teaching segment α in all alignment reference frames are recorded as analyzable behavioral data. Establish a time axis of length T, and denote it as the teaching time axis, where T is the teaching duration specified in the teaching framework; based on the start time and end time of all video data and audio data to be analyzed, mark the time domain occupied by all video data and audio data to be analyzed within the teaching time axis, and denote them as the video time domain and audio time domain respectively.

[0008] Furthermore, based on the alignment reference framework, all teaching segments are analyzed, and the text feature data, behavioral feature data, and time-series defined intervals for each teaching segment are obtained based on the analysis results. Based on the records of actionable analysis data in all alignment reference frames, the occurrence time of all actions in the actionable data is marked in the teaching time axis and recorded as action time points; when any action time point is outside all video time domains and audio time domains, the time interval obtained by merging the time intervals corresponding to all video time domains and audio time domains is recorded as the short domain interval of teaching segment α, and the interval formed by the minimum and maximum values ​​of the time values ​​corresponding to all action time points is recorded as the long domain interval. When all action time points are within all video and audio time domains, the time interval obtained by merging the time intervals corresponding to all video and audio time domains is denoted as the long domain interval of the teaching segment α, and the interval formed by the minimum and maximum values ​​of the time values ​​corresponding to all action time points is denoted as the short domain interval.

[0009] Furthermore, based on the alignment reference framework, all teaching segments are analyzed, and the text feature data, behavioral feature data, and time-series defined intervals for each teaching segment are obtained based on the analysis results. For any audio data to be analyzed corresponding to the teaching segment α: the audio data to be analyzed is converted into text, and the resulting text data is recorded as audio text; the audio text is segmented into words, and all the resulting knowledge point nouns are stored in a text set, where the text set is used to store words, and the knowledge point nouns are obtained from an educational knowledge base; Obtain the text set corresponding to all audio data to be analyzed in the teaching session α, and denote the union of all text sets as the audio text feature of the teaching session α.

[0010] Furthermore, based on the alignment reference framework, all teaching segments are analyzed, and the text feature data, behavioral feature data, and time-series defined intervals for each teaching segment are obtained based on the analysis results. For any video data to be analyzed corresponding to the teaching segment α: identify the sound source in the video data to be analyzed, and record the audio generated by the teacher and the audio generated by the student as teacher time-series audio and student time-series audio respectively; perform text processing on the teacher time-series audio and student time-series audio respectively, and record the resulting text data as teacher video text and student video text respectively; The teacher's video text and the student's video text were segmented into words respectively, and all the knowledge point nouns obtained were stored in text sets respectively.

[0011] Furthermore, based on the alignment reference framework, all teaching segments are analyzed, and the text feature data, behavioral feature data, and time-series defined intervals for each teaching segment are obtained based on the analysis results. Obtain the union of the text sets corresponding to the teacher video texts of all the video data to be analyzed, and denote it as the teacher text feature of teaching segment α; obtain the union of the text sets corresponding to the student video texts of all the video data to be analyzed, and denote it as the student text feature of teaching segment α. Audio text features, teacher text features, and student text features are denoted as the text feature data of teaching segment α.

[0012] Furthermore, based on the alignment reference framework, all teaching segments are analyzed, and the text feature data, behavioral feature data, and time-series defined intervals for each teaching segment are obtained based on the analysis results. For any analyzable behavioral data corresponding to teaching segment α: all teacher behaviors and student behaviors recorded in the analyzable behavioral data are recorded as segment behaviors of teaching segment α; Obtain all analyzable behavioral data for each stage of the lesson, and denote the set containing all stage behaviors as the behavioral characteristic data of the teaching stage α.

[0013] Furthermore, when aligning the temporal information data within the teaching framework, the temporal information data is extracted based on the textual feature data, behavioral feature data, and temporal limitation intervals of each teaching segment. Based on the data extraction results, the temporal domain of all teaching segments is obtained, including: When aligning time-series information data within the teaching framework, for any teaching segment α to be aligned: the long domain interval of teaching segment α is recorded as the filtering interval; all data in the time-series information data that are within the filtering interval are recorded as data to be filtered. For audio time-series data in the data to be screened: the audio time-series data is converted to text, and the resulting text is recorded as the audio text to be screened; when any word β1 in the audio text to be screened is the same as any word in the audio text features of the teaching segment α, the time when word β1 is generated in the audio time-series data is recorded as the audio recognition time. For video time series data in the data to be screened: identify the sound source in the video time series data, convert the audio generated by the teacher and the audio generated by the student into text, and record the resulting text as the teacher text to be screened and the student text to be screened, respectively. When any word β2 in the teacher's text to be screened is the same as any word in the teacher's text features of the teaching segment α, the time when word β2 is generated in the video time series data is recorded as the teacher recognition time. When any word β3 in the student text to be screened is the same as any word in the student text features of the teaching segment α, the time when word β3 is generated in the video time series data is recorded as the student recognition time.

[0014] Furthermore, when aligning time-series information data within the teaching framework, data extraction is performed on the time-series information data based on the text feature data, behavioral feature data, and time-series defined intervals of each teaching segment. The process of obtaining the segment time domain for all teaching segments based on the data extraction results also includes: For discrete data in the data to be screened: obtain all the behaviors recorded in the data and record them as discrete behaviors to be screened; when any discrete behavior to be screened is the same as any behavior in the behavioral feature data of the teaching segment α, record the occurrence event of the discrete behavior to be screened in the data to be screened as the discrete behavior time. Mark all audio recognition times, teacher recognition times, student recognition times, and discrete behavior times on the teaching timeline, and denote the closed interval formed by the minimum and maximum values ​​of all times as the information coverage interval. The union of the short domain interval and the information coverage interval of the teaching segment α is denoted as the segment time domain of the teaching segment α; the data in the time series information data that are in the segment time domain of the teaching segment α are all denoted as the data that the teaching segment α is aligned in the teaching framework. Obtain the time domain of all teaching segments.

[0015] The beneficial effects of this invention are as follows: First, this application obtains all teaching segments corresponding to the teaching framework and the time sequence information type corresponding to each teaching segment based on the alignment reference framework; then, it analyzes all teaching segments based on the alignment reference framework, and obtains the text feature data, behavioral feature data, and time sequence limitation interval for each teaching segment based on the analysis results. The advantage of this is that by obtaining the teaching segments and obtaining the text feature data, behavioral feature data, and time sequence limitation interval based on the time sequence information type of the teaching segments, it is possible to obtain the features of multiple types of time sequence data corresponding to each teaching segment in the teaching framework. This facilitates matching the corresponding data in the time sequence information data based on the text feature data, behavioral feature data, and time sequence limitation interval of the teaching segments during data alignment, thereby achieving more flexible time sequence information alignment that conforms to the changes in the teaching rhythm and avoiding the problem of inaccurate data after alignment due to mismatch of multiple types of time sequence data segments.

[0016] This application further aligns the time-series information data within the teaching framework by extracting data based on the textual feature data, behavioral feature data, and time-series defined intervals of each teaching segment. Based on the data extraction results, the time domain of all teaching segments is obtained. Finally, based on the time domains of all teaching segments, the time-series information data is aligned within the teaching framework. The advantage of this approach is that by obtaining the time domain of each teaching segment, the time domain corresponding to the data in the time-series information data that matches the teaching segment can be obtained. This avoids the problem of missing or misaligned data after alignment caused by dividing data intervals solely by fixed clock times. This ensures that after aligning the time-series information data within the teaching framework, the data corresponding to each teaching segment in the teaching framework is time-series information data generated by that teaching segment in the actual course, thereby improving the reliability of data alignment. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the steps of the method of the present invention; Figure 2 This is a schematic diagram illustrating the acquisition of the time domain of the components in this invention; Figure 3 This is a schematic diagram of the electronic device of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Example 1, please refer to Figure 1As shown, this application provides a method for aligning time-series information data in a teaching framework, including the following steps: Step S1: Obtain multiple teaching frameworks that have been aligned with time-series information data, and denot them as alignment reference frameworks; based on the alignment reference frameworks, obtain all teaching segments corresponding to the teaching frameworks and the time-series information type corresponding to each teaching segment; Step S1 includes: Step S101, for any obtained alignment reference frame: all segments in the alignment reference frame into which timing information data is inserted are recorded as teaching segments; based on time from first to last, all teaching segments in the alignment reference frame are obtained sequentially. In the specific implementation process, the teaching process may include classroom introduction, review of prior knowledge, teaching of new knowledge, explanation of examples, in-class exercises, classroom summary and homework assignment, etc. The specific teaching process can be set according to the teaching framework included in the actual data alignment. Step S102: For any teaching segment: the data type contained in the time-series information data corresponding to the teaching segment is recorded as the time-series information type of the teaching segment. The data type in the time-series information data includes continuous data type and discrete data type. Continuous data type includes video time-series data and audio time-series data. Discrete data type includes courseware interaction data, classroom interaction data, blackboard writing data and teaching operation data. In the specific implementation process, courseware interaction data may include: data corresponding to operations such as turning page in the courseware PPT, switching between pages, jumping to a specified page number, inserting external videos, and web page projection; classroom interaction data may include: data corresponding to operations such as answer submission, triggering a response, raising hands to sign in, voting and reporting, and quiz answer records; blackboard writing data may include: data corresponding to operations such as putting down a pen, lifting a pen, erasing, pen stroke coordinate trajectory, and marking the start and end times of each stroke; teaching operation data may include: data corresponding to operations such as starting and stopping recording, pausing teaching, resuming teaching, switching projection, and starting and stopping supporting experimental equipment; discrete data types are data recorded when behaviors that randomly occur in the classroom are recorded, and can be more specifically set according to the behaviors that will occur in the actual teaching process; Step S103: Sequentially obtain all teaching segments in all alignment reference frames, as well as the timing information type corresponding to each teaching segment.

[0020] Step S2: Analyze all teaching segments based on the alignment reference frame, and obtain the text feature data, behavioral feature data and time-series limited intervals for each teaching segment based on the analysis results. The time-series limited intervals include short-domain intervals and long-domain intervals. Step S2 includes: Step S201, for any teaching segment α obtained from the alignment reference frame: all video time-series data corresponding to teaching segment α in all alignment reference frames are recorded as video data to be analyzed; all audio time-series data corresponding to teaching segment α in all alignment reference frames are recorded as audio data to be analyzed; all data with discrete data type corresponding to teaching segment α in all alignment reference frames are recorded as analyzable behavioral data. Step S202: Establish a time axis of length T and denot it as the teaching time axis, where T is the teaching duration specified in the teaching framework; based on the start and end times of all video data and audio data to be analyzed, mark the time domain occupied by all video data and audio data to be analyzed within the teaching time axis and denot them as the video time domain and audio time domain, respectively. In the data analysis of this embodiment, for example, if the teaching time specified in the teaching framework is 40 minutes during a data analysis, then the value of T can be set to 40 minutes, and a teaching time axis with a length of 40 minutes can be established.

[0021] Step S2 also includes: Step S203, based on the records of the actionable analysis data in all alignment reference frames, marking the occurrence time of all actions in the actionable data in the teaching time axis and recording them as action time points; when any action time point is outside all video time domains and audio time domains, the time interval obtained by merging the time intervals corresponding to all video time domains and audio time domains is recorded as the short domain interval of the teaching segment α, and the interval formed by the minimum and maximum values ​​of the time values ​​corresponding to all action time points is recorded as the long domain interval; In the specific implementation process, by obtaining short-domain intervals and long-domain intervals, the time domain length corresponding to the teaching segment can be limited, so as to avoid the time domain corresponding to the teaching segment being infinitely magnified or infinitely shrunk in subsequent analysis, thereby affecting the accurate alignment of time series information data. Step S204: When all action time points are within all video time domains and audio time domains, the time interval obtained by merging the time intervals corresponding to all video time domains and audio time domains is recorded as the long domain interval of the teaching segment α, and the interval formed by the minimum and maximum values ​​of the time values ​​corresponding to all action time points is recorded as the short domain interval. In the data analysis of the embodiment, for example, in the data analysis of the teaching session "new knowledge teaching", the time domain after merging all video time domains is [15:00, 30:00], and the time domain after merging all audio time domains is [16:00, 31:00]. Among the time values ​​corresponding to all behavior time points, the minimum and maximum values ​​are 15:30 and 30:30, respectively. Through analysis, it can be found that all behavior time points are within all video time domains and audio time domains. Therefore, the long domain interval is [15:00, 31:00], and the short domain interval is [15:30, 30:30].

[0022] Step S2 also includes: Step S205, for any audio data to be analyzed corresponding to the teaching segment α: the audio data to be analyzed is converted into text, and the resulting text data is recorded as audio text; the audio text is segmented into words, and all the resulting knowledge point nouns are stored in a text set, wherein the text set is used to store words, and the knowledge point nouns are obtained from an educational knowledge base; Step S206: Obtain the text set corresponding to all audio data to be analyzed in the teaching session α, and denote the union of all text sets as the audio text feature of the teaching session α.

[0023] Step S2 also includes: Step S207, for any video data to be analyzed corresponding to the teaching segment α: identify the sound source in the video data to be analyzed, and record the audio generated by the teacher and the audio generated by the student as teacher time-series audio and student time-series audio respectively; perform text processing on the teacher time-series audio and student time-series audio respectively, and record the resulting text data as teacher video text and student video text respectively; In the specific implementation process, AI can be used to analyze the location of the sound in the video data to be analyzed, and match the location of the sound with the location of the teacher and the student. When the location of the sound coincides with the location of the teacher, it can be determined that the audio was produced by the teacher; when the location of the sound coincides with the location of the student, it can be determined that the audio was produced by the student. Step S208: Perform word segmentation on the teacher's video text and the student's video text respectively, and store all the knowledge point nouns obtained into the text set respectively.

[0024] Step S2 further includes: Step S209, obtaining the union of the text sets corresponding to the teacher video texts of all the video data to be analyzed, and denoting it as the teacher text feature of teaching segment α; obtaining the union of the text sets corresponding to the student video texts of all the video data to be analyzed, and denoting it as the student text feature of teaching segment α; Step S210: Record the audio text features, teacher text features, and student text features as the text feature data of the teaching segment α.

[0025] Step S2 also includes: Step S211, for any analyzable behavioral data corresponding to teaching segment α: all teacher behaviors and student behaviors recorded in the analyzable behavioral data are recorded as segment behaviors of teaching segment α; In the data analysis of this embodiment, for example, in a data analysis, all the behaviors of the teaching segment α are as follows: answer submission, answer trigger, raise hand to sign in, vote report, write down, lift pen, erase operation, record start and stop, teaching pause, resume class and screen projection switch. Therefore, if the teaching segment and behavior record are matched in the future, if the above behaviors exist, it can be determined that it is within the teaching segment α. Step S212: Obtain all segment behaviors that can be analyzed, and denote the set containing all segment behaviors as the behavioral feature data of teaching segment α.

[0026] Step S3: When aligning the time-series information data in the teaching framework, extract the time-series information data based on the text feature data, behavioral feature data, and time-series limited interval of each teaching segment, and obtain the segment time domain of all teaching segments based on the data extraction results. Based on the time domain of all teaching segments, time sequence information data is aligned within the teaching framework; Step S3 includes: Step S301, when aligning the time series information data in the teaching framework, for any teaching segment α to be aligned: the long domain interval of the teaching segment α is recorded as the filtering interval; all data in the time series information data that are within the filtering interval are recorded as data to be filtered. Step S302: For audio time series data in the data to be screened: the audio time series data is converted to text and the resulting text is recorded as the audio text to be screened; when any word β1 in the audio text to be screened is the same as any word in the audio text features of the teaching segment α, the time when word β1 is generated in the audio time series data is recorded as the audio recognition time. Step S303, for the video time series data in the data to be screened: identify the sound source in the video time series data, convert the audio generated by the teacher and the audio generated by the student into text, and record the resulting text as the teacher text to be screened and the student text to be screened, respectively. Step S304: When any word β2 in the teacher's text to be screened is the same as any word in the teacher's text features of the teaching segment α, the time when word β2 is generated in the video time series data is recorded as the teacher recognition time. Step S305: When any word β3 in the student text to be screened is the same as any word in the student text features of the teaching segment α, the time when word β3 is generated in the video time series data is recorded as the student recognition time.

[0027] Step S3 also includes: Step S306, for data of discrete data type in the data to be screened: obtain all behaviors recorded in the data and record them as discrete behaviors to be screened; when any discrete behavior to be screened is the same as any behavior in the behavioral feature data of the teaching segment α, record the occurrence event of the discrete behavior to be screened in the data to be screened as the discrete behavior time. Step S307: Mark all audio recognition times, teacher recognition times, student recognition times, and discrete behavior times on the teaching timeline, and record the closed interval formed by the minimum and maximum values ​​of all times as the information coverage interval; In the data analysis of this embodiment, for example, in the data analysis of the "new knowledge instruction" teaching segment, by obtaining the audio recognition time, teacher recognition time, student recognition time, and discrete behavior time in the data to be screened, the minimum and maximum values ​​of all the times are obtained, and the interval formed by the minimum and maximum values ​​is [15:00, 30:00], that is, the information coverage interval is [15:00, 30:00]. Through the above analysis, the short domain interval of "new knowledge instruction" is obtained as [15:30, 30:30]. Therefore, the time domain of the "new knowledge instruction" segment is [15:00, 30:30]. [15:00, 30:30] can be used to record the data in the time series information data at [15:00, 30:30] as the data aligned in the teaching framework for "New Knowledge Presentation". If the above method is not used, and data alignment is performed only by a fixed time domain or a fixed clock time, such as using a fixed time domain [15:00, 30:00] for data alignment, the data in the time domain [30:00, 30:30] in the time series information data will be missed during data alignment, resulting in data loss after data alignment and misalignment with data in other teaching segments. Step S308: The union of the short domain interval and the information coverage interval of the teaching segment α is denoted as the segment time domain of the teaching segment α; the data in the time series information data that are in the segment time domain of the teaching segment α are all denoted as the data that the teaching segment α is aligned in the teaching framework. Step S309: Obtain the time domain of all teaching segments.

[0028] Example 2, please refer to Figure 3 As shown, Figure 3A schematic diagram of an electronic device is provided, which may include a processor, a communication interface, a memory, and a communication bus. The processor, communication interface, and memory communicate with each other via the communication bus. The memory stores computer-readable instructions, and the processor can call these instructions. When the processor executes a computer-readable instruction, it performs steps as described in a method for aligning timing information data in a teaching framework, achieving the following functions: First, multiple teaching frameworks that have undergone timing information data alignment are acquired and designated as alignment reference frameworks. Based on the alignment reference framework, all teaching segments corresponding to the teaching framework and the timing information type corresponding to each teaching segment are acquired. Then, all teaching segments are analyzed based on the alignment reference framework, and the text feature data, behavioral feature data, and timing limit intervals for each teaching segment are acquired based on the analysis results. When aligning timing information data within the teaching framework, the timing information data is extracted based on the text feature data, behavioral feature data, and timing limit intervals for each teaching segment, and the segment time domain of all teaching segments is obtained based on the data extraction results. Finally, timing information data alignment is performed within the teaching framework based on the segment time domains of all teaching segments.

[0029] Furthermore, when the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0030] Example 3: This application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute a method for aligning the timing information data of a teaching framework provided by the above methods. The method includes: first, acquiring multiple teaching frameworks that have undergone timing information data alignment, and denoting them as alignment reference frameworks; acquiring all teaching segments corresponding to the teaching framework and the timing information type corresponding to each teaching segment based on the alignment reference framework; then analyzing all teaching segments based on the alignment reference framework, and acquiring text feature data, behavioral feature data, and timing limit intervals for each teaching segment based on the analysis results; when aligning the timing information data in the teaching framework, extracting the timing information data based on the text feature data, behavioral feature data, and timing limit intervals for each teaching segment, and acquiring the segment time domain of all teaching segments based on the data extraction results; finally, aligning the timing information data in the teaching framework based on the segment time domains of all teaching segments.

[0031] Example 4: This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it performs the steps of the above-described method for aligning time-series information data in a teaching framework to achieve the following functions: First, it acquires multiple teaching frameworks that have undergone time-series information data alignment, and denotes them as alignment reference frameworks; based on the alignment reference frameworks, it acquires all teaching segments corresponding to the teaching frameworks and the time-series information type corresponding to each teaching segment; then, it analyzes all teaching segments based on the alignment reference frameworks, and acquires the text feature data, behavioral feature data, and time-series limitation intervals of each teaching segment based on the analysis results; when aligning time-series information data in the teaching framework, it extracts time-series information data based on the text feature data, behavioral feature data, and time-series limitation intervals of each teaching segment, and acquires the segment time domain of all teaching segments based on the data extraction results; finally, it performs time-series information data alignment in the teaching framework based on the segment time domains of all teaching segments.

[0032] Based on the above description of the embodiments, the embodiments of the present invention can be provided as methods, systems, or computer program products. Based on this understanding, the above technical solutions, in essence or in terms of their contribution to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or certain parts of the embodiments.

[0033] In the embodiments provided in this application, it should be understood that the disclosed system or method can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules or units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple modules or units may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interfaces. The indirect coupling or communication connection between systems, modules, and units may be electrical, mechanical, or other forms.

[0034] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for aligning temporal information data in a teaching framework, characterized in that, Includes the following steps: Obtain multiple teaching frames that have been aligned with time-series information data, and denote them as alignment reference frames; Based on the alignment reference frame, obtain all teaching segments corresponding to the teaching framework and the timing information type corresponding to each teaching segment; All teaching segments are analyzed based on the alignment reference framework, and text feature data, behavioral feature data, and time-series limited intervals are obtained based on the analysis results. The time-series limited intervals include short-domain intervals and long-domain intervals. When aligning temporal information data within the teaching framework, data extraction is performed on the temporal information data based on the text feature data, behavioral feature data, and temporal limitation intervals of each teaching segment. The segment temporal domain of all teaching segments is then obtained based on the data extraction results. This includes: when aligning temporal information data within the teaching framework, for any teaching segment α to be aligned: the long domain interval of teaching segment α is recorded as the filtering interval; all data within the filtering interval in the temporal information data is recorded as data to be filtered; for audio temporal data in the data to be filtered: the audio temporal data is transcribed into text. The text is processed and recorded as the audio text to be screened. When any word β1 in the audio text to be screened is the same as any word in the audio text features of the teaching segment α, the time when word β1 is generated in the audio time series data is recorded as the audio recognition time. For the video time series data in the data to be screened: the sound source in the video time series data is identified, and the audio generated by the teacher and the audio generated by the student are converted into text, and the resulting text is recorded as the teacher text to be screened and the student text to be screened, respectively. When any word β2 in the teacher text to be screened is the same as any word in the audio text features of the teaching segment α, the time when word β1 is generated in the audio time series data is recorded as the audio recognition time. When any word in the teacher's text features of teaching segment α is the same as any word in the video time series data, the time when word β2 is generated is recorded as the teacher recognition time; when any word β3 in the student's text to be screened is the same as any word in the student's text features of teaching segment α, the time when word β3 is generated is recorded as the student recognition time; for data of discrete data type in the data to be screened: obtain all behaviors recorded in the data and record them as discrete behaviors to be screened; when any discrete behavior to be screened is the same as any word in the behavioral feature data of teaching segment α, the time when word β3 is generated is recorded as the student recognition time; To ensure consistency, the occurrence events of the discrete behaviors to be screened recorded in the data to be screened are denoted as discrete behavior times; all audio recognition times, teacher recognition times, student recognition times, and discrete behavior times are marked on the teaching timeline, and the closed interval formed by the minimum and maximum values ​​of all times is denoted as the information coverage interval; the union of the short domain interval of teaching segment α and the information coverage interval is denoted as the segment time domain of teaching segment α; all data in the time series information data that are in the segment time domain of teaching segment α are denoted as the data aligned by teaching segment α in the teaching framework; the segment time domains of all teaching segments are obtained; Based on the time domain of all teaching segments, time sequence information data is aligned within the teaching framework.

2. The method for aligning temporal information data in a teaching framework according to claim 1, characterized in that, Based on the alignment reference frame, all teaching segments corresponding to the teaching framework and the timing information types corresponding to each teaching segment are obtained, including: For any obtained alignment reference frame: the stage in the alignment reference frame where time sequence information data is inserted is recorded as a teaching stage; based on time from first to last, all teaching stages in the alignment reference frame are obtained sequentially. For any teaching segment: the data type contained in the time-series information data corresponding to the teaching segment is recorded as the time-series information type of the teaching segment. The data types in the time-series information data include continuous data types and discrete data types. Continuous data types include video time-series data and audio time-series data. Discrete data types include courseware interaction data, classroom interaction data, blackboard writing data, and teaching operation data. Sequentially obtain all teaching segments in all aligned reference frames, as well as the timing information type corresponding to each teaching segment.

3. The method for aligning temporal information data in a teaching framework according to claim 2, characterized in that, All teaching segments were analyzed based on an alignment reference framework, and the text feature data, behavioral feature data, and time-series defined intervals for each teaching segment were obtained based on the analysis results: For any teaching segment α obtained from the alignment reference frame: all video time-series data corresponding to teaching segment α in all alignment reference frames are recorded as video data to be analyzed; all audio time-series data corresponding to teaching segment α in all alignment reference frames are recorded as audio data to be analyzed; all data of discrete data type corresponding to teaching segment α in all alignment reference frames are recorded as analyzable behavioral data. Establish a time axis of length T, and denote it as the teaching time axis, where T is the teaching duration specified in the teaching framework; based on the start time and end time of all video data and audio data to be analyzed, mark the time domain occupied by all video data and audio data to be analyzed within the teaching time axis, and denote them as the video time domain and audio time domain respectively.

4. The method for aligning temporal information data in a teaching framework according to claim 3, characterized in that, Based on the alignment reference framework, all teaching segments are analyzed, and the text feature data, behavioral feature data, and time-series defined intervals for each teaching segment are obtained based on the analysis results. Based on the records of actionable analysis data in all alignment reference frames, the occurrence time of all actions in the actionable data is marked in the teaching time axis and recorded as action time points; when any action time point is outside all video time domains and audio time domains, the time interval obtained by merging the time intervals corresponding to all video time domains and audio time domains is recorded as the short domain interval of teaching segment α, and the interval formed by the minimum and maximum values ​​of the time values ​​corresponding to all action time points is recorded as the long domain interval. When all action time points are within all video and audio time domains, the time interval obtained by merging the time intervals corresponding to all video and audio time domains is denoted as the long domain interval of the teaching segment α, and the interval formed by the minimum and maximum values ​​of the time values ​​corresponding to all action time points is denoted as the short domain interval.

5. The method for aligning temporal information data in a teaching framework according to claim 4, characterized in that, Based on the alignment reference framework, all teaching segments are analyzed, and the text feature data, behavioral feature data, and time-series defined intervals for each teaching segment are obtained based on the analysis results. For any audio data to be analyzed corresponding to the teaching segment α: the audio data to be analyzed is converted into text, and the resulting text data is recorded as audio text; the audio text is segmented into words, and all the resulting knowledge point nouns are stored in a text set, where the text set is used to store words, and the knowledge point nouns are obtained from an educational knowledge base; Obtain the text set corresponding to all audio data to be analyzed in the teaching session α, and denote the union of all text sets as the audio text feature of the teaching session α.

6. The method for aligning temporal information data in a teaching framework according to claim 5, characterized in that, Based on the alignment reference framework, all teaching segments are analyzed, and the text feature data, behavioral feature data, and time-series defined intervals for each teaching segment are obtained based on the analysis results. For any video data to be analyzed corresponding to the teaching segment α: identify the sound source in the video data to be analyzed, and record the audio generated by the teacher and the audio generated by the student as teacher time-series audio and student time-series audio respectively; perform text processing on the teacher time-series audio and student time-series audio respectively, and record the resulting text data as teacher video text and student video text respectively; The teacher's video text and the student's video text were segmented into words respectively, and all the knowledge point nouns obtained were stored in text sets respectively.

7. The method for aligning temporal information data in a teaching framework according to claim 6, characterized in that, Based on the alignment reference framework, all teaching segments are analyzed, and the text feature data, behavioral feature data, and time-series defined intervals for each teaching segment are obtained based on the analysis results. Obtain the union of the text sets corresponding to the teacher video texts of all the video data to be analyzed, and denote it as the teacher text feature of teaching segment α; obtain the union of the text sets corresponding to the student video texts of all the video data to be analyzed, and denote it as the student text feature of teaching segment α. Audio text features, teacher text features, and student text features are denoted as the text feature data of teaching segment α.

8. The method for aligning temporal information data in a teaching framework according to claim 7, characterized in that, Based on the alignment reference framework, all teaching segments are analyzed, and the text feature data, behavioral feature data, and time-series defined intervals for each teaching segment are obtained based on the analysis results. For any analyzable behavioral data corresponding to teaching segment α: all teacher behaviors and student behaviors recorded in the analyzable behavioral data are recorded as segment behaviors of teaching segment α. Obtain all analyzable behavioral data for each stage of the lesson, and denote the set containing all stage behaviors as the behavioral characteristic data of the teaching stage α.

Citation Information

Patent Citations

  • Audio and video timestamp processing method and device, electronic equipment and storage medium

    CN112423075A

  • Smart classroom management system based on NB-IoT (Narrow Band Internet of Things)

    CN116385221A

  • Methods and systems for processing subtitles and performing audio processing based on subtitles

    WO2026122594A1