A Smart Teaching and Research Activity Management Method and System Supporting Multimodal Input

By establishing a pointer association mechanism between audio streams, video streams, and comment texts in teaching and research activities, chained annotations are constructed and composite teaching activity curves and rhythm analysis maps are generated. This solves the problem of fragmented multimodal data and improves the comprehensive utilization efficiency and analytical depth of teaching and research activities.

CN121936735BActive Publication Date: 2026-05-26MINXI VOCATIONAL & TECHN COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MINXI VOCATIONAL & TECHN COLLEGE
Filing Date
2026-03-26
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

The lack of a unified correlation mechanism for multimodal data in the current teaching and research activity management leads to the inability to accurately correspond comment texts with teaching scene fragments, resulting in low comprehensive utilization. Furthermore, the analysis can only provide single-dimensional data statistics, making it difficult to reveal changes in teaching rhythm and stage characteristics.

Method used

By establishing a pointer association mechanism among audio streams, video streams, and comment text, chained annotations are constructed to generate composite teaching activity curves and rhythm analysis maps, thereby achieving deep fusion and temporal segmentation analysis of multimodal data.

Benefits of technology

It enables efficient and convenient retrospective analysis of multimodal information, and the generated comprehensive analysis report not only presents the fluctuations in teaching rhythm but also includes qualitative feedback, meeting the process evaluation needs of smart teaching research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121936735B_ABST
    Figure CN121936735B_ABST
Patent Text Reader

Abstract

This invention relates to the field of educational management technology, and discloses a method and system for managing intelligent teaching and research activities that supports multimodal input. The method includes: collecting audio streams, video streams, and comment texts of teaching and research activities; writing the comment texts into pointer fields pointing to speech segments and behavior tag indices to obtain chained annotations; establishing and encapsulating mutual pointers between the audio stream, video stream, and chained annotations based on time identifiers to obtain a teaching and research dataset; constructing a composite teaching activity curve with the speech text time period as the horizontal axis and the composite value of behavior frequency and speech transition number as the vertical axis; using the peaks, troughs, and slope change points of the curve as rhythm feature points, and matching them with a behavioral semantic database to obtain a rhythm analysis graph; using the graph as the time axis skeleton, embedding the chained annotations into the target time interval according to the pointer direction to obtain a comprehensive analysis report; this invention can improve the efficiency of intelligent teaching and research activity management that supports multimodal input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of educational management technology, and in particular to a smart teaching and research activity management method and system that supports multimodal input. Background Technology

[0002] In existing teaching and research activity management technologies, audio streams, video streams, and comment texts are typically collected and stored as independent modal data. Due to the lack of a unified data association mechanism, a precise correspondence cannot be established between comment texts and specific teaching scene segments. This forces users to manually switch and search between audio, video, and text when reviewing teaching and research activities. This fragmented data organization results in low utilization of multimodal information and fails to intuitively present the specific teaching behavior corresponding to the textual feedback.

[0003] Furthermore, existing technologies, when analyzing the process of teaching and research activities, often only provide single-dimensional data statistics, such as the number of times a speaker speaks or the frequency of a behavior, making it difficult to reveal the implicit rhythmic changes and stage characteristics of teaching activities. Due to the lack of in-depth mining and semantic interpretation of the temporal distribution of teaching behaviors, the generated teaching and research reports usually remain at the level of data listing, failing to provide teachers with effective feedback on the control of teaching rhythm and stage division, and thus failing to meet the actual needs of intelligent teaching and research for process evaluation and precise reflection. Summary of the Invention

[0004] This invention provides a smart teaching and research activity management method and system that supports multimodal input, the main purpose of which is to solve the problem of low efficiency in smart teaching and research activity management that supports multimodal input.

[0005] To achieve the above objectives, this invention provides a smart teaching and research activity management method that supports multimodal input, comprising:

[0006] Based on the user time stamps of the teaching and research activities, the audio stream, video stream, and comment text of the teaching and research activities are collected respectively.

[0007] The comment text is written into a pointer field pointing to the speech segment index in the audio stream and the behavior tag index in the video stream to obtain the chained comments of the teaching and research activity;

[0008] Based on the time marker of the teaching and research activity, pointers are established between the audio stream, the video stream and the chained annotations, and the data streams after the pointers are established are encapsulated to obtain the teaching and research dataset of the teaching and research activity.

[0009] Using the time period of the speech text in the teaching and research dataset as the horizontal axis and the composite value generated by fusing the behavior frequency and the number of speech transitions in the teaching and research dataset within the time period as the vertical axis, a composite teaching activity curve of the teaching and research activity is constructed.

[0010] The peaks, troughs, and slope change points of the composite teaching activity curve are used as rhythm feature points. The behavioral labels corresponding to the rhythm feature points are matched with the behavioral semantic database of the teaching and research activities to obtain the rhythm analysis map of the teaching and research activities.

[0011] Using the rhythm analysis graph as the time axis framework and segmentation basis, the chain annotations are embedded into the target time interval of the rhythm analysis graph according to the pointer direction to obtain a comprehensive analysis report of the teaching and research activities.

[0012] In a preferred embodiment, the user time stamp based on the teaching and research activity collects the audio stream, video stream, and comment text of the teaching and research activity, including:

[0013] The creation time and user identifier of the teaching and research activity are used as channel association seeds and written into the header information field of the acquisition link corresponding to the teaching and research activity to obtain the audio acquisition channel, video acquisition channel and text acquisition channel of the teaching and research activity.

[0014] The audio stream, video stream, and comment text of the teaching and research activities are collected through the acquisition channel, the video acquisition channel, and the text acquisition channel.

[0015] In a preferred embodiment, the step of writing the comment text into pointer fields pointing to the speech segment index in the audio stream and the behavior tag index in the video stream to obtain the chained annotations of the teaching and research activity includes:

[0016] Based on the global time anchor point of the comment text, in the audio stream's speech segments and the video stream's classroom event sequence, respectively, the target speech segments and target behavior tags with start times less than or equal to the global time anchor point and end times greater than or equal to the global time anchor point are searched to obtain the speech segment index and behavior tag index of the teaching and research activity.

[0017] Write the speech segment index and the behavior tag index into the speech pointer field and behavior pointer field of the comment text data structure to obtain the chained annotation of the teaching and research activity.

[0018] In a preferred embodiment, based on the time identifier of the teaching and research activity, pointers are established between the audio stream, the video stream, and the chained annotations, and the data streams after the pointers are established are encapsulated to obtain the teaching and research dataset of the teaching and research activity, including:

[0019] Based on the time interval of the speech segment in the audio stream, the behavior labels and corresponding behavior label indices located within the time interval are traversed in the classroom event sequence of the video stream, and a downlink pointer pointing to the behavior label index is created in the data structure of the speech segment.

[0020] Based on the time interval of the behavior label, find the chained comments whose synchronization time point is within the time interval and their corresponding comment identifiers in the chained comments, and create an uplink pointer to the comment identifier in the data structure of the behavior label;

[0021] A source backtracking pointer is established pointing to the corresponding speech segment in the audio stream based on the speech segment index in the speech pointer field, and a target mapping pointer is established pointing to the corresponding behavior label in the classroom event sequence based on the behavior label index in the behavior pointer field.

[0022] The voice text, the classroom event sequence, and the chained annotations are stored together. The downlink pointer, the uplink pointer, the source backtracking pointer, and the target mapping pointer are written into the storage structure of the teaching and research activity as the association between the voice text, the classroom event sequence, and the chained annotations, thus obtaining the teaching and research dataset of the teaching and research activity.

[0023] In a preferred embodiment, the step of constructing a composite teaching activity curve for the teaching and research activity, using the time period of the speech text in the teaching and research dataset as the horizontal axis and the composite value generated by fusing the corresponding values ​​of the behavior frequency and the number of speech transitions in the teaching and research dataset within the time period as the vertical axis, includes:

[0024] The total duration of the audio text in the teaching and research dataset is divided into continuous time periods, and the total number of behavior labels whose start time is located in the continuous time period in the classroom event sequence in the teaching and research dataset is taken as the behavior frequency of the continuous time period.

[0025] The number of times the speaker identifier of the speech text in the teaching and research dataset changes between adjacent speech segments is taken as the number of speech transitions in the continuous time period.

[0026] Using the continuous time period as the horizontal axis value and the frequency of the row as the left vertical axis value, the left data point corresponding to the continuous time period is obtained;

[0027] Using the continuous time period as the horizontal axis value and the number of speech transitions as the right vertical axis value, the right data point corresponding to the continuous time period is obtained;

[0028] By connecting all the left and right data points in chronological order, the frequency curve and interaction conversion curve of the teaching and research activities are obtained.

[0029] The behavior frequency curve and the interaction conversion curve are integrated into a composite teaching activity curve for the teaching and research activities.

[0030] In a preferred embodiment, integrating the behavior frequency curve and the interaction conversion curve into a composite teaching activity curve for the teaching and research activity includes:

[0031] The frequency of behaviors in each time period of the behavior frequency curve and the number of speech transitions in each time period of the interaction transition curve are normalized to obtain the frequency scale value and transition scale value of the corresponding time period in the teaching and research activity.

[0032] The Euclidean distance between the frequency scale value and the transformation scale value is calculated based on the same time period of the teaching and research activities to obtain the composite teaching index of the teaching and research activities.

[0033] Using the continuous time period of the teaching and research activities as the horizontal axis and the composite teaching indicators as the vertical axis, the data points are connected sequentially in chronological order to obtain the composite teaching activity curve of the teaching and research activities.

[0034] In a preferred embodiment, the step of using the peaks, troughs, and slope change points of the composite teaching activity curve as rhythm feature points, and matching the behavioral labels corresponding to the rhythm feature points with the behavioral semantic database of the teaching and research activity to obtain the rhythm analysis map of the teaching and research activity, includes:

[0035] The local maxima and local minima of the composite teaching activity curve are taken as peaks and troughs, the points where the sign of the first derivative changes in the composite teaching activity curve are taken as slope change points, and the peaks, troughs and slope change points are all taken as rhythm feature points.

[0036] Taking the time position of the rhythm feature point on the time axis of the audio text of the teaching and research activity as the center, extend forward to the position of the preceding rhythm feature point and backward to the position of the subsequent rhythm feature point. If there is no rhythm feature point before the time position, extend forward to the starting point of the audio text in the teaching and research activity. If there is no rhythm feature point after the time position, extend backward to the ending point of the audio text in the teaching and research activity to obtain the feature point time window of the teaching and research activity.

[0037] Extract behavioral labels whose starting time is within the time window of the feature point from the classroom event sequence of the teaching and research activities to obtain a set of behavioral labels corresponding to the rhythm feature point;

[0038] Calculate the matching degree between the behavior tag set and the teaching stage template in the behavior semantic library of teaching and research activities, and take the stage name of the teaching stage template corresponding to the maximum value of the matching degree as the teaching semantic tag of the rhythm feature point;

[0039] The teaching semantic tags are labeled to the corresponding rhythm feature points in the composite teaching activity curve to obtain the rhythm analysis map of the teaching and research activity.

[0040] In a preferred embodiment, the formula for calculating the matching degree includes:

[0041]

[0042] in, For the set of behavior labels and the first The matching degree of the template for each teaching stage. This serves as an index for the template of the aforementioned teaching stage. For the behavior type in the behavior tag set and the first The total number of behavior tags that match the standard behavior types in each teaching stage template. The sequence number of the behavior tag in the behavior tag set that matches the teaching stage template. For the first behavior tag set The behavior type corresponding to each matching behavior tag in the first... The preset weight values ​​corresponding to each teaching stage template For the first behavior tag set The actual number of times the behavior type corresponding to each matching behavior label appears within the time window of the feature point. It is a natural constant. The time decay coefficient, For the first behavior tag set The timestamp of each matching behavior tag. The time position of the rhythmic feature point. The total duration of the time window for the feature points. For the first The total number of all standard behavior types in each teaching stage template This serves as an index for the standard behavior types in the teaching stage template. To indicate the first In the first teaching stage template Preset weight values ​​for each standard behavior type For the behavior type in the behavior tag set and the first The number of behavior types that match the standard behavior types in the templates for each teaching stage. For the first The total number of standard behavior types in each teaching stage template.

[0043] In a preferred embodiment, the step of using the rhythm analysis graph as the time axis framework and segmentation basis, and embedding the chained annotations into the target time interval of the rhythm analysis graph according to the pointer direction, to obtain a comprehensive analysis report of the teaching and research activity, includes:

[0044] The time pointer labels of the chained annotations are compared with the temporal axis of the rhythm analysis graph to determine the target temporal interval to which the annotation nodes in the rhythm analysis graph belong.

[0045] Following the logical pointer direction of the chained annotations, the annotation nodes are sequentially filled into the annotation display sub-area of ​​the target time interval to obtain a comprehensive analysis report of the teaching and research activities.

[0046] To address the aforementioned problems, the present invention also provides a smart teaching and research activity management system supporting multimodal input, the system comprising:

[0047] The data acquisition module collects the audio stream, video stream, and comment text of the teaching and research activities based on the user time stamps of the teaching and research activities.

[0048] The chained annotation module writes the comment text into pointer fields pointing to the speech segment index in the audio stream and the behavior tag index in the video stream, thereby obtaining the chained annotation of the teaching and research activity;

[0049] The teaching and research dataset module establishes mutually pointing pointers between the audio stream, the video stream, and the chained annotations based on the time identifier of the teaching and research activity, and encapsulates the data streams after the pointers are established to obtain the teaching and research dataset of the teaching and research activity.

[0050] The composite teaching activity curve module uses the time period of the audio text in the teaching and research dataset as the horizontal axis and the composite value generated by fusing the behavior frequency and the number of speech transitions in the teaching and research dataset during the time period as the vertical axis to construct the composite teaching activity curve of the teaching and research activity.

[0051] The rhythm analysis graph module uses the peaks, troughs, and slope change points of the composite teaching activity curve as rhythm feature points, and matches the behavioral labels corresponding to the rhythm feature points with the behavioral semantic database of the teaching and research activity to obtain the rhythm analysis graph of the teaching and research activity.

[0052] The comprehensive analysis report module uses the rhythm analysis graph as the time axis framework and segmentation basis, and embeds the chain annotations into the target time interval of the rhythm analysis graph according to the pointer direction to obtain the comprehensive analysis report of the teaching and research activities.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] 1. By establishing a unified pointer association mechanism for audio streams, video streams, and comment text, deep fusion and mutual indexing of multimodal data are achieved. During the data acquisition phase, channel association seeds ensure that all data belongs to the same teaching and research activity. In the data processing process, downlink pointers, uplink pointers, source backtracking pointers, and target mapping pointers are constructed layer by layer, ultimately forming a structured teaching and research dataset. This data organization allows any comment to be directly located to its corresponding audio segment and behavioral label. When reviewing teaching and research activities, users can freely jump between audio, video, and text along the pointer direction, greatly improving the comprehensive backtracking efficiency and ease of use of multimodal information.

[0055] 2. By performing time-series segmentation and multi-dimensional numerical fusion on the teaching and research dataset, a composite teaching activity curve is constructed, and a rhythm analysis map with teaching semantic labels is further generated. The composite calculation based on behavior frequency and speech transition frequency comprehensively reflects the dynamic changes in teaching activity, while the matching of rhythm feature points with the behavioral semantic database transforms the numerical curve into a teaching stage division with educational implications. Chained annotations are finally categorized and filled according to the time intervals of the rhythm analysis map. The generated comprehensive analysis report presents both quantitative fluctuations in teaching rhythm and integrates qualitative feedback content corresponding to each stage, providing a complete interpretation path from process data to teaching principles for teaching and research activities. Attached Figure Description

[0056] Figure 1 This is a flowchart illustrating a smart teaching and research activity management method supporting multimodal input, provided in an embodiment of the present invention.

[0057] Figure 2 This is a functional module diagram of a smart teaching and research activity management system supporting multimodal input provided in an embodiment of the present invention;

[0058] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0059] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0060] This application provides a smart teaching and research activity management method that supports multimodal input. The executing entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the smart teaching and research activity management method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0061] Reference Figure 1 The diagram shown is a flowchart illustrating a smart teaching and research activity management method supporting multimodal input according to an embodiment of the present invention. In this embodiment, the smart teaching and research activity management method supporting multimodal input includes:

[0062] In this embodiment of the invention, when the user time stamp based on the teaching and research activity is used to collect the audio stream, video stream, and comment text of the teaching and research activity, it is specifically used for:

[0063] The creation time and user identifier of the teaching and research activity are used as channel association seeds and written into the header information field of the acquisition link corresponding to the teaching and research activity to obtain the audio acquisition channel, video acquisition channel and text acquisition channel of the teaching and research activity.

[0064] The audio stream, video stream, and comment text of the teaching and research activities are collected through the acquisition channel, the video acquisition channel, and the text acquisition channel.

[0065] Specifically, the creation time of a teaching and research activity is automatically recorded by the system when the activity is initiated, and the user identifier of the initiator is also obtained. These two pieces of information together constitute the channel association seed. The system pre-establishes a collection link for each teaching and research activity. This collection link is a logical channel management unit responsible for coordinating the collection process of audio, video, and text data. The channel association seed is written into the header information field of the collection link. The header information field is the starting part of the data structure of the collection link and is used to store metadata related to the activity.

[0066] Specifically, after the audio acquisition channel is activated, it connects to the sound pickup equipment at the teaching and research activity site, continuously receives audio signals and converts them into digital audio stream data. The audio stream is stored as continuous speech segments in chronological order. Simultaneously, the video acquisition channel activates the on-site camera equipment to capture images of the entire teaching and research activity, generating video stream data containing timestamps. Behavioral tags corresponding to the sequence of classroom events are labeled in the video stream using a behavior recognition algorithm.

[0067] Furthermore, the acquisition chain generates three independent sub-channels based on the association seeds in the header information: an audio acquisition channel, a video acquisition channel, and a text acquisition channel. The audio acquisition channel is specifically used to manage audio data acquisition tasks, the video acquisition channel is specifically used to manage video data acquisition tasks, and the text acquisition channel is specifically used to manage comment text acquisition tasks. The header information of each sub-channel inherits the original channel association seeds, ensuring that all subsequently acquired data can be traced back to the same teaching and research activity.

[0068] Furthermore, the text acquisition channel connects to the comment text input interface of the teaching and research platform, collecting comment texts posted by participants during teaching and research activities in real time. Each comment text is accompanied by its timestamp and the poster's identifier. The audio acquisition channel, video acquisition channel, and text acquisition channel work in parallel, continuously outputting corresponding audio streams, video streams, and comment texts according to the timeline of the teaching and research activities. These raw data streams all carry header information for their respective channels, thus maintaining a correlation with the creation time of the teaching and research activities and the user's identifier.

[0069] In summary, by combining creation time and user identifier into a channel association seed and writing it into the header information of the acquisition link, the system establishes a unique and traceable source identifier for all subsequently acquired data. This header information field becomes the metadata foundation for each acquisition channel, ensuring that audio, video, and text acquisition channels inherit the same activity and user identifiers during generation. This design achieves unified management of multimodal data sources, enabling subsequently acquired audio streams, video streams, and comment texts to be directly associated with the same teaching and research activity through the header information field, avoiding potential attribution confusion during multi-channel data acquisition. Simultaneously, because each channel is written with the same association seed during generation, the system can quickly identify the teaching and research activity to which it belongs simply by reading the header information of each data packet, without relying on external association tables. This self-contained data structure simplifies data management and subsequent processing, providing a reliable identification foundation for the alignment and fusion of cross-modal data.

[0070] In summary, the system achieved synchronous capture of multimodal data from teaching and research activities by operating three independent acquisition channels in parallel. The audio acquisition channel focused on continuous acquisition of sound signals, generating an audio stream containing complete speech content; the video acquisition channel focused on continuous acquisition of video signals, generating a video stream containing classroom behavior and scenes; and the text acquisition channel focused on real-time capture of text input, generating comment text containing participant feedback. These three channels operated in parallel without interference, ensuring that different types of data were recorded completely at their most suitable acquisition frequency and format. Each channel continuously attached header information fields during the acquisition process, ensuring that each audio stream, video stream, and comment text carried the creation time and user identifier of the teaching and research activity. This identification-based acquisition method ensured that the data was inextricably linked to the teaching and research activity from the moment it was generated. The system ultimately obtained three types of raw data streams that were time-synchronized and of a unified source. These data streams formed the foundational material for subsequent chained annotation construction, teaching and research dataset encapsulation, and composite teaching activity curve analysis, providing complete data support for a comprehensive reconstruction of the teaching and research activity process.

[0071] In this embodiment of the invention, when writing the comment text into pointer fields pointing to the speech segment index in the audio stream and the behavior tag index in the video stream to obtain the chained annotations of the teaching and research activity, it is specifically used for:

[0072] Based on the global time anchor point of the comment text, in the audio stream's speech segments and the video stream's classroom event sequence, respectively, the target speech segments and target behavior tags with start times less than or equal to the global time anchor point and end times greater than or equal to the global time anchor point are searched to obtain the speech segment index and behavior tag index of the teaching and research activity.

[0073] Write the speech segment index and the behavior tag index into the speech pointer field and behavior pointer field of the comment text data structure to obtain the chained annotation of the teaching and research activity.

[0074] Specifically, the system extracts the timestamp associated with the generation of the comment text from its data structure; this timestamp is called the global time anchor. The audio stream consists of multiple consecutive speech segments, each with its start and end times recorded at the time of generation. These speech segments are stored in the audio stream's speech segment list in chronological order. In the video stream, a sequence of classroom events is marked using a behavior recognition algorithm. Each classroom event corresponds to a behavior label, which also has its start and end times recorded. These classroom events are stored in the video stream's classroom event sequence in chronological order.

[0075] Specifically, after obtaining the speech segment index and behavior tag index, the system locates the data structure of the comment text being processed. The comment text data structure reserves speech pointer and behavior pointer fields during creation, both of which are initially empty. The system writes the speech segment index value to the speech pointer field and the behavior tag index value to the behavior pointer field. This writing operation is performed by assigning values ​​to the data structure fields, that is, directly storing the index value into the memory location of the corresponding field.

[0076] Furthermore, the system first iterates through the list of audio segments in the audio stream. For each segment, it reads its start and end times, comparing the start and end times with the global time anchor. If the start time is less than or equal to the global time anchor and the end time is greater than or equal to the global time anchor, the segment is the target segment. The system obtains a unique identifier for this segment in the segment list; this identifier is the segment index. The system then iterates through the sequence of classroom events in the video stream. For each classroom event, it reads the start and end times of its behavior label, comparing the start and end times with the global time anchor. If the start time is less than or equal to the global time anchor and the end time is greater than or equal to the global time anchor, the behavior label is the target behavior label. The system obtains a unique identifier for this behavior label in the classroom event sequence; this identifier is the behavior label index.

[0077] Furthermore, after the writing process is complete, the data structure of the comment text contains pointers to specific speech segments in the audio stream and pointers to specific behavior tags in the video stream. Such comment text is called chained annotation for teaching and research activities. Chained annotation thus possesses the ability to associate with specific content in the audio and video streams. Its internal data fields explicitly record the indexes of the pointed-to speech segments and behavior tags, which can be used later to quickly locate and trace back the corresponding multimedia segments.

[0078] In summary, by using a global time anchor as a time benchmark, the system achieves precise alignment between comment text and content segments occurring synchronously in the audio and video streams. This lookup mechanism ensures that each comment text can find the audio segment and the classroom event that was happening at the time it was generated, thus establishing a natural time-based association between the comment text and the multimedia content. After obtaining the audio segment index and behavior tag index, the system lays the positional foundation for subsequent cross-modal data fusion. These two indexes become the bridge connecting text comments and specific teaching segments, enabling previously isolated comment texts to be located within multimedia content, providing core support for constructing deeply correlated teaching and research data.

[0079] In summary, the newly added voice pointer and behavior pointer fields in the chained annotation data structure allow each annotation to directly point to its corresponding audio segment and video behavior tag. This built-in pointer design transforms annotations from mere text attached to a timeline into key nodes connecting multimodal data. When it's necessary to revisit the teaching scenario corresponding to the annotation, the system can quickly locate the specific voice segment in the audio stream using the voice segment index in the voice pointer field, and quickly locate the specific classroom event in the video stream using the behavior tag index in the behavior pointer field. This direct pointer access method significantly improves the efficiency of data retrieval and presentation, providing a data structure foundation for generating comprehensive analysis reports containing deep correlations.

[0080] In this embodiment of the invention, the step of establishing mutually pointing pointers between the audio stream, the video stream, and the chained annotations based on the time identifier of the teaching and research activity, and encapsulating the data streams after establishing the pointers to obtain the teaching and research dataset of the teaching and research activity, is specifically used for:

[0081] Based on the time interval of the speech segment in the audio stream, the behavior labels and corresponding behavior label indices located within the time interval are traversed in the classroom event sequence of the video stream, and a downlink pointer pointing to the behavior label index is created in the data structure of the speech segment.

[0082] Based on the time interval of the behavior label, find the chained comments whose synchronization time point is within the time interval and their corresponding comment identifiers in the chained comments, and create an uplink pointer to the comment identifier in the data structure of the behavior label;

[0083] A source backtracking pointer is established pointing to the corresponding speech segment in the audio stream based on the speech segment index in the speech pointer field, and a target mapping pointer is established pointing to the corresponding behavior label in the classroom event sequence based on the behavior label index in the behavior pointer field.

[0084] The voice text, the classroom event sequence, and the chained annotations are stored together. The downlink pointer, the uplink pointer, the source backtracking pointer, and the target mapping pointer are written into the storage structure of the teaching and research activity as the association between the voice text, the classroom event sequence, and the chained annotations, thus obtaining the teaching and research dataset of the teaching and research activity.

[0085] Specifically, the system reads the data structure of the current speech segment from the audio stream. This structure stores the start and end times of the speech segment. These two time values ​​constitute the time interval of the speech segment. The system then traverses the classroom event sequence in the video stream. The classroom event sequence is an ordered list in which each element contains a behavior label, the start and end times of the behavior label, and the unique behavior label index of the behavior label in the sequence.

[0086] Specifically, the system obtains the data structure of the current behavior label from the classroom event sequence of the video stream. This structure stores the start and end times of the behavior label. These two time values ​​constitute the time interval of the behavior label. The system then iterates through all generated chained annotations. The data structure of each chained annotation stores its global time anchor point at the time of its generation. The system judges each chained annotation by comparing whether its global time anchor point is greater than or equal to the start time of the current behavior label and less than or equal to the end time of the current behavior label. When both conditions are met, the chained annotation is considered to have its synchronization time point within the time interval of the current behavior label.

[0087] Specifically, the system iterates through all generated chained comments. The data structure of each chained comment already contains the previously written speech pointer field and behavior pointer field. The speech pointer field stores the speech segment index, and the behavior pointer field stores the behavior tag index. The system accesses the speech segment list of the audio stream based on the speech segment index in the speech pointer field. When the speech segment object with the exact same index value is located in the speech segment list, the system creates a source backtrack pointer field in the data structure of the chained comment and assigns the memory address or reference of the located speech segment object to the source backtrack pointer field, thereby establishing a source backtrack pointer from the chained comment to the corresponding speech segment in the audio stream.

[0088] Specifically, the system prepares to persistently store all speech-text data in the audio stream, all classroom event sequence data in the video stream, and all chained annotation data. This data is written into a storage structure pre-allocated for the teaching and research activity. This storage structure is a dedicated database or file system directory used to store all relevant data for the teaching and research activity. During the writing process, the system collects all established pointer relationships, including downlink pointers created in the speech segment data structure, uplink pointers created in the behavior label data structure, source backlink pointers created in the chained annotation data structure, and target mapping pointers created in the chained annotation data structure. The system uses these pointers as metadata representing the association between speech-text, classroom event sequences, and chained annotations, and writes them into the storage structure of the teaching and research activity.

[0089] Furthermore, the system judges each classroom event by comparing whether the start time of the behavior label is less than or equal to the end time of the audio segment and whether the end time of the behavior label is greater than or equal to the start time of the audio segment. When both conditions are met, the behavior label is considered to be within the time interval of the current audio segment. The system records the behavior label index corresponding to these behavior labels. After traversal, the system allocates a downlink pointer field in the data structure of the current audio segment. This field is designed as a list structure that can store multiple pointers. The system writes all the recorded behavior label indices into the downlink pointer field in sequence, with each index value stored as a reference. Thus, a downlink pointer pointing to the relevant behavior label in the sequence of classroom events is successfully created in the audio segment.

[0090] Furthermore, the system records the comment identifiers corresponding to these chained comments. The comment identifier is a unique number for each chained comment in the system. After traversal, the system allocates an uplink pointer field in the data structure of the current behavior label. This field is also designed as a list structure that can store multiple pointers. The system writes all the recorded comment identifiers into this uplink pointer field in sequence. Each comment identifier is stored as a reference, thus successfully creating an uplink pointer in the behavior label that points to the relevant comment in the chained comment set.

[0091] Furthermore, at the same time, the system accesses the classroom event sequence of the video stream based on the behavior label index in the behavior pointer field, locates the behavior label object with the same index value in the classroom event sequence, creates a target mapping pointer field in the chained annotation data structure, and assigns the memory address or reference of the located behavior label object to the target mapping pointer field, thereby establishing a target mapping pointer from the chained annotation to the corresponding behavior label in the classroom event sequence.

[0092] Furthermore, the specific writing method is to record the identifiers of other objects that each data object points to in the metadata area, or to store all pointer mappings in an independent association table. When all data and pointer relationships are written, the complete data set stored in the storage structure is defined as the teaching and research dataset of the teaching and research activity.

[0093] In summary, by filtering classroom event sequences through the time intervals of audio segments, the system accurately captures all teaching behaviors occurring concurrently with each audio segment and stores the indices of these behavior tags as pointers within the audio segment's data structure. This operation allows each audio segment to directly access all behavior tags occurring within its corresponding time period, establishing a downlink relationship from audio content to video behavior. When subsequent analysis of classroom interactions corresponding to a particular audio segment is needed, the system can quickly locate the relevant behavior tags using the downlink pointers within the audio segment, without needing to re-traverse the entire classroom event sequence. This pre-established link significantly improves the retrieval efficiency of cross-modal data and provides a direct access path for the temporal analysis of teaching behaviors.

[0094] In summary, this operation establishes an upward link from video actions to textual comments, enabling each action tag to directly access real-time feedback from participants. When it is necessary to assess the acceptance of a particular teaching action or analyze the discussion it sparks, the system can quickly access related chained comments through the upward pointers in the action tag. This link tightly connects objective teaching actions with subjective commentary text, providing direct semantic support for evaluating teaching effectiveness.

[0095] In summary, source backtracking pointers allow chained annotations to directly locate the audio segment at the time of their generation, while target mapping pointers allow them to directly locate the video action tag at the time of their generation. These two types of pointers enable chained annotations to have bidirectional access capabilities to both audio and video. When a user views a chained annotation, the system can quickly play the corresponding audio segment using the source backtracking pointer and quickly display the corresponding classroom event using the target mapping pointer. This bidirectional pointer design creates an interactive connection between text annotations and multimedia content, greatly enhancing the readability and explorability of teaching and research data.

[0096] In summary, this dataset not only contains the original audio streams, video streams, and commentary text, but more importantly, it includes four types of relationships: downlink pointers, uplink pointers, source backlink pointers, and target mapping pointers. These pointers form a complete mesh index structure between the audio text, classroom event sequences, and chained annotations. When conducting subsequent teaching analysis, the system can directly access any related data segment from this teaching research dataset via pointers, without needing to perform time alignment and lookup operations again. This associative storage method integrates fragmented multimodal data into an organic whole, providing a unified data foundation for subsequently constructing composite teaching activity curves, generating rhythm analysis maps, and producing comprehensive analysis reports, ensuring that all analytical operations are based on complete and interconnected data.

[0097] In this embodiment of the invention, when constructing the composite teaching activity curve of the teaching and research activity by using the time period of the speech text in the teaching and research dataset as the horizontal axis and the composite value generated by fusing the corresponding values ​​of the behavior frequency and the number of speech transitions in the teaching and research dataset within the time period as the vertical axis, it is specifically used for:

[0098] The total duration of the audio text in the teaching and research dataset is divided into continuous time periods, and the total number of behavior labels whose start time is located in the continuous time period in the classroom event sequence in the teaching and research dataset is taken as the behavior frequency of the continuous time period.

[0099] The number of times the speaker identifier of the speech text in the teaching and research dataset changes between adjacent speech segments is taken as the number of speech transitions in the continuous time period.

[0100] Using the continuous time period as the horizontal axis value and the frequency of the row as the left vertical axis value, the left data point corresponding to the continuous time period is obtained;

[0101] Using the continuous time period as the horizontal axis value and the number of speech transitions as the right vertical axis value, the right data point corresponding to the continuous time period is obtained;

[0102] By connecting all the left and right data points in chronological order, the frequency curve and interaction conversion curve of the teaching and research activities are obtained.

[0103] The behavior frequency curve and the interaction conversion curve are integrated into a composite teaching activity curve for the teaching and research activities.

[0104] Specifically, the system reads all audio text data from the teaching and research dataset. These audio texts are arranged in chronological order, and each audio text segment contains a start time and an end time. The system takes the start time of the first audio text as the starting point of the total duration and the end time of the last audio text as the ending point of the total duration. The time interval between the starting point and the ending point is determined as the total duration of the audio text. The system divides this total duration into multiple continuous and non-overlapping time segments according to a pre-set fixed time length.

[0105] Specifically, the system iterates through the classroom event sequence in the teaching and research dataset. Each behavior label in the classroom event sequence records its start time and end time. For each behavior label, the system extracts its start time and compares it with each time period in the list of consecutive time periods to find the time period whose start time falls within the start and end time interval of that time period.

[0106] Specifically, the system processes the speech and text data in the teaching and research dataset based on a pre-defined list of consecutive time periods. Each speech segment in the speech and text data records the speaker's identifier. The system arranges all speech segments in chronological order, with adjacent speech segments being consecutive in time. For each consecutive time period, the system locates the start and end times of that time period.

[0107] Specifically, the system uses each time period in the list of consecutive time periods as the horizontal axis value. The horizontal axis value is usually represented by the start time or time number of that time period. For each consecutive time period, the system retrieves the behavior frequency of that time period that was previously calculated and stored, and uses this behavior frequency as the left vertical axis value corresponding to that time period. The system combines the horizontal axis value and the left vertical axis value into a data point, which is called a left data point. The system generates a left data point for each consecutive time period in chronological order, and saves these left data points sequentially into a data point set.

[0108] Specifically, the system extracts all left data points and arranges them in chronological order from front to back. Starting from the first left data point, adjacent left data points are connected sequentially with straight line segments. That is, the first left data point is connected to the second left data point, and then the second left data point is connected to the third left data point, and so on until the last left data point. The continuous broken line formed after all left data points are connected is the frequency curve of teaching and research activities.

[0109] Specifically, the system integrates the behavior frequency curve and the interaction conversion curve into a new curve. For each consecutive time period, the system obtains the behavior frequency and the number of speech conversions in that time period. The system finds the maximum and minimum values ​​of behavior frequency and speech conversions in all time periods. For the behavior frequency of each time period, the system subtracts the minimum value and divides it by the difference between the maximum and minimum values ​​to obtain the normalized value of the behavior frequency in that time period.

[0110] Furthermore, each time period is defined by its start and end times. These consecutive time periods cover the entire audio-visual timeline of the teaching and research activity. There are no gaps between adjacent time periods, and all time periods are of equal length. After the division is completed, the system obtains a list of consecutive time periods arranged in chronological order.

[0111] Furthermore, the system maintains a counter for this time period. Each time a behavior label whose start time belongs to this time period is found, the counter value for this time period is incremented by one. After all behavior labels have been traversed, the counter value for each consecutive time period is the total number of behavior labels whose start time is within this time period. The system defines this total number as the behavior frequency of this consecutive time period and stores each time period in association with its corresponding behavior frequency.

[0112] Furthermore, the system iterates through all audio segments within the time period, extracts the speaker identifiers of these audio segments, and checks whether the speaker identifiers of two adjacent audio segments are the same. If the speaker identifiers of adjacent audio segments are different, it is determined that a speech transition has occurred. The system counts the number of times the speaker identifiers of all adjacent audio segments change within the time period. This cumulative number is the number of speech transitions in the continuous time period. The system associates and stores each time period with its corresponding number of speech transitions.

[0113] Furthermore, the system also uses each time period in the list of consecutive time periods as the horizontal axis value. For each consecutive time period, the system retrieves the number of speech transitions that was previously calculated and stored for that time period, and uses this number of speech transitions as the right vertical axis value corresponding to that time period. The system combines the horizontal axis value and the right vertical axis value into a data point, which is called a right data point. The system generates a right data point for each consecutive time period in chronological order, and saves these right data points sequentially into another data point set.

[0114] Furthermore, the system simultaneously extracts all right data points and arranges them in chronological order from front to back. Starting from the first right data point, adjacent right data points are connected sequentially with straight line segments. The continuous broken line formed after all right data points are connected is the interactive transformation curve of the teaching and research activity.

[0115] Furthermore, for the number of speech transitions in each time period, the minimum value is subtracted and then divided by the difference between the maximum and minimum values ​​to obtain the normalized value of the number of speech transitions in that time period. The system adds these two normalized values ​​to obtain the composite value for that time period. The system uses the continuous time period as the horizontal axis and the calculated composite value as the vertical axis to generate a new data point. All the new data points are connected sequentially with straight line segments according to the time order. The final continuous broken line is the composite teaching activity curve of the teaching and research activity.

[0116] In summary, the system transforms the originally continuous teaching activities into quantifiable time units by dividing the complete audio duration into equal-length continuous time segments. The total number of behavioral tags within each time unit is accurately counted as the behavior frequency. This quantification process allows the occurrence density of various behaviors in teaching activities to be numerically presented, with behavior frequency becoming the core indicator for measuring the intensity of teaching activities in each time segment. The system obtains a behavior frequency sequence arranged in chronological order, which directly reflects the distribution pattern of teaching behaviors on the time axis, providing basic data support for subsequent analysis of fluctuations in the teaching rhythm. At the same time, this segmented statistical method eliminates the influence of random fluctuations in the timing of behavior occurrences in the original classroom event sequence, making the overall situation of teaching activities clearer and more discernible.

[0117] In summary, the number of speaker transitions directly reflects the level of classroom interaction. A higher value indicates more frequent speaker changes and more thorough interaction within that time period, while a lower value suggests that the speaker remains relatively fixed, possibly engaged in lecturing or focused work. The system obtained a chronological sequence of speaker transition frequencies, which serves as a quantitative indicator of classroom interaction intensity. This provides an objective basis for subsequent analysis of the dynamic changes in teaching interaction patterns. This statistical method based on changes in speaker identifiers avoids interference from subjective judgment and ensures the accuracy and repeatability of the data.

[0118] In summary, this coordinate-based processing transforms the temporal distribution characteristics of teaching activities into graphically representable data units, providing raw material for subsequent plotting of behavior frequency curves. The set of data points on the left fully preserves all information about the frequency of behaviors over time; each point corresponds to a specific time location and a quantified behavior density value. This one-to-one data structure ensures that subsequent curve connection operations can accurately reconstruct the temporal trajectory of behavior frequency changes. Similarly, converting the intensity distribution of classroom interaction into graphically representable data units provides raw material for subsequent plotting of interaction transition curves. The set of data points on the right fully preserves all information about the number of speaking transitions over time; each point corresponds to a specific time location and a quantified interaction intensity value. This processing ensures that subsequent curve connection operations can accurately reconstruct the temporal trajectory of classroom interaction intensity changes.

[0119] In summary, the system connects the left-hand data points arranged chronologically by straight line segments, forming a continuous broken line, the behavior frequency curve. This curve visually illustrates the fluctuation of behavior frequency over time; the rising segment indicates an increase in teaching activity density, while the falling segment indicates a decrease. Simultaneously, the system connects the right-hand data points arranged chronologically by straight line segments, forming another continuous broken line, the interaction transition curve. This curve visually illustrates the fluctuation of speaking transitions over time; the rising segment indicates increased classroom interaction, while the falling segment indicates decreased classroom interaction. These two curves transform the originally discrete data points into a continuous visual graph, allowing educational researchers to clearly observe the overall rhythmic changes in teaching activities and classroom interactions. The peaks, troughs, and slopes of the curves become key visual clues for subsequent analysis of teaching characteristics.

[0120] In summary, by integrating two indicators with different dimensions—behavior frequency and speaker switching frequency—a single composite teaching activity curve is generated. This curve comprehensively reflects information from two dimensions: teaching activity density and classroom interaction intensity. Each composite teaching indicator value simultaneously includes the frequency of behavior occurrence and speaker switching frequency within that time period. The composite teaching activity curve eliminates the tedious comparative analysis required when observing two separate curves, condensing the core characteristics of teaching rhythm into an easily interpretable curve. Every fluctuation in the curve comprehensively reflects the combined changes in teaching behavior and classroom interaction. By observing the composite teaching activity curve, teaching researchers can quickly grasp the rhythmic context of the entire teaching research activity. The peaks of the curve correspond to the peak periods of teaching activities and interactions, while the troughs correspond to relatively calm periods. This comprehensive presentation provides a unified numerical basis for subsequent extraction of rhythmic feature points, greatly simplifying the process of teaching rhythm analysis.

[0121] In this embodiment of the invention, when integrating the behavior frequency curve and the interaction conversion curve into a composite teaching activity curve for the teaching and research activity, it is specifically used for:

[0122] The frequency of behaviors in each time period of the behavior frequency curve and the number of speech transitions in each time period of the interaction transition curve are normalized to obtain the frequency scale value and transition scale value of the corresponding time period in the teaching and research activity.

[0123] The Euclidean distance between the frequency scale value and the transformation scale value is calculated based on the same time period of the teaching and research activities to obtain the composite teaching index of the teaching and research activities.

[0124] Using the continuous time period of the teaching and research activities as the horizontal axis and the composite teaching indicators as the vertical axis, the data points are connected sequentially in chronological order to obtain the composite teaching activity curve of the teaching and research activities.

[0125] Specifically, the system iterates through all consecutive time periods, collecting the frequency values ​​of behaviors across all time periods, and then identifies the maximum and minimum behavior frequencies. For each consecutive time period, the system subtracts the minimum behavior frequency from the behavior frequency of that period, and then divides this difference by the difference between the maximum and minimum behavior frequencies. The result obtained is the frequency scale value for that time period.

[0126] Specifically, the system treats the same continuous time period of teaching and research activities as a processing unit, retrieving the frequency scale value and transformation scale value of that time period from storage. The system considers the frequency scale value as a value in one dimension and the transformation scale value as a value in another dimension; these two values ​​together constitute a point in a two-dimensional space.

[0127] Specifically, the system uses a list of pre-defined consecutive time periods for teaching and research activities as the basis for the horizontal axis, with each consecutive time period identified by its start time or sequential number. The system uses the calculated composite teaching indicators as the vertical axis value, generating a data point for each consecutive time period. The horizontal axis of this data point represents the current consecutive time period, and the vertical axis represents the composite teaching indicator corresponding to that time period.

[0128] Further, as before, the system iterates through all consecutive time periods, collecting the speech transition counts for each period, and identifying the maximum and minimum speech transition counts. For each consecutive time period, the system subtracts the minimum speech transition count from the total speech transition count for that period, then divides this difference by the difference between the maximum and minimum speech transition counts. The result is the transition scale value for that time period. The system stores the calculated frequency scale value and transition scale value for each consecutive time period.

[0129] Furthermore, the system calculates the distance from this point to the origin. Specifically, it squares both the frequency scale value and the transformation scale value, adds the results of these two square operations, and then takes the square root of the sum. The final result is the composite teaching indicator for that time period. The system processes each consecutive time period in the same way, calculating the composite teaching indicator for each time period sequentially, and stores the correspondence between each time period and its composite teaching indicator.

[0130] Furthermore, the system arranges all generated data points in chronological order according to consecutive time periods. Starting with the data points from the first time period, a straight line segment connects the first data point to the second data point, then connects the second data point to the third data point, and so on, until the data points from the last time period are connected to the data points preceding them. The continuous broken line formed after all the straight line segments are connected is the composite teaching activity curve of the teaching and research activity.

[0131] In summary, normalization transforms two indicators with different dimensions and numerical ranges—behavior frequency and speech transition count—into a unified scale value. The original values ​​of behavior frequency may range from zero to tens, and the original values ​​of speech transition count may range from zero to a dozen or so. After normalization, both are mapped to a standard range of zero to one. The frequency scale value eliminates the absolute numerical differences in behavior frequency caused by variations in time period length or activity type, while the transformation scale value eliminates the absolute numerical differences in speech transition count, making these two previously incomparable indicators comparable. This scaling process lays the foundation for subsequent integrated calculations, ensuring that behavior frequency and speech transition count have equal contribution weights in the composite teaching indicators. It avoids the problem of one indicator dominating the calculation results due to an excessively large numerical range. Furthermore, the frequency scale value and transformation scale value for each time period retain the changing trends and relative relationships of the original data.

[0132] In summary, the frequency scale value and the transformation scale value are considered as two coordinate components on a two-dimensional plane. By calculating the distance from the point formed by these two components to the origin, a comprehensive numerical index is obtained. This composite teaching index reflects the combined level of behavioral density and interaction intensity within a given time period. When both the behavioral frequency scale value and the speech transformation scale value are high, the composite teaching index is large, indicating that the teaching activities during that time period are intensive and the interaction is active. When both are low, the composite teaching index is small, indicating that the teaching activities during that time period are sparse and the interaction is bland. When one is high and the other is low, the composite teaching index is at a moderate level, indicating that one aspect of the teaching activities during that time period is prominent while the other is weak. The Euclidean distance calculation method ensures that the composite teaching index can reflect the contribution of both dimensions in a balanced way. It will not cause the index to become zero due to one dimension being zero, nor will it overemphasize the influence of one dimension. This comprehensive index provides a unified data foundation for subsequent curve plotting.

[0133] In summary, the composite teaching indicators for each time period are transformed into two-dimensional coordinate points, and these points are connected by straight line segments in chronological order to form a continuous broken line. This composite teaching activity curve comprehensively presents the overall rhythmic changes of teaching behavior and classroom interaction, with each fluctuation representing a comprehensive fluctuation in teaching activity. Compared to observing behavior frequency curves and interaction transition curves separately, the composite teaching activity curve eliminates the cognitive burden of simultaneously tracking two curves and psychologically integrating them, condensing complex teaching rhythm information into a single visual graph. By observing the composite teaching activity curve, teaching researchers can quickly grasp the activity change patterns of the entire teaching research activity. The peaks of the curve correspond to high-intensity periods with relatively dense teaching behavior and interaction, while the troughs correspond to calmer periods with relatively sparse teaching behavior and interaction. The rising and falling segments of the curve reflect the acceleration and deceleration of the teaching rhythm. This curve becomes the core basis for subsequently extracting rhythmic feature points and dividing teaching stages, providing direct visual input for the quantitative analysis and semantic interpretation of teaching rhythm.

[0134] In this embodiment of the invention, when the peaks, troughs, and slope change points of the composite teaching activity curve are used as rhythm feature points, and the behavioral labels corresponding to the rhythm feature points are matched with the behavioral semantic database of the teaching and research activity to obtain the rhythm analysis map of the teaching and research activity, the specific method is as follows:

[0135] The local maxima and local minima of the composite teaching activity curve are taken as peaks and troughs, the points where the sign of the first derivative changes in the composite teaching activity curve are taken as slope change points, and the peaks, troughs and slope change points are all taken as rhythm feature points.

[0136] Taking the time position of the rhythm feature point on the time axis of the audio text of the teaching and research activity as the center, extend forward to the position of the preceding rhythm feature point and backward to the position of the subsequent rhythm feature point. If there is no rhythm feature point before the time position, extend forward to the starting point of the audio text in the teaching and research activity. If there is no rhythm feature point after the time position, extend backward to the ending point of the audio text in the teaching and research activity to obtain the feature point time window of the teaching and research activity.

[0137] Extract behavioral labels whose starting time is within the time window of the feature point from the classroom event sequence of the teaching and research activities to obtain a set of behavioral labels corresponding to the rhythm feature point;

[0138] Calculate the matching degree between the behavior tag set and the teaching stage template in the behavior semantic library of teaching and research activities, and take the stage name of the teaching stage template corresponding to the maximum value of the matching degree as the teaching semantic tag of the rhythm feature point;

[0139] The teaching semantic tags are labeled to the corresponding rhythm feature points in the composite teaching activity curve to obtain the rhythm analysis map of the teaching and research activity.

[0140] Specifically, the system identifies local maxima and local minima from the composite teaching activity curve. The composite teaching activity curve is composed of a series of ordinate values ​​corresponding to consecutive time periods. The system traverses all data points. For each data point, it compares its ordinate value with the ordinate values ​​of the two adjacent data points. If the ordinate value of a data point is greater than the ordinate value of both the preceding and following data points, then the data point is identified as a local maximum, i.e., a peak. If the ordinate value of a data point is less than the ordinate value of both the preceding and following data points, then the data point is identified as a local minimum, i.e., a trough.

[0141] Specifically, the system processes each rhythmic feature point using its time position as the center point. First, the system arranges all rhythmic feature points in chronological order. For the currently processed rhythmic feature point, the system searches for the time position of its preceding rhythmic feature point. If a preceding rhythmic feature point exists, the forward-extending boundary is determined as the time position of that preceding rhythmic feature point. If no preceding rhythmic feature point exists, the forward-extending boundary is determined as the starting time point of the audio text in the teaching and research activity.

[0142] Specifically, the system extracts a set of behavioral labels corresponding to each rhythm feature point from the classroom event sequence in the teaching and research dataset. The classroom event sequence contains multiple behavioral labels, each recording its start time. For the feature point time window of the current rhythm feature point, the system iterates through all behavioral labels in the classroom event sequence.

[0143] Specifically, the behavioral semantic database stores multiple teaching stage templates. Each teaching stage template contains a set of standard behavior types and a preset weight value corresponding to each standard behavior type. The system calculates the matching degree for each teaching stage template. The specific calculation process is as follows: For each behavior tag in the behavior tag set, it is determined whether the behavior type of the behavior tag exists in the standard behavior types of the current teaching stage template. If it exists, the behavior tag is considered a matching behavior tag. The system records all matching behavior tags. For each matching behavior tag, the preset weight value corresponding to its behavior type in the teaching stage template is obtained. At the same time, the actual number of times the behavior tag appears within the feature point time window is obtained, as well as the specific timestamp of the occurrence of the behavior tag is also obtained.

[0144] Specifically, the system annotates the calculated teaching semantic tags onto the corresponding rhythmic feature points in the composite teaching activity curve. The composite teaching activity curve is a line graph with continuous time periods on the x-axis and composite teaching indicators on the y-axis. Each rhythmic feature point has its specific time position coordinates on the curve. The system adds a label at the time position of each rhythmic feature point on the curve, with the label content being the teaching semantic tag for that rhythmic feature point.

[0145] Furthermore, the system simultaneously identifies slope change points. For each line segment between two adjacent data points, its slope is calculated as the difference between the ordinate and the abscissa. When the slope of two adjacent line segments changes from positive to negative or vice versa, the data point connecting these two line segments is the point where the sign of the first derivative changes, and this point is identified as a slope change point. The system collects all identified peaks, troughs, and slope change points, defining these points as rhythmic characteristic points of the teaching and research activities, with each rhythmic characteristic point corresponding to a specific time position.

[0146] Furthermore, the system then searches for the time position of the next rhythmic feature point. If a subsequent rhythmic feature point exists, the forward extension boundary is defined as the time position of that subsequent rhythmic feature point. If no subsequent rhythmic feature point exists, the forward extension boundary is defined as the end time point of the audio text in the teaching and research activity. The system defines the time interval between the forward extension boundary and the backward extension boundary as the feature point time window of that rhythmic feature point. The start time of this window is the forward extension boundary, and the end time is the backward extension boundary.

[0147] Furthermore, the system checks each behavior label's start time to see if it is greater than or equal to the start time of the feature point's time window and less than or equal to the end time of the feature point's time window. If this condition is met, the behavior label's start time falls within the feature point's time window, and the system adds the behavior label to the behavior label set corresponding to the current rhythm feature point. After traversal, the system obtains a behavior label set for each rhythm feature point, which contains all behavior labels whose start times fall within the feature point's time window.

[0148] Further, the system calculates the time decay factor for each matching behavior tag. This factor is determined based on the time difference between the occurrence timestamp of the behavior tag and the time position of the rhythm feature point, as well as the total duration of the feature point's time window. The smaller the time difference, the closer the decay factor is to one; the larger the time difference, the smaller the decay factor. The system multiplies the preset weight value, the actual occurrence frequency, and the time decay factor for each matching behavior tag, and then sums the products of all matching behavior tags to obtain the first value. Simultaneously, the system counts the number of behavior types matching the current teaching stage template in the behavior tag set, obtains the total number of standard behavior types in the current teaching stage template, and the preset weight value for each standard behavior type, and sums the preset weight values ​​of all standard behavior types to obtain the second value. The system divides the first value by the second value, and then multiplies the quotient by the ratio of the number of matching behavior types to the total number of types, finally obtaining the matching degree of the teaching stage template. The system performs the above calculations for each teaching stage template, obtaining a series of matching degree values, selecting the largest matching degree value, and using the stage name of the teaching stage template corresponding to the largest value as the teaching semantic tag for the current rhythm feature point.

[0149] Furthermore, after the annotation is completed, each rhythmic feature point on the entire composite teaching activity curve is accompanied by its corresponding teaching semantic label. These labels are arranged in chronological order, forming a semantic description of the rhythm of the teaching and research activity. The resulting annotated composite teaching activity curve is the rhythm analysis map of the teaching and research activity.

[0150] In summary, by identifying local maxima and minima on the curve, the system accurately captures the critical moments when the activity level of teaching activities reaches its peak and trough. These peaks correspond to periods of high activity, while the troughs correspond to periods of calm. The system also identifies points where the sign of the first derivative changes (i.e., slope change points) by calculating the changes in the slope of adjacent line segments. These points correspond to the turning points where the teaching rhythm shifts from acceleration to deceleration or vice versa. By using these three types of points collectively as rhythmic feature points, the system obtains a complete description of the key points of the composite teaching activity curve. These feature points cover all important changes in the teaching rhythm, providing accurate positioning benchmarks for subsequent division of teaching stages. Each feature point corresponds to a specific turning point or extreme moment in the teaching rhythm.

[0151] In summary, each rhythmic feature point is used as the core, and adjacent feature points serve as natural boundaries to define a dedicated time analysis interval for each feature point. This division ensures that each feature point's time window corresponds to a complete teaching rhythm segment. The starting point of the window is the previous rhythmic change point or activity start point, and the ending point is the next rhythmic change point or activity end point. The feature point's time window includes not only the moment the feature point itself exists but also the preceding and following context of the teaching rhythm segment it represents. This allows subsequent behavioral analysis of each feature point to be conducted within its complete context, avoiding the contextual gaps caused by isolated analysis of a single moment. The length of each window is dynamically determined by the spacing between adjacent rhythmic feature points, reflecting the actual duration of the teaching segment.

[0152] In summary, by using feature point time windows as filtering criteria, all teaching behaviors occurring within a specific time window are accurately captured from the classroom event sequence. The start time of each behavior label is used as a criterion to ensure that the extracted behavior labels actually occurred within the teaching segment represented by the feature point. The behavior label set comprehensively records the various types of teaching behaviors that appeared within this time period and their frequency, providing objective data input for subsequent semantic matching. Since each feature point has its own exclusive time window, the system generates an independent behavior label set for each rhythm feature point. These sets reflect the differences in the behavioral composition of different teaching segments. The behavior label set corresponding to the peak point may contain more interactive behaviors, while the behavior label set corresponding to the trough point may contain more lecturing behaviors. This difference provides a basis for subsequent identification of teaching stages.

[0153] In summary, the system systematically compares the actual teaching behaviors within each feature point's time window with preset teaching stage templates, calculating the matching degree to find the most similar teaching stage template for each behavior tag set. The matching degree calculation comprehensively considers multiple factors, including whether the behavior type matches, the preset weight of the matching behavior, the actual frequency of the matching behavior within the window, the proximity of the matching behavior's occurrence time to the feature point's time, and the breadth of coverage of the matching behavior types. This multi-factor comprehensive evaluation ensures the accuracy and reliability of the matching results; the more core behaviors with high weights appear, the closer their occurrence time is to the feature point, and the more comprehensive the coverage of behavior types, the higher the matching degree. Finally, the system selects the stage name of the teaching stage template with the highest matching degree as its teaching semantic label for each rhythm feature point. These labels, such as introduction stage, lecture stage, interaction stage, and summary stage, transform abstract numerical curve features into concrete teaching stage semantics, giving the analysis results of teaching rhythm a clear educational meaning.

[0154] In summary, directly labeling each rhythmic feature point with its corresponding teaching semantic tag on the composite teaching activity curve imbues the curve, which was originally composed solely of numerical values, with a semantic interpretation. The rhythm analysis graph, while retaining all quantitative information from the original composite teaching activity curve, adds teaching stage name tags at each peak, trough, and slope change point. These tags are arranged chronologically, fully presenting the evolution of the teaching stages from start to finish of the teaching research activity. By observing the rhythm analysis graph, teaching researchers can clearly see which period belongs to the introduction stage, which to the lecture stage, which to the interaction stage, which to the summary stage, and the turning points between each stage. If a peak in the graph is labeled as an interaction stage, it indicates high teaching activity and a significant amount of interactive behavior during that period; if a trough is labeled as a lecture stage, it indicates low teaching activity and a predominantly lecture-based approach during that period. This semantic visualization elevates the teaching rhythm analysis from numerical observation to a level of teaching understanding, providing a structured framework organized by teaching stage for the subsequent generation of a comprehensive analysis report.

[0155] In this embodiment of the invention, the formula for calculating the matching degree is specifically used for:

[0156]

[0157] in, For the set of behavior labels and the first The matching degree of the template for each teaching stage. This serves as an index for the template of the aforementioned teaching stage. For the behavior type in the behavior tag set and the first The total number of behavior tags that match the standard behavior types in each teaching stage template. The sequence number of the behavior tag in the behavior tag set that matches the teaching stage template. For the first behavior tag set The behavior type corresponding to each matching behavior tag in the first... The preset weight values ​​corresponding to each teaching stage template For the first behavior tag set The actual number of times the behavior type corresponding to each matching behavior label appears within the time window of the feature point. It is a natural constant. The time decay coefficient, For the first behavior tag set The timestamp of each matching behavior tag. The time position of the rhythmic feature point. The total duration of the time window for the feature points. For the first The total number of all standard behavior types in each teaching stage template This serves as an index for the standard behavior types in the teaching stage template. To indicate the first In the first teaching stage template Preset weight values ​​for each standard behavior type For the behavior type in the behavior tag set and the first The number of behavior types that match the standard behavior types in the templates for each teaching stage. For the first The total number of standard behavior types in each teaching stage template.

[0158] Specifically, the behavior tag set is a collection of all behavior tags whose start time falls within the feature point time window, extracted from the classroom event sequence. The teaching stage template is a standard teaching stage description pre-existing in the behavior semantic library. Each template contains multiple standard behavior types and their corresponding preset weight values. For each matching behavior tag, its corresponding behavior type has a preset weight value in the teaching stage template. This weight value is set by experts and stored in the behavior semantic library when the template is created. The actual number of times the behavior type corresponding to each matching behavior tag appears within the feature point time window is obtained by statistically analyzing the total number of times that behavior type appears within the feature point time window in the classroom event sequence. The time decay coefficient is a preset positive number used to control the rate of time decay; this coefficient is preset according to the actual needs of the teaching and research activities. The occurrence timestamp of each matching behavior tag is directly obtained from the record of that behavior tag in the classroom event sequence. The time position of the rhythm feature point is the specific moment corresponding to the feature point extracted from the composite teaching activity curve. The total duration of the feature point time window is calculated from the boundary time difference extending forward and backward from the feature point. The preset weight value of each standard behavior type is also stored in the template.

[0159] Furthermore, the significance of the calculation formula lies in quantifying the similarity between the behavior tag set and the template for a specific teaching stage. This formula calculates a value between zero and one by comprehensively considering the weight, frequency of occurrence, temporal proximity, and type coverage of the matching behavior tags. The larger the value, the higher the degree of matching between the behavior tag set and the template for that teaching stage. The numerator is a weighted sum for each matching behavior tag, where the contribution of each matching behavior tag is determined by its preset weight value, actual frequency of occurrence, and a time decay factor. The time decay factor is calculated based on the proximity of the occurrence time of the behavior tag to the temporal location of the rhythmic feature point. The closer the occurrence time is to the rhythmic feature point, the closer the factor is to one, and the greater the contribution; the farther the occurrence time is from the rhythmic feature point, the smaller the factor is, and the smaller the contribution. The denominator is the first... The sum of the preset weights of all standard behavior types in each teaching stage template is used to normalize the numerator, eliminating the influence of differences in the total weights of different templates. The entire fraction is multiplied by the ratio of the number of matched behavior types to the total number of standard behavior types in the template. This ratio reflects the completeness of the behavior tag set in terms of type coverage of the template. The more types covered, the larger the ratio, and the higher the matching degree.

[0160] In general, when more behavior labels in the behavior label set match the template's standard behavior types, the number of terms in the numerator increases, and the matching degree improves. The larger the preset weight value of the matching behavior label, or the more times it actually appears within the feature point's time window, the larger the value of the corresponding term in the numerator, and the higher the matching degree. The closer the occurrence time of the matching behavior label is to the time position of the rhythm feature point, the closer the time decay factor is to one, the larger the value of the corresponding term in the numerator, and the higher the matching degree. The higher the proportion of behavior types covered by the matching behavior labels to the total number of types in the template, i.e., the more comprehensive the matching types, the larger the final ratio, and the higher the matching degree. Conversely, if the number of matching behavior labels is small, the weight is low, the frequency of occurrence is low, the time is far from the feature point, or the coverage of types is incomplete, the matching degree will decrease. Therefore, this formula can effectively identify the teaching stage template that best matches the teaching behaviors actually occurring within the feature point's time window.

[0161] In this embodiment of the invention, when using the rhythm analysis graph as the time axis framework and segmentation basis, and embedding the chain annotations into the target time interval of the rhythm analysis graph according to the pointer direction to obtain the comprehensive analysis report of the teaching and research activity, it is specifically used for:

[0162] The time pointer labels of the chained annotations are compared with the temporal axis of the rhythm analysis graph to determine the target temporal interval to which the annotation nodes in the rhythm analysis graph belong.

[0163] Following the logical pointer direction of the chained annotations, the annotation nodes are sequentially filled into the annotation display sub-area of ​​the target time interval to obtain a comprehensive analysis report of the teaching and research activities.

[0164] Specifically, the system reads the time pointer tag from the data structure of each chained annotation. The time pointer tag is a global time anchor point recorded when the chained annotation is generated. This anchor point is accurate to the millisecond level and represents the teaching and research activity moment corresponding to the generation of the annotation. The rhythm analysis map contains a complete time axis. Each scale on the time axis corresponds to the time point of the audio text of the teaching and research activity. The time axis has been divided into multiple continuous time intervals by rhythm feature points. Each time interval is defined by a start time and an end time. The start time is the time position of the previous rhythm feature point or the beginning point of the audio text, and the end time is the time position of the next rhythm feature point or the end point of the audio text.

[0165] Specifically, the system uses a rhythm analysis graph as its basic framework. Within each time interval of the rhythm analysis graph, a sub-region for displaying annotations is pre-defined. This sub-region is used to centrally display all annotation nodes belonging to that time interval. The system traverses each time interval to obtain all chained annotations corresponding to that time interval. These chained annotations have already had their affiliation determined through comparison in the previous step.

[0166] Furthermore, the system iterates through all chained annotations. For each chained annotation, its time pointer label is compared with each time interval on the temporal axis of the rhythm analysis graph. It determines whether the time pointer label is greater than or equal to the start time of a certain time interval and less than or equal to the end time of that time interval. When a time interval that meets this condition is found, that time interval is determined as the target time interval to which the chained annotation belongs. The system records the correspondence between each chained annotation and its target time interval. At this point, all chained annotations have found the time period to which they should belong in the rhythm analysis graph.

[0167] Furthermore, the system reads the logical pointer directions from these chained annotations. These logical pointer directions are the natural order formed during the generation of the chained annotations based on their pointer relationships with the audio and video streams. Specifically, each chained annotation may contain pointers to other annotations, or be arranged in chronological order. The system sorts these chained annotations according to their logical pointer directions. After sorting, the system sequentially fills the sorted annotation nodes into the annotation display sub-area of ​​the current time interval. During filling, each annotation node is presented with its original text content, while retaining its built-in pointers to the audio and video segments. After all time intervals are filled, each time segment of the rhythm analysis graph is accompanied by sequentially arranged chained annotations. At this point, the entire document becomes a comprehensive analysis report of the teaching and research activities.

[0168] In summary, by utilizing the global time anchors stored in the chained annotations and precisely matching them with the temporal axes already divided by rhythmic feature points in the rhythm analysis graph, the system identifies the appropriate teaching stage interval for each chained annotation. This comparison mechanism categorizes the annotation texts, which were originally scattered across the time axis, according to teaching stages, ensuring that each annotation falls within a specific semantic interval such as the introduction stage, lecture stage, interaction stage, or summary stage. When the annotation's time pointer falls between the start and end times of a temporal interval, the annotation is identified as an annotation node for that interval. This time anchor-based attribution ensures that the alignment of annotations with teaching stages is completely objective and accurate, avoiding potential biases caused by subjective division. Through this operation, the system provides a clear structural framework for subsequent organization of annotation content by stage, with each annotation node obtaining its specific positional identifier within the teaching rhythm.

[0169] In summary, based on the logical order established by pointers between chained annotations, the annotation nodes within each target time interval are sorted according to their inherent relationships and then sequentially filled into the annotation display sub-area reserved for that interval in the rhythm analysis graph. The direction of the logical pointers may be based on the generation time order of the annotations or the order of mutual references between annotations. This sorting method ensures that the presentation order of the annotations is consistent with the logical context that occurs during the teaching process. After filling, each teaching stage in the rhythm analysis graph is accompanied by all the annotations generated in that stage, and the annotations are arranged in their logical order, forming a complete feedback record within the stage. The final rhythm analysis graph and the orderly arranged chained annotations below together constitute a comprehensive analysis report. This report not only shows the fluctuations in the teaching rhythm but also simultaneously presents the specific feedback content of the participants in each teaching stage, integrating quantitative teaching rhythm analysis with qualitative interpretation of commentary texts, providing teaching researchers with complete teaching review materials that are both holistic and detailed.

[0170] Compared with the prior art, the present invention has the following beneficial effects:

[0171] 1. By establishing a unified pointer association mechanism for audio streams, video streams, and comment text, deep fusion and mutual indexing of multimodal data are achieved. During the data acquisition phase, channel association seeds ensure that all data belongs to the same teaching and research activity. In the data processing process, downlink pointers, uplink pointers, source backtracking pointers, and target mapping pointers are constructed layer by layer, ultimately forming a structured teaching and research dataset. This data organization allows any comment to be directly located to its corresponding audio segment and behavioral label. When reviewing teaching and research activities, users can freely jump between audio, video, and text along the pointer direction, greatly improving the comprehensive backtracking efficiency and ease of use of multimodal information.

[0172] 2. By performing time-series segmentation and multi-dimensional numerical fusion on the teaching and research dataset, a composite teaching activity curve is constructed, and a rhythm analysis map with teaching semantic labels is further generated. The composite calculation based on behavior frequency and speech transition frequency comprehensively reflects the dynamic changes in teaching activity, while the matching of rhythm feature points with the behavioral semantic database transforms the numerical curve into a teaching stage division with educational implications. Chained annotations are finally categorized and filled according to the time intervals of the rhythm analysis map. The generated comprehensive analysis report presents both quantitative fluctuations in teaching rhythm and integrates qualitative feedback content corresponding to each stage, providing a complete interpretation path from process data to teaching principles for teaching and research activities.

[0173] like Figure 2 The diagram shown is a functional block diagram of a smart teaching and research activity management system that supports multimodal input, provided in an embodiment of the present invention.

[0174] The intelligent teaching and research activity management system 100 supporting multimodal input described in this invention can be installed in an electronic device. Depending on the functions implemented, the intelligent teaching and research activity management system 100 supporting multimodal input may include a data acquisition module 101, a chained annotation module 102, a teaching and research dataset module 103, a composite teaching activity curve module 104, a rhythm analysis graph module 105, and a comprehensive analysis report module 106. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and can perform fixed functions, and are stored in the memory of the electronic device.

[0175] In this embodiment, the functions of each module / unit are as follows:

[0176] The data acquisition module collects the audio stream, video stream, and comment text of the teaching and research activities based on the user time stamps of the teaching and research activities.

[0177] The chained annotation module writes the comment text into pointer fields pointing to the speech segment index in the audio stream and the behavior tag index in the video stream, thereby obtaining the chained annotation of the teaching and research activity;

[0178] The teaching and research dataset module establishes mutually pointing pointers between the audio stream, the video stream, and the chained annotations based on the time identifier of the teaching and research activity, and encapsulates the data streams after the pointers are established to obtain the teaching and research dataset of the teaching and research activity.

[0179] The composite teaching activity curve module uses the time period of the audio text in the teaching and research dataset as the horizontal axis and the composite value generated by fusing the behavior frequency and the number of speech transitions in the teaching and research dataset during the time period as the vertical axis to construct the composite teaching activity curve of the teaching and research activity.

[0180] The rhythm analysis graph module uses the peaks, troughs, and slope change points of the composite teaching activity curve as rhythm feature points, and matches the behavioral labels corresponding to the rhythm feature points with the behavioral semantic database of the teaching and research activity to obtain the rhythm analysis graph of the teaching and research activity.

[0181] The comprehensive analysis report module uses the rhythm analysis graph as the time axis framework and segmentation basis, and embeds the chain annotations into the target time interval of the rhythm analysis graph according to the pointer direction to obtain the comprehensive analysis report of the teaching and research activities.

[0182] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0183] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0184] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0185] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0186] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A smart teaching and research activity management method supporting multimodal input, characterized in that, The method includes: Based on the user time stamps of the teaching and research activities, the audio stream, video stream, and comment text of the teaching and research activities are collected respectively. Based on the global time anchor point of the comment text, in the audio stream's speech segments and the video stream's classroom event sequence, find the target speech segments and target behavior tags whose start time is less than or equal to the global time anchor point and whose end time is greater than or equal to the global time anchor point, and obtain the speech segment index and behavior tag index of the teaching and research activity; Write the audio segment index and the behavior tag index into the audio pointer field and behavior pointer field of the comment text to obtain the chained annotation of the teaching and research activity; Based on the time marker of the teaching and research activity, pointers are established between the audio stream, the video stream and the chained annotations, and the data streams after the pointers are established are encapsulated to obtain the teaching and research dataset of the teaching and research activity. Using the time period of the speech text in the teaching and research dataset as the horizontal axis and the composite value generated by fusing the behavior frequency and the number of speech transitions in the teaching and research dataset within the time period as the vertical axis, a composite teaching activity curve of the teaching and research activity is constructed. The peaks, troughs, and slope change points of the composite teaching activity curve are used as rhythm feature points. The behavioral labels corresponding to these rhythm feature points are matched with the behavioral semantic database of the teaching and research activity to obtain a rhythm analysis map of the teaching and research activity, including: The local maxima and local minima of the composite teaching activity curve are taken as peaks and troughs, the points where the sign of the first derivative changes in the composite teaching activity curve are taken as slope change points, and the peaks, troughs and slope change points are taken as rhythm feature points. Taking the time position of the rhythm feature point on the time axis of the audio text of the teaching and research activity as the center, extend forward to the position of the preceding rhythm feature point and backward to the position of the subsequent rhythm feature point. If there is no rhythm feature point before the time position, extend forward to the starting point of the audio text in the teaching and research activity. If there is no rhythm feature point after the time position, extend backward to the ending point of the audio text in the teaching and research activity to obtain the feature point time window of the teaching and research activity. Extract behavioral labels whose starting time is within the time window of the feature point from the classroom event sequence of the teaching and research activities to obtain a set of behavioral labels corresponding to the rhythm feature point; Calculate the matching degree between the set of behavioral labels and the teaching stage templates in the behavioral semantic database of teaching and research activities, and take the stage name of the teaching stage template corresponding to the maximum value of the matching degree as the teaching semantic label of the rhythm feature point; The teaching semantic tags are labeled to the corresponding rhythm feature points in the composite teaching activity curve to obtain the rhythm analysis map of the teaching and research activity. Using the rhythm analysis graph as the time axis framework and segmentation basis, the chain annotations are embedded into the target time interval of the rhythm analysis graph according to the pointer direction to obtain a comprehensive analysis report of the teaching and research activities.

2. The intelligent teaching and research activity management method supporting multimodal input as described in claim 1, characterized in that, The user time stamp based on the teaching and research activities collects the audio stream, video stream, and comment text of the teaching and research activities, including: The creation time and user identifier of the teaching and research activity are used as channel association seeds and written into the header information field of the acquisition link corresponding to the teaching and research activity to obtain the audio acquisition channel, video acquisition channel and text acquisition channel of the teaching and research activity. The audio stream, video stream, and comment text of the teaching and research activities are collected through the acquisition channel, the video acquisition channel, and the text acquisition channel.

3. The intelligent teaching and research activity management method supporting multimodal input as described in claim 1, characterized in that, Based on the time identifier of the teaching and research activity, pointers are established between the audio stream, the video stream, and the chained annotations, and the data streams after the pointers are established are encapsulated to obtain the teaching and research dataset of the teaching and research activity, including: Based on the time interval of the speech segment in the audio stream, the behavior labels and corresponding behavior label indices located within the time interval are traversed in the classroom event sequence of the video stream, and a downlink pointer pointing to the behavior label index is created in the data structure of the speech segment. Based on the time interval of the behavior label, find the chained comments whose synchronization time point is within the time interval and their corresponding comment identifiers in the chained comments, and create an uplink pointer to the comment identifier in the data structure of the behavior label; A source backtracking pointer is established pointing to the corresponding speech segment in the audio stream based on the speech segment index in the speech pointer field, and a target mapping pointer is established pointing to the corresponding behavior label in the classroom event sequence based on the behavior label index in the behavior pointer field. The voice text, the classroom event sequence, and the chained annotations are stored together. The downlink pointer, the uplink pointer, the source backtracking pointer, and the target mapping pointer are written into the storage structure of the teaching and research activity as the association between the voice text, the classroom event sequence, and the chained annotations, thus obtaining the teaching and research dataset of the teaching and research activity.

4. The intelligent teaching and research activity management method supporting multimodal input as described in claim 1, characterized in that, The composite teaching activity curve for the teaching and research activity is constructed by using the time period of the speech text in the teaching and research dataset as the horizontal axis and the composite value generated by fusing the corresponding values ​​of the behavior frequency and the number of speech transitions in the teaching and research dataset within the time period as the vertical axis. This includes: The total duration of the audio text in the teaching and research dataset is divided into continuous time periods, and the total number of behavior labels whose start time is located in the continuous time period in the classroom event sequence in the teaching and research dataset is taken as the behavior frequency of the continuous time period. The number of times the speaker identifier of the speech text in the teaching and research dataset changes between adjacent speech segments is taken as the number of speech transitions in the continuous time period. Using the continuous time period as the horizontal axis value and the frequency of the row as the left vertical axis value, the left data point corresponding to the continuous time period is obtained; Using the continuous time period as the horizontal axis value and the number of speech transitions as the right vertical axis value, the right data point corresponding to the continuous time period is obtained; By connecting all the left and right data points in chronological order, the frequency curve and interaction conversion curve of the teaching and research activities are obtained. The behavior frequency curve and the interaction conversion curve are integrated into a composite teaching activity curve for the teaching and research activities.

5. The intelligent teaching and research activity management method supporting multimodal input as described in claim 4, characterized in that, The process of integrating the behavior frequency curve and the interaction conversion curve into a composite teaching activity curve for the teaching and research activities includes: The frequency of behaviors in each time period of the behavior frequency curve and the number of speech transitions in each time period of the interaction transition curve are normalized to obtain the frequency scale value and transition scale value of the corresponding time period in the teaching and research activity. The Euclidean distance between the frequency scale value and the transformation scale value is calculated based on the same time period of the teaching and research activities to obtain the composite teaching index of the teaching and research activities. Using the continuous time period of the teaching and research activities as the horizontal axis and the composite teaching indicators as the vertical axis, the data points are connected sequentially in chronological order to obtain the composite teaching activity curve of the teaching and research activities.

6. The intelligent teaching and research activity management method supporting multimodal input as described in claim 1, characterized in that, The formula for calculating the matching degree includes: in, For the set of behavior labels and the first The matching degree of the template for each teaching stage. This serves as an index for the template of the aforementioned teaching stage. For the behavior type in the behavior tag set and the first The total number of behavior tags that match the standard behavior types in each teaching stage template. The sequence number of the behavior tag in the behavior tag set that matches the teaching stage template. For the first behavior tag set The behavior type corresponding to each matching behavior tag in the first... The preset weight values ​​corresponding to each teaching stage template For the first behavior tag set The actual number of times the behavior type corresponding to each matching behavior label appears within the time window of the feature point. It is a natural constant. The time decay coefficient, For the first behavior tag set The timestamp of each matching behavior tag. The time position of the rhythmic feature point. The total duration of the time window for the feature points. For the first The total number of all standard behavior types in each teaching stage template This serves as an index for the standard behavior types in the teaching stage template. For the first In the first teaching stage template Preset weight values ​​for each standard behavior type For the behavior type in the behavior tag set and the first The number of behavior types that match the standard behavior types in the templates for each teaching stage. For the first The total number of standard behavior types in each teaching stage template.

7. The intelligent teaching and research activity management method supporting multimodal input as described in claim 1, characterized in that, Using the rhythm analysis graph as the timeline framework and segmentation basis, the chained annotations are embedded into the target time interval of the rhythm analysis graph according to the pointer direction, resulting in a comprehensive analysis report of the teaching and research activity, including: The time pointer labels of the chained annotations are compared with the temporal axis of the rhythm analysis graph to determine the target temporal interval to which the annotation nodes in the rhythm analysis graph belong. Following the logical pointer direction of the chained annotations, the annotation nodes are sequentially filled into the annotation display sub-area of ​​the target time interval to obtain a comprehensive analysis report of the teaching and research activities.

8. A smart teaching and research activity management system supporting multimodal input, used to implement the smart teaching and research activity management method supporting multimodal input as described in any one of claims 1-7, characterized in that, The system includes: The data acquisition module collects the audio stream, video stream, and comment text of the teaching and research activities based on the user time stamps of the teaching and research activities. The chained annotation module, based on the global time anchor point of the comment text, searches for target speech segments and target behavior tags in the speech segments of the audio stream and the classroom event sequence of the video stream, with a start time less than or equal to the global time anchor point and an end time greater than or equal to the global time anchor point, to obtain the speech segment index and behavior tag index of the teaching and research activity; Write the audio segment index and the behavior tag index into the audio pointer field and behavior pointer field of the comment text to obtain the chained annotation of the teaching and research activity; The teaching and research dataset module establishes mutually pointing pointers between the audio stream, the video stream, and the chained annotations based on the time identifier of the teaching and research activity, and encapsulates the data streams after the pointers are established to obtain the teaching and research dataset of the teaching and research activity. The composite teaching activity curve module uses the time period of the audio text in the teaching and research dataset as the horizontal axis and the composite value generated by fusing the behavior frequency and the number of speech transitions in the teaching and research dataset during the time period as the vertical axis to construct the composite teaching activity curve of the teaching and research activity. The rhythm analysis graph module uses the peaks, troughs, and slope change points of the composite teaching activity curve as rhythm feature points. It matches the behavioral labels corresponding to these rhythm feature points with the behavioral semantic database of the teaching and research activity to obtain the rhythm analysis graph of the teaching and research activity, including: The local maxima and local minima of the composite teaching activity curve are taken as peaks and troughs, the points where the sign of the first derivative changes in the composite teaching activity curve are taken as slope change points, and the peaks, troughs and slope change points are taken as rhythm feature points. Taking the time position of the rhythm feature point on the time axis of the audio text of the teaching and research activity as the center, extend forward to the position of the preceding rhythm feature point and backward to the position of the subsequent rhythm feature point. If there is no rhythm feature point before the time position, extend forward to the starting point of the audio text in the teaching and research activity. If there is no rhythm feature point after the time position, extend backward to the ending point of the audio text in the teaching and research activity to obtain the feature point time window of the teaching and research activity. Extract behavioral labels whose starting time is within the time window of the feature point from the classroom event sequence of the teaching and research activities to obtain a set of behavioral labels corresponding to the rhythm feature point; Calculate the matching degree between the set of behavioral labels and the teaching stage templates in the behavioral semantic database of teaching and research activities, and take the stage name of the teaching stage template corresponding to the maximum value of the matching degree as the teaching semantic label of the rhythm feature point; The teaching semantic tags are labeled to the corresponding rhythm feature points in the composite teaching activity curve to obtain the rhythm analysis map of the teaching and research activity. The comprehensive analysis report module uses the rhythm analysis graph as the time axis framework and segmentation basis, and embeds the chain annotations into the target time interval of the rhythm analysis graph according to the pointer direction to obtain the comprehensive analysis report of the teaching and research activities.