Multi-service bearing teaching video analysis method and system based on cross-modal data

By integrating multimodal data acquisition, preprocessing and fusion technologies in the teaching video analysis system, combined with semantic analysis and knowledge graph construction, the problem of insufficient integration of multimodal data in the existing technology is solved, efficient teaching video analysis and personalized teaching suggestions are achieved, and teaching quality and learning effect are significantly improved.

CN120014510APending Publication Date: 2025-05-16SUZHOU ZHIHAI YINGHUI TECHNOLOGY CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510069637.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing teaching video analysis technology cannot effectively integrate multimodal data in videos, resulting in limited level of intelligent analysis and insufficient personalized teaching service capabilities.

Method used

Provide a multi-service carrying teaching video analysis method and system based on cross-modal data. Through modules such as data acquisition, preprocessing, cross-modal data fusion, semantic analysis, knowledge graph construction and personalized teaching suggestions generation, it realizes the integration and in-depth analysis of multi-modal data of teaching videos.

Benefits of technology

Through in-depth analysis of cross-modal data and the construction of knowledge graphs, the depth and intelligence level of teaching video analysis are significantly improved, personalized teaching suggestions are provided, and teaching quality and learning effect are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014510A_ABST
    Figure CN120014510A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-service bearing teaching video analysis method and system based on cross-modal data, and belongs to the technical field of intelligent education. According to the system, through cooperative work of the data acquisition module and the data preprocessing module, multi-modal data, including video pictures, audio signals, subtitle texts and behavior data, from teaching videos can be effectively acquired and cleaned, and synchronism and high quality of the data are ensured; and reliable data input is provided for subsequent cross-modal data fusion and knowledge graph construction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent education technology, and in particular to a multi-service carrying teaching video analysis method and system based on cross-modal data. Background Art

[0002] With the popularization of digital education, teaching videos have gradually become an important carrier of teaching.

[0003] However, existing teaching video analysis technologies cannot effectively integrate multimodal data (such as vision, audio, text, etc.) in videos, resulting in limited intelligent analysis levels and insufficient personalized teaching service capabilities.

[0004] Therefore, there is an urgent need for an innovative technology to achieve cross-modal data fusion and in-depth analysis of teaching videos, so as to improve teaching quality and learning outcomes. Summary of the invention

[0005] The embodiment of the present application provides a method and system for analyzing multi-service-bearing teaching videos based on cross-modal data. The technical solution is as follows:

[0006] According to one aspect of the present application, a multi-service bearing teaching video analysis system based on cross-modal data is provided, the system comprising:

[0007] The data acquisition module is used to collect multimodal data in teaching videos, including video images, audio, subtitle text and user interaction behavior data;

[0008] Data preprocessing module, used to clean, align and standardize the collected multimodal data;

[0009] A cross-modal data fusion module, used to realize feature extraction and cross-modal fusion of the multi-modal data using a deep learning model;

[0010] Teaching video semantic analysis module, which is used to perform semantic recognition and analysis of knowledge points in teaching videos based on cross-modal fusion data;

[0011] The course knowledge graph construction module is used to generate the course knowledge graph based on the identified knowledge points and their associations;

[0012] Personalized teaching suggestion generation module, which is used to generate personalized teaching suggestions based on learners' behavior data, knowledge graphs, and learning preferences;

[0013] The continuous optimization module is used to dynamically optimize the system's models, parameters and strategies to adapt to changing teaching needs.

[0014] Optionally, the data acquisition module includes:

[0015] A video processing unit for extracting video frames and their visual features;

[0016] An audio processing unit, used for extracting audio signals and their features;

[0017] Subtitle parsing unit, used to extract subtitle text from the video and perform semantic analysis;

[0018] The behavior data collection unit is used to record the learners' interactive behavior data in the teaching video, including clicks, pauses and replays.

[0019] Optionally, in working state, the data preprocessing module includes:

[0020] Dynamic cleaning algorithm unit, used to identify and remove invalid or low-quality data;

[0021] The time series synchronization unit is used to align timestamps and normalize sequences of multimodal data.

[0022] Optionally, the cross-modal data fusion module includes:

[0023] Visual feature extraction model, used to extract edge and texture visual features from video frames;

[0024] Audio feature extraction model, used to extract pitch, frequency, and speech rate information from audio;

[0025] Text feature extraction model, used to extract keywords and semantic information from subtitle text;

[0026] Cross-modal fusion model realizes the fusion of multimodal data based on attention mechanism or feature splicing method.

[0027] Optionally, the course knowledge graph construction module includes:

[0028] Knowledge point extraction unit, used to identify core knowledge points in the video;

[0029] Knowledge point association analysis unit, used to explore the sequence, causal relationship and parallel relationship between knowledge points;

[0030] The knowledge graph generation unit generates a course knowledge graph based on knowledge points and their associations, and displays it through visualization technology.

[0031] According to another aspect of the present application, a method for analyzing a multi-service-bearing teaching video based on cross-modal data is provided. The method is used in the multi-service-bearing teaching video analysis system based on cross-modal data described above. The method includes:

[0032] The multimodal data in the teaching video is obtained through the data acquisition module, including video frames, audio signals, subtitle texts and interactive behavior data;

[0033] Using a time series synchronization algorithm to achieve alignment and standardization of the multimodal data;

[0034] Extract visual, audio, and text features from the multimodal data based on the deep learning model, and achieve cross-modal data fusion through a multimodal attention mechanism;

[0035] The fused data is semantically understood through a large model to identify knowledge points in the teaching video and their relationships.

[0036] Generate a knowledge graph based on the identified knowledge points and their associations, and display it through visualization technology;

[0037] Combine learners' behavior data and knowledge graphs to generate personalized learning paths and resource recommendations for learners;

[0038] Incremental training and parameter adjustment of large models based on new data and user feedback to improve analysis results.

[0039] Optionally, extracting visual, audio, and text features from the multimodal data based on the deep learning model, and implementing cross-modal data fusion through a multimodal attention mechanism, includes:

[0040] Extract key visual features from teaching video images through visual feature extraction model;

[0041] Extract the intonation and frequency features in the teaching video through the audio feature extraction model;

[0042] Extract semantic information of subtitle text through natural language processing model;

[0043] Unified encoding and representation of multimodal data is achieved through cross-modal attention mechanism.

[0044] Optionally, the combining of learner behavior data and knowledge graph to generate personalized learning paths and resource recommendations for learners includes:

[0045] Build user profiles based on learners’ behavioral data;

[0046] Utilize the nodes and association information in the knowledge graph to plan learning paths for learners and recommend relevant learning resources.

[0047] According to another aspect of the present application, a computer-readable storage medium is provided, which stores at least one instruction, and the at least one instruction is used to be executed by a processor to implement the multi-service carrying teaching video analysis method based on cross-modal data as described in the above aspect.

[0048] In the embodiment of the present application, through the collaborative work of the data acquisition module and the data preprocessing module, the present invention can effectively acquire and clean the multimodal data from the teaching video, including video images, audio signals, subtitle texts and behavioral data, thereby ensuring the synchronization and high quality of the data, and providing reliable data input for subsequent cross-modal data fusion and knowledge graph construction. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 A schematic diagram of the structure of a multi-service bearing teaching video analysis system based on cross-modal data provided by an exemplary embodiment of the present invention;

[0050] Figure 2 A flowchart of a multi-service bearing teaching video analysis method based on cross-modal data is provided for an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0051] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.

[0052] The term "multiple" as used herein refers to two or more than two. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the related objects are in an "or" relationship.

[0053] Example 1

[0054] Please refer to Figure 1 , which shows a schematic diagram of the structure of a multi-service bearing teaching video analysis system based on cross-modal data provided by an exemplary embodiment of the present application. The system includes:

[0055] The data acquisition module is used to collect multimodal data in teaching videos, including video images, audio, subtitle text and user interaction behavior data.

[0056] Optionally, the data acquisition module includes: a video processing unit for extracting video frames and their visual features; an audio processing unit for extracting audio signals and their features; a subtitle parsing unit for extracting subtitle text from the video and performing semantic analysis; and a behavior data acquisition unit for recording learners' interactive behavior data in the teaching video, including clicks, pauses, and replays.

[0057] The video processing unit uses computer vision technology (such as the OpenCV library) to extract image information frame by frame from the teaching video and identify the scene, characters and graphic features in the video. At the same time, it extracts key visual features, including object edges, textures and colors, through pre-trained deep learning models (such as YOLO or ResNet).

[0058] Among them, the audio processing unit uses an audio processing library (such as Librosa) to sample, segment and extract features of the audio signals in the teaching video, extracting information including intonation, frequency, speaking speed, etc. for audio semantic understanding and analysis.

[0059] Among them, the subtitle parsing unit uses optical character recognition technology (OCR) to extract subtitle text in the video, and uses a natural language processing model (such as BERT) to perform semantic analysis on the subtitle text to obtain the core keywords and concepts in the text.

[0060] The behavior data collection unit integrates the API interface of the online platform to record the interactive behavior data of learners in the video (such as clicks, pauses and replays). The behavior data is extracted through log analysis or database query to analyze learning preferences and behavior patterns.

[0061] The data collection module realizes the unified management of multimodal data sources by synchronously collecting video, audio, subtitles and behavioral data, ensuring the consistency of timestamps and content integrity of each modality. This module is based on multimodal data collection technology and provides comprehensive basic data support for subsequent analysis.

[0062] The data preprocessing module is used to clean, align and standardize the collected multimodal data.

[0063] Optionally, in working state, the data preprocessing module includes: a dynamic cleaning algorithm unit for identifying and eliminating invalid or low-quality data; a time series synchronization unit for performing timestamp alignment and sequence standardization on multimodal data.

[0064] The dynamic cleaning algorithm unit uses a rule-based dynamic cleaning algorithm and anomaly detection techniques (such as cluster analysis and outlier detection based on statistical distribution) to automatically remove blurry images in videos, background noise in audio, and redundant or inaccurate content in subtitles. In one example, Fourier transform is applied to audio data to analyze the spectrum information and filter out low-quality noise segments.

[0065] Among them, the time series synchronization unit uses a timestamp alignment algorithm (such as the dynamic time warping algorithm DTW) to align and standardize multimodal data according to the time dimension. For example, the subtitle text and video frame are aligned through timestamps to achieve a one-to-one correspondence between the content; audio features are matched with image features to ensure synchronization of audio and video data.

[0066] Therefore, data cleaning and time series synchronization ensure the accuracy, synchronization and quality of multimodal data before analysis through technical processing of data integrity and consistency, laying a solid foundation for subsequent cross-modal data fusion.

[0067] The cross-modal data fusion module is used to realize feature extraction and cross-modal fusion of multimodal data using deep learning models.

[0068] Optionally, the cross-modal data fusion module includes: a visual feature extraction model for extracting edge and texture visual features from video frames; an audio feature extraction model for extracting pitch, frequency, and speaking speed information from audio; a text feature extraction model for extracting keywords and semantic information from subtitle text; and a cross-modal fusion model for realizing the fusion of multimodal data based on an attention mechanism or a feature splicing method.

[0069] The teaching video semantic analysis module is used to perform semantic recognition and analysis of knowledge points in teaching videos based on cross-modal fusion data.

[0070] The course knowledge graph construction module is used to generate the course knowledge graph based on the identified knowledge points and their associations.

[0071] Optionally, the course knowledge graph construction module includes: a knowledge point extraction unit, used to identify core knowledge points in the video; a knowledge point association analysis unit, used to explore the sequence, causal relationship and parallel relationship between knowledge points; a knowledge graph generation unit, which generates a course knowledge graph based on knowledge points and their associations, and displays it through visualization technology.

[0072] The personalized teaching suggestion generation module is used to generate personalized teaching suggestions based on learners' behavioral data, knowledge graphs, and learning preferences.

[0073] The continuous optimization module is used to dynamically optimize the system's models, parameters, and strategies to adapt to changing teaching needs.

[0074] In summary, in the embodiments of the present application, through the collaborative work of the data acquisition module and the data preprocessing module, the present invention can effectively acquire and clean the multimodal data from the teaching video, including video images, audio signals, subtitle texts and behavioral data, ensuring the synchronization and high quality of the data, and providing reliable data input for subsequent cross-modal data fusion and knowledge graph construction.

[0075] Example 2

[0076] According to another aspect of the present application, Figure 2 As shown, a method for analyzing teaching videos with multiple services carried based on cross-modal data is provided. The method is used in the above-mentioned teaching video analysis system with multiple services carried based on cross-modal data. The method includes:

[0077] Step 201, obtaining multimodal data in the teaching video through a data acquisition module, including video frames, audio signals, subtitle texts and interactive behavior data.

[0078] Multimodal data is extracted from the teaching video through the data acquisition module in Example 1. Video frame information is extracted using a video processing library (such as OpenCV), audio signals are extracted using Librosa, subtitles are parsed using OCR, and user behavior data is collected using an API.

[0079] Among them, multimodal data acquisition is based on computer vision, speech processing and natural language processing technologies, which realizes the efficient acquisition of different modal data sources and provides multi-dimensional input for subsequent processing.

[0080] Step 202: aligning and standardizing multimodal data using a time series synchronization algorithm.

[0081] Use the dynamic time warping (DTW) algorithm to calibrate timestamps to ensure temporal consistency among video frames, audio clips, and subtitle segments. Use normalization techniques (such as Min-Max normalization or Z-Score normalization) to standardize the collected behavioral data and feature values ​​so that data from different modalities can be fused at the same scale.

[0082] Among them, through time series modeling and normalization algorithms, the inconsistency and feature scale differences of multimodal data in the time dimension are solved, laying the foundation for the unified representation of cross-modal data.

[0083] Step 203: extract visual, audio, and text features from the multimodal data based on the deep learning model, and achieve cross-modal data fusion through a multimodal attention mechanism.

[0084] In one possible implementation, key visual features in the teaching video screen are extracted through a visual feature extraction model, intonation and frequency features in the teaching video are extracted through an audio feature extraction model, semantic information of the subtitle text is extracted through a natural language processing model, and unified encoding and representation of multimodal data are achieved through a cross-modal attention mechanism.

[0085] In one example, a deep learning model (such as ResNet) is used to extract features from video frames, extracting visual features such as edges and textures; Librosa is used to frame the audio signal and extract features such as Mel-frequency cepstral coefficients (MFCC) and volume; a pre-trained natural language processing model (such as BERT or RoBERTa) is used to extract semantic information from subtitle text, including keywords and conceptual relationships; and a cross-modal attention mechanism is used to achieve a unified representation of the data by weighted integration of features from each modality.

[0086] Among them, the deep learning model is used to extract features from video, audio and text respectively, and a unified representation of data is achieved based on the cross-modal attention mechanism or feature splicing strategy, ensuring the efficient fusion of multimodal information.

[0087] Step 204, semantic understanding of the fused data is performed through a large model to identify knowledge points in the teaching video and the relationships between them.

[0088] In one possible implementation, a large model (such as Transformer) is used to perform semantic analysis on the fused data to identify the core knowledge points in the teaching video; and an association analysis algorithm (such as association rule mining or graph algorithm) is applied to analyze the sequence, causal relationship, and parallel relationship between knowledge points.

[0089] Among them, the semantic understanding and association analysis technology based on the deep learning model is used to mine the semantic and structured information of knowledge points in teaching videos.

[0090] Step 205, generate a knowledge graph based on the identified knowledge points and their associations, and display it through visualization technology.

[0091] In one possible implementation, the results of knowledge point recognition are used to construct a course knowledge graph, which is stored and managed based on a graph database (such as Neo4j), and a graph visualization tool (such as D3.js or Gephi) is used to generate an intuitive knowledge graph view.

[0092] Among them, knowledge points and their relationships are stored in a graph database, and the graph structure is optimized based on graph algorithms to display the logical connections and learning paths between knowledge in a visual way.

[0093] Step 206, combining the learner's behavior data and knowledge graph to generate personalized learning paths and resource recommendations for the learner.

[0094] In one possible implementation, a user profile is constructed based on the learner's behavioral data, and the nodes and associated information in the knowledge graph are used to plan a learning path for the learner and recommend relevant learning resources.

[0095] For example, build user portraits based on knowledge graphs, analyze learners' learning preferences and behavior patterns, use collaborative filtering algorithms or content-based recommendation algorithms, and combine learner portraits and knowledge graphs to generate personalized learning paths and resource recommendation lists.

[0096] Among them, the association information and behavioral data of the knowledge graph are combined, and the recommendation algorithm is used to achieve targeted resource recommendation and learning path planning.

[0097] Step 207, incremental training and parameter adjustment of the large model based on the newly added data and user feedback to improve the analysis effect.

[0098] In one possible implementation, new teaching video data and learner feedback are collected regularly, and transfer learning or incremental training techniques are used to adjust parameters and optimize performance of the large model.

[0099] Among them, through incremental training and dynamic adjustment mechanisms, the semantic understanding ability and cross-modal fusion effect of the large model are improved, maintaining the efficiency and adaptability of the system.

[0100] To sum up, the embodiments of the present application can effectively extract knowledge points and their correlations in teaching videos through unified collection, alignment, fusion and semantic analysis of multimodal data, generate intuitive course knowledge graphs, and provide learners with personalized learning paths and resource recommendations, thereby significantly improving the depth and intelligence level of teaching video analysis.

[0101] An embodiment of the present application also provides a computer-readable medium storing at least one instruction, wherein the at least one instruction is loaded and executed by the processor to implement the multi-service carrying teaching video analysis method based on cross-modal data as described in the above embodiments.

[0102] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A multi-service carrying teaching video analysis system based on cross-modal data, characterized in that: The system comprises: The data acquisition module is used to collect multimodal data in teaching videos, including video images, audio, subtitle text and user interaction behavior data; Data preprocessing module, used to clean, align and standardize the collected multimodal data; A cross-modal data fusion module, used to realize feature extraction and cross-modal fusion of the multi-modal data using a deep learning model; Teaching video semantic analysis module, which is used to perform semantic recognition and analysis of knowledge points in teaching videos based on cross-modal fusion data; The course knowledge graph construction module is used to generate the course knowledge graph based on the identified knowledge points and their associations; Personalized teaching suggestion generation module, which is used to generate personalized teaching suggestions based on learners' behavior data, knowledge graphs, and learning preferences; The continuous optimization module is used to dynamically optimize the system's models, parameters, and strategies to adapt to changing teaching needs.

2. The system according to claim 1, characterized in that The data acquisition module comprises: A video processing unit for extracting video frames and their visual features; An audio processing unit, used for extracting audio signals and their features; Subtitle parsing unit, used to extract subtitle text from the video and perform semantic analysis; The behavior data collection unit is used to record the learners' interactive behavior data in the teaching video, including clicks, pauses and replays.

3. The system according to claim 1, characterized in that In the working state, the data preprocessing module includes: Dynamic cleaning algorithm unit, used to identify and remove invalid or low-quality data; The time series synchronization unit is used to align timestamps and normalize sequences of multimodal data.

4. The system according to claim 1, characterized in that The cross-modal data fusion module includes: Visual feature extraction model, used to extract edge and texture visual features from video frames; Audio feature extraction model, used to extract pitch, frequency, and speech rate information from audio; Text feature extraction model, used to extract keywords and semantic information from subtitle text; Cross-modal fusion model realizes the fusion of multimodal data based on attention mechanism or feature splicing method.

5. The system according to claim 1, characterized in that The course knowledge graph construction module includes: Knowledge point extraction unit, used to identify core knowledge points in the video; Knowledge point association analysis unit, used to explore the sequence, causal relationship and parallel relationship between knowledge points; The knowledge graph generation unit generates a course knowledge graph based on knowledge points and their associations, and displays it through visualization technology.

6. A multi-service bearing teaching video analysis method based on cross-modal data, characterized in that: The method is used in the multi-service bearing teaching video analysis system based on cross-modal data according to any one of claims 1 to 5, and the method comprises: The multimodal data in the teaching video is obtained through the data acquisition module, including video frames, audio signals, subtitle texts and interactive behavior data; Using a time series synchronization algorithm to achieve alignment and standardization of the multimodal data; Extract visual, audio, and text features from the multimodal data based on the deep learning model, and achieve cross-modal data fusion through a multimodal attention mechanism; The fused data is semantically understood through a large model to identify knowledge points in the teaching video and their relationships. Generate a knowledge graph based on the identified knowledge points and their associations, and display it through visualization technology; Combine learners' behavior data and knowledge graphs to generate personalized learning paths and resource recommendations for learners; Incremental training and parameter adjustment of large models based on new data and user feedback to improve analysis results.

7. The method according to claim 6, characterized in that The extracting visual, audio, and text features of the multimodal data based on the deep learning model, and realizing cross-modal data fusion through a multimodal attention mechanism, includes: Extract key visual features from teaching video images through visual feature extraction model; Extract the intonation and frequency features in the teaching video through the audio feature extraction model; Extract semantic information of subtitle text through natural language processing model; Unified encoding and representation of multimodal data is achieved through cross-modal attention mechanism.

8. The method according to claim 6 or 7, characterized in that: The above method combines learners' behavior data and knowledge graph to generate personalized learning paths and resource recommendations for learners, including: Build user profiles based on learners’ behavioral data; Utilize the nodes and association information in the knowledge graph to plan learning paths for learners and recommend relevant learning resources.

Citation Information

Cited By

  • Online teaching interaction method based on multi-modal knowledge graph, medium and equipment

    CN120339011A

  • Multi-modal visual analysis system for teaching

    CN120339924A

  • Image analysis method and device, equipment and medium

    CN120611775A

  • Teaching video quality evaluation method based on multi-modal semantic understanding and knowledge graph automatic construction

    CN122388411A