Education video content intelligent analysis and labeling system based on deep learning

Through the combination of deep learning, graph neural network and reinforcement learning, the automation, intelligent analysis and labeling of educational video content is realized, the problems of low efficiency and poor accuracy in the existing technology are solved, and the intelligence and management efficiency of labeling results are improved.

CN120279459AInactive Publication Date: 2025-07-08CHONGQING NORMAL UNIVERSITY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510340293.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing educational video content analysis and labeling methods rely on manual operations, are inefficient and prone to errors, and are difficult to accurately understand and label video content, and lack an intelligent automated optimization mechanism.

Method used

The intelligent analysis and labeling system of educational video content based on deep learning is adopted, combined with deep learning, graph neural network and reinforcement learning, video scene features are automatically extracted, accurate content descriptions are generated through semantic reasoning, and labeling is optimized using knowledge graphs, and continuous optimization is carried out in combination with user feedback.

Benefits of technology

It improves the automation, accuracy and intelligence level of educational video content analysis and labeling, reduces manual intervention, improves the accuracy and efficiency of labeling results, and supports the intelligent management and retrieval of educational video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279459A_ABST
    Figure CN120279459A_ABST
Patent Text Reader

Abstract

The invention provides an education video content intelligent analysis and labeling system based on deep learning, and belongs to the technical field of video analysis, and the system comprises a data processing module which is used for extracting a plurality of scene features based on preprocessed education video data through a pre-trained deep learning model; the feature screening module inputs the extracted scene features into a feature fusion network, and performs feature screening and fusion based on a preset condition to obtain fused features; the content reasoning module is used for constructing a final graph structure of the video content through a graph neural network based on the fusion features, performing semantic reasoning in combination with a preset teaching logic rule, and generating semantic description of the video content; the annotation and optimization module is used for generating a video annotation through a reinforcement learning model based on a knowledge graph and semantic description to obtain a final annotation result; and the storage and retrieval module is used for associatively storing the final labeling result and the preprocessed education video data, and constructing a retrieval index. And the accuracy and the intelligent level of the labeling result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video analysis, and particularly to an intelligent analysis and annotation system for educational video content based on deep learning. Background Art

[0002] With the rapid development of online education, educational videos have become important teaching resources. However, existing video content analysis and annotation methods rely on manual operations, which are inefficient and error-prone. Traditional video analysis systems are difficult to accurately understand and annotate video content, and lack intelligent automated optimization mechanisms. In the prior art, although there are some deep learning-based solutions, there are still problems such as low accuracy, slow processing speed, and difficulty in self-learning and optimization in feature extraction, semantic reasoning, and annotation optimization of video scenes.

[0003] The prior art cannot efficiently and accurately extract scene features in videos, and it is difficult to generate accurate semantic descriptions through intelligent reasoning. In addition, the optimization of annotation results mostly relies on manual intervention, and there is a lack of a mechanism for iterative optimization based on user feedback. Therefore, how to improve the automation, accuracy, and intelligence level of educational video content analysis and annotation remains an urgent problem to be solved.

[0004] Therefore, the present invention provides an intelligent analysis and annotation system for educational video content based on deep learning. Summary of the Invention

[0005] The present invention provides an intelligent analysis and annotation system for educational video content based on deep learning, which provides an efficient and intelligent educational video content analysis and annotation system by combining deep learning, graph neural networks, and reinforcement learning. Compared with the prior art, this system can automatically extract scene features in videos, generate accurate content descriptions based on semantic reasoning, and optimize annotations through a knowledge graph. In addition, the system is continuously optimized through user feedback, improving the accuracy and intelligence level of annotation results, solving the problems of low efficiency and poor accuracy of traditional manual annotation, and helping to accelerate the intelligent management and retrieval of educational video content.

[0006] The present invention provides an intelligent analysis and annotation system for educational video content based on deep learning, including:

[0007] A data processing module: obtaining educational video data, performing preprocessing, and extracting a plurality of scene features based on the preprocessed educational video data using a pre-trained deep learning model;

[0008] A feature screening module: inputting the extracted scene features into a feature fusion network, performing feature screening and fusion based on preset conditions to obtain fusion features;

[0009] Content Inference Module: Based on the fused features, construct the final graph structure of the video content through a graph neural network, and perform semantic reasoning in combination with preset teaching logic rules to generate a semantic description of the video content;

[0010] Annotation and Optimization Module: Based on the knowledge graph and semantic description, generate video annotations through a reinforcement learning model, and obtain user feedback data to iteratively optimize the annotation results to obtain the final annotation results;

[0011] Storage and Retrieval Module: Associatively store the final annotation results with the preprocessed educational video data, and construct a retrieval index based on the knowledge graph and semantic description.

[0012] Preferably, the data processing module includes:

[0013] Video Segmentation Unit: Split the educational video data into several sub-video segments;

[0014] Video Processing Unit: Based on the dynamic saliency detection algorithm using the optical flow method, extract several key frames of significant teaching content from each sub-video segment, and then preprocess each key frame to obtain a set of key frames;

[0015] Feature Extraction Unit: Based on the preprocessed set of key frames, use a pre-trained deep learning model to extract several features respectively.

[0016] Preferably, the video segmentation unit includes:

[0017] Initial Segmentation Sub-unit: Initially segment the educational video data according to preset initial segmentation parameters to obtain a set of initial sub-video segments;

[0018] Video Analysis Sub-unit: Analyze each initial sub-video segment in the set of initial sub-video segments to obtain the pixel difference and optical flow change data of each adjacent frame corresponding to all frames of each initial sub-video segment;

[0019] Video Frame Analysis Sub-unit: Obtain the inter-frame difference degree of each sub-video segment based on the pixel difference and optical flow change data of each adjacent frame corresponding to all frames of each initial sub-video segment;

[0020] Video Marking Sub-unit: If the inter-frame difference degree of each frame of each sub-video segment exceeds the preset inter-frame difference threshold, mark the frame with the inter-frame difference degree exceeding the preset inter-frame difference threshold as the initial segmentation point;

[0021] Feature Determination Sub-unit: For each sub-video segment in the set of initial sub-video segments, extract key frames and determine the scene features;

[0022] Fine segmentation subunit: Determine the similarity of scene features through a pre-trained convolutional neural network. If the similarity of scene features between adjacent sub-video segments in the initial set of sub-video segments is lower than the preset scene change sensitivity, insert a scene change segmentation point at this position;

[0023] Segment generation subunit: Combine the preliminary segmentation points and scene change segmentation points to generate a number of sub-video segments.

[0024] Preferably, the feature screening module includes:

[0025] Model construction unit: Construct a scene classification model according to the preset teaching scene classification rules;

[0026] Probability acquisition unit: Input all scene features into the scene classification model respectively to obtain the probability distribution of each scene feature under each teaching scene;

[0027] Score determination unit: Based on the probability distribution, obtain the scene correlation score of each scene feature through weighted summation;

[0028] Feature screening unit: Screen out the scene features with scores higher than the preset score threshold based on the scene correlation scores of each scene feature and the preset score threshold;

[0029] Knowledge graph construction unit: Construct a knowledge graph according to the preset knowledge point association rules;

[0030] Feature mapping unit: Map all scene features into the knowledge graph to determine the association degree between each scene feature and the knowledge points;

[0031] Feature detection unit: Feature consistency detection based on the semantic similarity threshold:

[0032] Matrix determination unit: Determine the semantic similarity matrix between each scene feature based on the cosine similarity algorithm;

[0033] Feature marking unit: Mark the elements in the semantic similarity matrix that are lower than the preset semantic similarity threshold as inconsistent feature pairs based on the preset semantic similarity threshold;

[0034] Feature alignment unit: Perform semantic alignment on the inconsistent feature pairs through an adversarial generation network to generate consistent features;

[0035] Weight determination unit: Input the consistent features into the attention mechanism model to determine the attention weights of each consistent feature;

[0036] Feature determination unit: Perform weighted summation on all consistent features according to the attention weights of all consistent features to obtain the preliminary fusion features;

[0037] Feature Optimization Unit: Input the preliminary fusion features into the multi-modal fusion network for optimization to generate fusion features.

[0038] Preferably, the content reasoning module includes:

[0039] Semantic Analysis Unit: Perform semantic analysis on the fusion features to obtain the semantic information of the video content and define the set of node types;

[0040] Feature Decomposition Unit: Decompose the fusion features to obtain the feature subsets corresponding to the node types in the set of node types;

[0041] Feature Mapping Unit: Map the feature subsets corresponding to the node types in the set of node types to the low-dimensional semantic space through the semantic embedding model to generate node embedding vectors;

[0042] Logical Analysis Unit: Perform logical analysis on the fusion features to obtain the logical relationships of the video content, and then define the set of edge types;

[0043] Relationship Extraction Unit: Extract relationships from the fusion features to obtain the relationship subsets corresponding to the edge types in the set of edge types;

[0044] Relationship Mapping Unit: Map the relationship subsets corresponding to the edge types in the set of edge types to the low-dimensional relationship space through the relationship embedding model to generate edge embedding vectors;

[0045] Vector Combination Unit: Combine the node embedding vectors and the edge embedding vectors to construct the initial graph structure.

[0046] Preferably, the content reasoning module further includes:

[0047] Knowledge Reasoning Unit: Perform knowledge reasoning on the initial graph structure based on the preset knowledge reasoning rules to generate the reasoned graph structure;

[0048] Structure Update Unit: Optimize the graph structure of the reasoned graph structure, define the update rules of the knowledge graph according to the dynamic changes of the video content, and then perform dynamic update on the optimized graph structure to generate the updated graph structure;

[0049] Graph Compression Unit: Compress the updated graph structure to generate the final graph structure;

[0050] Semantic Reasoning Unit: Perform semantic reasoning on the final graph structure based on the preset teaching logic rules to generate the semantic description of the video content.

[0051] Preferably, the semantic reasoning module includes:

[0052] Combined Representation Sub-Unit: Represent the preset teaching logic rules as a rule set;

[0053] Content matching subunit: Traverse the nodes and edges in the final graph structure to match the content related to the rule set;

[0054] Semantic reasoning subunit: Perform semantic reasoning on the content related to the rule set by a rule-based reasoning engine;

[0055] Description generation subunit: Generate semantic descriptions based on the semantic reasoning results and a preset description template.

[0056] Preferably, the annotation and optimization module includes:

[0057] Data acquisition unit: Acquire user feedback data;

[0058] Policy adjustment unit: Input the final graph structure and semantic descriptions into a reinforcement learning model, and the reinforcement learning model dynamically adjusts the annotation policy according to the user feedback data;

[0059] Perform several rounds of iterative optimization through the adjusted annotation policy to generate the final annotation result.

[0060] Compared with the prior art, the beneficial effects of the present application are as follows:

[0061] By combining deep learning, graph neural networks, and reinforcement learning, an efficient and intelligent educational video content analysis and annotation system is provided. Compared with the prior art, this system can automatically extract scene features in the video, generate accurate content descriptions based on semantic reasoning, and optimize annotations through a knowledge graph. In addition, the system is continuously optimized through user feedback, improving the accuracy and intelligence level of the annotation results, solving the problems of low efficiency and poor accuracy in traditional manual annotation, and helping to accelerate the intelligent management and retrieval of educational video content. Description of the Drawings

[0062] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0063] Figure 1 It is a schematic structural diagram of an intelligent educational video content analysis and annotation system based on deep learning provided by an embodiment of the present invention. Detailed Embodiments

[0064] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0065] Embodiment 1:

[0066] The embodiment of the present invention provides an intelligent analysis and annotation system for educational video content based on deep learning, as Figure 1 shown, including:

[0067] A data processing module: acquiring educational video data, performing preprocessing, and extracting a plurality of scene features based on the preprocessed educational video data by using a pre-trained deep learning model;

[0068] A feature screening module: inputting the extracted scene features into a feature fusion network, performing feature screening and fusion based on preset conditions, and obtaining fused features;

[0069] A content reasoning module: based on the fused features, constructing a final graph structure of the video content through a graph neural network, and performing semantic reasoning in combination with preset teaching logic rules to generate a semantic description of the video content;

[0070] An annotation and optimization module: generating video annotations through a reinforcement learning model based on a knowledge graph and the semantic description, acquiring user feedback data, and then iteratively optimizing the annotation results to obtain final annotation results;

[0071] A storage and retrieval module: associatively storing the final annotation results with the preprocessed educational video data, and constructing a retrieval index based on the knowledge graph and the semantic description.

[0072] In this embodiment, the preprocessing includes video segmentation, key frame extraction, and noise filtering to obtain preprocessed video data;

[0073] In this embodiment, extracting several scene features includes: visual feature C1, audio feature C2, and text feature C3. Among them, C1 extracts the spatial features of video frames through a convolutional neural network, C2 extracts the spectral features of audio signals through a temporal convolutional network, and C3 extracts the text information in the video through a natural language processing model; Visual feature C1: Extracts the spatial features of key frames through a pre-trained convolutional neural network (CNN), and combines a temporal convolutional network (TCN) to capture the temporal relationship between frames; Audio feature C2: Converts the audio signal through a Mel spectrogram and uses a convolutional neural network to extract spectral features. At the same time, the audio is converted into text through a speech recognition model to extract speech semantic features; Text feature C3: Extracts the semantic embedding representation of subtitle text through a pre-trained natural language processing model (such as BERT), and combines a keyword extraction algorithm to obtain the core knowledge points in the text. The multimodal features C1, C2, and C3 respectively represent the visual, audio, and text information of the video.

[0074] In this embodiment, the preset conditions include teaching scene classification rules, knowledge point association rules, and semantic similarity thresholds.

[0075] The beneficial effects of the above technical solutions are: By combining deep learning, graph neural networks, and reinforcement learning, an efficient and intelligent educational video content analysis and annotation system is provided. Compared with the prior art, this system can automatically extract scene features in the video, generate accurate content descriptions based on semantic reasoning, and optimize the annotation through a knowledge graph. In addition, the system is continuously optimized through user feedback, improving the accuracy and intelligence level of the annotation results, solving the problems of low efficiency and poor accuracy of traditional manual annotation, and helping to accelerate the intelligent management and retrieval of educational video content.

[0076] Embodiment 2:

[0077] The embodiment of the present invention provides an intelligent analysis and annotation system for educational video content based on deep learning, and a data processing module, including:

[0078] Video segmentation unit: Segments the educational video data into several sub-video segments;

[0079] Video processing unit: Extracts several key frames of significant teaching content from each sub-video segment based on the dynamic saliency detection algorithm based on optical flow method, and then preprocesses each key frame to obtain a key frame set;

[0080] Feature extraction unit: Based on the preprocessed key frame set, uses a pre-trained deep learning model to extract several features respectively.

[0081] In this embodiment, preprocessing each key frame involves performing noise reduction and speech enhancement on the audio stream to extract clear speech signals; performing word segmentation, stop word removal, and semantic normalization on the subtitle text to obtain structured text data; the preprocessed video data is denoted as B, and B includes a key frame set, processed audio signals, and structured text data.

[0082] The beneficial effects of the above technical solution are as follows: By means of video segmentation and the optical flow method dynamic saliency detection algorithm, key frames in educational videos are accurately extracted, effectively improving the recognition accuracy of significant teaching content. Compared with the prior art, this method can automatically extract important content from sub-video segments and use a deep learning model to extract high-quality features, realizing the intelligent analysis of video content. This method not only improves the processing efficiency but also reduces manual intervention, enhancing the accuracy and intelligence level of educational video annotation and analysis.

[0083] Embodiment 3:

[0084] The embodiment of the present invention provides an intelligent analysis and annotation system for educational video content based on deep learning. The video segmentation unit includes:

[0085] The preliminary segmentation sub-unit: preliminarily segment the educational video data according to preset initial segmentation parameters to obtain an initial sub-video segment set;

[0086] The video analysis sub-unit: analyze each initial sub-video segment in the initial sub-video segment set to obtain the pixel difference and optical flow change data of each adjacent frame corresponding to all frames of each initial sub-video segment;

[0087] The video frame analysis sub-unit: obtain the inter-frame difference degree of each sub-video segment based on the pixel difference and optical flow change data of each adjacent frame corresponding to all frames of each initial sub-video segment;

[0088] The video marking sub-unit: if the inter-frame difference degree of each frame of each sub-video segment exceeds the preset inter-frame difference threshold, mark the frames with the inter-frame difference degree exceeding the preset inter-frame difference threshold as preliminary segmentation points;

[0089] The feature determination sub-unit: for each sub-video segment in the initial sub-video segment set, extract key frames and determine scene features;

[0090] The fine segmentation sub-unit: determine the similarity of scene features through a pre-trained convolutional neural network. If the scene feature similarity of adjacent sub-video segments in the initial sub-video segment set is lower than the preset scene change sensitivity, insert a scene change segmentation point at this position;

[0091] The segment generation sub-unit: combine the preliminary segmentation points and scene change segmentation points to generate several sub-video segments.

[0092] In this embodiment, the initialization segmentation parameters include a time interval threshold T, a scene change sensitivity, and a semantic consistency threshold.

[0093] In this embodiment, the video frame analysis sub-unit evaluates the inter-frame difference degree of each sub-video segment by calculating the pixel difference and optical flow change data of adjacent frames. The specific steps are as follows: Pixel difference calculation: For all frames of each initial sub-video segment, calculate the pixel value difference between adjacent frames frame by frame to generate a pixel difference matrix. Optical flow change calculation: Use an optical flow algorithm (such as Lucas-Kanade or Farneback) to extract the motion vectors between adjacent frames and calculate the optical flow change intensity. Inter-frame difference degree calculation: Combine the pixel difference and optical flow change data and calculate the inter-frame difference degree of each sub-video segment through weighted summation or a fusion model (such as a neural network). Difference degree analysis: Judge whether the inter-frame difference degree is significant according to a preset threshold to identify scene switching or content changes. This process can effectively capture the dynamic changes of video content and provide key data support for subsequent segmentation and annotation.

[0094] In this embodiment, a pre-trained convolutional neural network is used to calculate the scene feature similarity of adjacent sub-video segments to determine the scene change segmentation point. The specific steps are as follows: Scene feature extraction: For each segment in the set of initial sub-video segments, extract the key frames and generate scene feature vectors through a pre-trained CNN (such as ResNet or VGG). Similarity calculation: Calculate the cosine similarity or Euclidean distance between the scene feature vectors of adjacent sub-video segments to obtain the scene feature similarity. Segmentation point insertion: If the scene feature similarity of adjacent segments is lower than a preset scene change sensitivity (such as 0.7), then insert a scene change segmentation point at this position. Example: Suppose the video contains two scenes, "classroom lecture" and "experimental demonstration". After extracting features through the CNN, it is found that the scene feature similarity between the "classroom lecture" segment and the "experimental demonstration" segment is 0.5, which is lower than the preset sensitivity of 0.7. Then insert a segmentation point between the two to achieve precise segmentation of the scene.

[0095] The beneficial effects of the above technical solution are: Through a refined video segmentation method, combined with pixel differences, optical flow changes, and deep learning models, it automatically analyzes and identifies key frames and scene changes in the video, significantly improving the accuracy and intelligence level of video segmentation. Compared with the prior art, this method can not only effectively extract important video segments, but also perform fine segmentation according to the similarity of inter-frame differences and scene features, reducing manual intervention, improving the processing efficiency and annotation accuracy of educational videos, and optimizing the management and analysis process of video content.

[0096] Example 4:

[0097] The embodiment of the present invention provides an intelligent analysis and annotation system for educational video content based on deep learning. The feature screening module includes:

[0098] Model construction unit: Construct a scene classification model according to the preset teaching scene classification rules;

[0099] Probability acquisition unit: Input all scene features into the scene classification model respectively to obtain the probability distribution of each scene feature under each teaching scene;

[0100] Score determination unit: Based on the probability distribution, obtain the scene correlation score of each scene feature through weighted summation;

[0101] Feature screening unit: Screen out the scene features with scores higher than the preset score threshold based on the scene correlation scores of each scene feature and the preset score threshold;

[0102] Knowledge graph construction unit: Construct a knowledge graph according to the preset knowledge point association rules;

[0103] Feature mapping unit: Map all scene features into the knowledge graph to determine the association degree between each scene feature and the knowledge points;

[0104] Feature detection unit: Feature consistency detection based on the semantic similarity threshold:

[0105] Matrix determination unit: Determine the semantic similarity matrix between each scene feature based on the cosine similarity algorithm;

[0106] Feature marking unit: Mark the elements in the semantic similarity matrix that are lower than the preset semantic similarity threshold as inconsistent feature pairs based on the preset semantic similarity threshold;

[0107] Feature alignment unit: Perform semantic alignment on the inconsistent feature pairs through an adversarial generation network to generate consistent features;

[0108] Weight determination unit: Input the consistent features into the attention mechanism model to determine the attention weights of each consistent feature;

[0109] Feature determination unit: Perform weighted summation on all consistent features according to the attention weights of all consistent features to obtain the preliminary fusion features;

[0110] Feature optimization unit: Input the preliminary fusion features into the multi-modal fusion network for optimization to generate fusion features.

[0111] In this embodiment, the knowledge graph: The nodes represent knowledge points, and the edges represent the logical relationships between knowledge points.

[0112] In this embodiment, the preset knowledge point association rules are used to define the logical relationships between knowledge points, such as hierarchical relationships, dependency relationships, and parallel relationships. For example: Hierarchical relationship: Knowledge point A (Newton's first law) is the basis of knowledge point B (Newton's second law), and A must be explained before B. Dependency relationship: Knowledge point C (fundamental theorem of calculus) depends on knowledge point D (theory of limits), and D is the predecessor of C. Parallel relationship: Knowledge point E (theorem of kinetic energy) and knowledge point F (theorem of momentum) belong to the same knowledge module and can be explained in parallel. These rules are represented by a graph structure or logical expressions and are used to guide the construction of the knowledge graph and semantic reasoning to ensure the logic and coherence of the teaching content.

[0113] In this embodiment, the multi-modal fusion network consists of multiple bidirectional LSTM layers connected by residual connections and is used to capture the temporal dependency relationships of features.

[0114] The beneficial effects of the above technical solution are as follows: By combining deep learning, probability analysis, and semantic alignment technologies, an efficient feature screening and optimization method is provided. Compared with the prior art, this method can accurately screen features related to the teaching scenario and establish the association between features and knowledge points through the knowledge point graph. In addition, through semantic alignment by the adversarial generation network and weighting by the attention mechanism model, the fusion quality of features is further improved. Finally, by optimizing features through the multi-modal fusion network, the accuracy and intelligence level of educational video content analysis are improved, manual intervention is reduced, and the overall performance of the system is enhanced.

[0115] Embodiment 5:

[0116] The embodiment of the present invention provides an intelligent analysis and annotation system for educational video content based on deep learning. The content reasoning module includes:

[0117] Semantic analysis unit: Perform semantic analysis on the fused features to obtain the semantic information of the video content and define the set of node types;

[0118] Feature decomposition unit: Decompose the fused features to obtain the feature subsets corresponding to the node types in the set of node types;

[0119] Feature mapping unit: Map the feature subsets corresponding to the node types in the set of node types to a low-dimensional semantic space through a semantic embedding model to generate node embedding vectors;

[0120] Logical analysis unit: Perform logical analysis on the fused features to obtain the logical relationships of the video content, and then define the set of edge types;

[0121] Relationship extraction unit: Extract relationships from the fused features to obtain the relationship subsets corresponding to the edge types in the set of edge types;

[0122] Relationship mapping unit: Maps the relationship subset corresponding to the edge types in the edge type set to a low-dimensional relationship space through a relationship embedding model to generate edge embedding vectors;

[0123] Vector combination unit: Combines the node embedding vectors and the edge embedding vectors to construct an initial graph structure.

[0124] In this embodiment, the node type set includes knowledge point nodes, teaching object nodes, scenario nodes, and time nodes.

[0125] In this embodiment, the edge type set includes the hierarchical relationship between knowledge points, the association relationship between teaching objects and knowledge points, the inclusion relationship between scenarios and knowledge points, and the time sequence relationship between time nodes.

[0126] In this embodiment, the nodes in the initial graph structure represent knowledge points, teaching objects, scenarios, and time, and the edges represent the relationships between them.

[0127] The beneficial effects of the above technical solution are: By combining semantic analysis, feature decomposition, and relationship extraction, a high-efficiency content reasoning module is constructed using deep learning technology. Compared with the prior art, this method can map video content to a low-dimensional space through a semantic embedding model, achieve accurate node and edge feature representations, and extract deep relationships in video content through logical analysis and relationship mapping. This method improves the semantic reasoning ability of video content, enhances the understanding of complex teaching logic and video structure, thereby generating a more accurate and rich video content graph structure to support more efficient automatic annotation and intelligent analysis.

[0128] Embodiment 6:

[0129] The embodiment of the present invention provides an intelligent analysis and annotation system for educational video content based on deep learning. The content reasoning module further includes:

[0130] Knowledge reasoning unit: Performs knowledge reasoning on the initial graph structure based on preset knowledge reasoning rules to generate an inferred graph structure;

[0131] Structure update unit: Optimizes the graph structure of the inferred graph structure, defines update rules for the knowledge graph according to the dynamic changes of video content, and then dynamically updates the optimized graph structure to generate an updated graph structure;

[0132] Graph compression unit: Compresses the updated graph structure to generate a final graph structure;

[0133] Semantic reasoning unit: Performs semantic reasoning on the final graph structure based on preset teaching logic rules to generate a semantic description of video content.

[0134] In this embodiment, optimizing the graph structure after reasoning includes: a) removing redundant nodes and edges through knowledge distillation technology to generate a refined graph structure G2; b) adding missing nodes and edges through knowledge completion technology to generate a complete graph structure G3.

[0135] In this embodiment, the updated graph structure can reflect the real-time changes of the video content.

[0136] The beneficial effects of the above technical solutions are as follows: By combining knowledge reasoning and structure update technology, the reasoning process of video content is optimized. Compared with the prior art, this method can not only perform in-depth reasoning on the initial graph structure based on knowledge reasoning rules, but also update the graph structure in real time according to the dynamic changes of video content to ensure the timeliness and accuracy of the reasoning results. Precise semantic descriptions are generated through graph compression and semantic reasoning, improving the intelligent analysis and annotation ability of video content, enhancing the system's understanding and adaptation ability to complex teaching content, reducing manual intervention, and improving the automation level.

[0137] Embodiment 7:

[0138] The embodiment of the present invention provides an intelligent analysis and annotation system for educational video content based on deep learning, and a semantic reasoning unit, including:

[0139] Combined representation subunit: representing the preset teaching logic rules as a rule set;

[0140] Content matching subunit: traversing the nodes and edges in the final graph structure to match the content related to the rule set;

[0141] Semantic reasoning subunit: performing semantic reasoning on the content related to the rule set based on a rule-based inference engine;

[0142] Description generation subunit: generating a semantic description based on the semantic reasoning result and a preset description template.

[0143] In this embodiment, the preset teaching logic rules are represented as a rule set, where the form of each preset teaching logic rule is "premise → conclusion". The form of each rule r∈R is "premise → conclusion". For example: Rule 1: If knowledge point A is the predecessor of knowledge point B, then A must be explained before B. Rule 2: If knowledge point C belongs to a high level of difficulty, then C must appear after the basic knowledge is explained. Rule 3: If teaching object D is related to knowledge point E, then D needs to appear when explaining E.

[0144] In this embodiment, matching the content related to the rule set, for example: matching the hierarchical relationship between knowledge points and identifying the predecessor and successor nodes. Matching the association relationship between knowledge points and teaching objects and identifying the role of teaching objects in the explanation of knowledge points.

[0145] In this embodiment, the description template definition includes the following parts: Knowledge point description: Describes the content of the knowledge point and its position in the course. Teaching object description: Describes the role of the teaching object in the explanation of the knowledge point. Scenario description: Describes the ways of explaining the knowledge point in different scenarios. Logical relationship description: Describes the logical relationships between knowledge points (such as predecessors and successors, difficulty grading, etc.).

[0146] In this embodiment, semantic descriptions are generated based on the semantic reasoning results and the preset description template. For example: Knowledge point A is the basis of knowledge point B, and A needs to be explained first. A belongs to basic knowledge and is taught through classroom lectures. Knowledge point B belongs to advanced knowledge and needs to be explained after A. B is taught through experimental demonstrations. Teaching object C plays a key role in the explanation of B.

[0147] The beneficial effects of the above technical solution are: By combining teaching logic rules and semantic reasoning technology of deep learning, the intelligent level of educational video content analysis is improved. Compared with the prior art, this method can automatically match the nodes and edges in the graph structure, perform semantic reasoning based on preset rules, and generate accurate semantic descriptions. Through the combination of the inference engine and the description template, the system can efficiently generate semantic descriptions that meet teaching requirements, greatly improving the automation degree of educational video annotation, reducing manual intervention, and enhancing the accuracy and adaptability of the system.

[0148] Embodiment 8:

[0149] The embodiment of the present invention provides an intelligent analysis and annotation system for educational video content based on deep learning, and an annotation and optimization module, including:

[0150] Data acquisition unit: Acquires user feedback data;

[0151] Strategy adjustment unit: Inputs the final graph structure and semantic description into the reinforcement learning model, and the reinforcement learning model dynamically adjusts the annotation strategy according to the user feedback data;

[0152] Through several rounds of iterative optimization with the adjusted annotation strategy, the final annotation result is generated.

[0153] The beneficial effects of the above technical solution are: By combining reinforcement learning and user feedback to optimize the annotation strategy, the adaptive optimization of educational video content annotation is achieved. Compared with the prior art, this method dynamically adjusts the annotation strategy through user feedback data and performs multiple rounds of iterative optimization, thereby improving the accuracy and intelligent level of the annotation result. Through continuous adjustment of the reinforcement learning model, the system can continuously adapt to different user needs, optimize the annotation effect, reduce manual intervention, and significantly improve the automation and accuracy of educational video annotation.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent analysis and annotation system for educational video content based on deep learning, characterized in that Including: Data processing module: Obtain educational video data, perform preprocessing, and extract several scene features based on the preprocessed educational video data using a pre-trained deep learning model. Feature screening module: Input the extracted scene features into a feature fusion network, perform feature screening and fusion based on preset conditions, and obtain fused features. Content reasoning module: Based on the fused features, construct the final graph structure of the video content through a graph neural network, and perform semantic reasoning in combination with preset teaching logic rules to generate a semantic description of the video content. Annotation and optimization module: Based on the knowledge graph and semantic description, generate video annotations through a reinforcement learning model, obtain user feedback data, and then iteratively optimize the annotation results to obtain the final annotation results. Storage and retrieval module: Associatively store the final annotation results with the preprocessed educational video data, and construct a retrieval index based on the knowledge graph and semantic description.

2. The intelligent analysis and annotation system for educational video content based on deep learning according to claim 1, wherein, The data processing module includes: Video segmentation unit: Segment the educational video data into several sub-video segments. Video processing unit: Extract several key frames of significant teaching content from each sub-video segment based on a dynamic saliency detection algorithm based on optical flow method, and then preprocess each key frame to obtain a set of key frames. Feature extraction unit: Based on the set of preprocessed key frames, use a pre-trained deep learning model to extract several features respectively.

3. The intelligent analysis and annotation system for educational video content based on deep learning according to claim 2, wherein, The video segmentation unit includes: Initial segmentation subunit: Perform initial segmentation on the educational video data according to preset initial segmentation parameters to obtain a set of initial sub-video segments. Video analysis subunit: Analyze each initial sub-video segment in the set of initial sub-video segments to obtain the pixel difference and optical flow change data of each adjacent frame corresponding to all frames of each initial sub-video segment. Video frame analysis subunit: Obtain the inter-frame difference degree of each sub-video segment based on the pixel difference and optical flow change data of each adjacent frame corresponding to all frames of each initial sub-video segment. Video marking subunit: If the inter-frame difference degree of each frame of each sub-video segment exceeds a preset inter-frame difference threshold, mark the frame with the inter-frame difference degree exceeding the preset inter-frame difference threshold as an initial segmentation point. Feature determination subunit: For each sub-video segment in the set of initial sub-video segments, extract key frames and determine scene features. Fine segmentation subunit: Determine the similarity of scene features through a pre-trained convolutional neural network. If the scene feature similarity of adjacent sub-video segments in the set of initial sub-video segments is lower than a preset scene change sensitivity, insert a scene change segmentation point at this position. Segment generation subunit: Combine the initial segmentation points and scene change segmentation points to generate several sub-video segments.

4. The intelligent analysis and annotation system for educational video content based on deep learning according to claim 1, wherein The feature screening module includes: Model construction unit: Construct a scene classification model according to preset teaching scene classification rules. Probability acquisition unit: Input all scene features into the scene classification model respectively to obtain the probability distribution of each scene feature under each teaching scene. Score determination unit: Based on the probability distribution, obtain the scene correlation score of each scene feature through weighted summation. Feature Screening Unit: Screen out the scenario features with scores higher than the preset score threshold based on the scenario correlation scores of each scenario feature and the preset score threshold; Knowledge Graph Construction Unit: Construct a knowledge graph according to the preset knowledge point association rules; Feature Mapping Unit: Map all scenario features into the knowledge graph to determine the association degree between each scenario feature and the knowledge points; Feature Detection Unit: Feature consistency detection based on the semantic similarity threshold: Matrix Determination Unit: Determine the semantic similarity matrix between each scenario feature based on the cosine similarity algorithm; Feature Marking Unit: Mark the elements in the semantic similarity matrix that are lower than the preset semantic similarity threshold as inconsistent feature pairs based on the preset semantic similarity threshold; Feature Alignment Unit: Semantically align the inconsistent feature pairs through a generative adversarial network to generate consistent features; Weight Determination Unit: Input the consistent features into the attention mechanism model to determine the attention weights of each consistent feature; Feature Determination Unit: Perform weighted summation on all consistent features according to the attention weights of all consistent features to obtain the preliminary fusion features; Feature Optimization Unit: Input the preliminary fusion features into the multi-modal fusion network for optimization to generate fusion features.

5. The intelligent analysis and annotation system for educational video content based on deep learning according to claim 1, characterized in that Content Reasoning Module, including: Semantic Analysis Unit: Perform semantic analysis on the fusion features to obtain the semantic information of the video content and define the node type set; Feature Decomposition Unit: Decompose the fusion features to obtain the feature subsets corresponding to the node types in the node type set; Feature Mapping Unit: Map the feature subsets corresponding to the node types in the node type set into the low-dimensional semantic space through the semantic embedding model to generate node embedding vectors; Logical Analysis Unit: Perform logical analysis on the fusion features to obtain the logical relationship of the video content, and then define the edge type set; Relationship Extraction Unit: Extract the relationships corresponding to the edge types in the edge type set from the fusion features; Relationship Mapping Unit: Map the relationship subsets corresponding to the edge types in the edge type set into the low-dimensional relationship space through the relationship embedding model to generate edge embedding vectors; Vector Combination Unit: Combine the node embedding vectors and the edge embedding vectors to construct the initial graph structure.

6. The intelligent analysis and annotation system for educational video content based on deep learning according to claim 5, characterized in that, Content Reasoning Module, further including: Knowledge Reasoning Unit: Perform knowledge reasoning on the initial graph structure based on the preset knowledge reasoning rules to generate the reasoned graph structure; Structure Update Unit: Optimize the graph structure of the reasoned graph structure, define the update rules of the knowledge graph according to the dynamic changes of the video content, and then perform dynamic update on the optimized graph structure to generate the updated graph structure; Graph Compression Unit: Compress the updated graph structure to generate the final graph structure; Semantic Reasoning Unit: Perform semantic reasoning on the final graph structure based on the preset teaching logic rules to generate the semantic description of the video content.

7. The intelligent analysis and annotation system for educational video content based on deep learning according to claim 6, wherein Semantic Reasoning Unit, including: Combined Representation Sub-unit: Represent the preset teaching logic rules as a rule set; Content Matching Sub-unit: Traverse the nodes and edges in the final graph structure to match the content related to the rule set; Semantic Reasoning Sub-unit: The rule-based inference engine performs semantic reasoning on the content related to the rule set; Description Generation Sub-unit: Generates semantic descriptions based on the semantic reasoning results and a preset description template.

8. The intelligent analysis and annotation system for educational video content based on deep learning according to claim 1, characterized in that, Annotation and Optimization Module, including: Data Acquisition Unit: Acquires user feedback data; Strategy Adjustment Unit: Inputs the final graph structure and semantic descriptions into the reinforcement learning model, and the reinforcement learning model dynamically adjusts the annotation strategy according to the user feedback data; Performs several rounds of iterative optimization through the adjusted annotation strategy to generate the final annotation result.

Citation Information

Cited By

  • Audio and video scene intelligent switching optimization method and system combined with mode recognition

    CN120856941A

  • Intelligent video content extraction and rapid positioning system based on multi-modal fusion and space-time perception

    CN121353850A

  • Course video key frame intelligent identification method based on AI visual attention mechanism

    CN121884250A

  • Course video key frame intelligent identification method based on AI visual attention mechanism

    CN121884250B