Intelligent management method for AI big data
By annotating key points of teaching behaviors, training knowledge point association models, and building user profiles, the problem of insufficient association between teaching behaviors and knowledge points in AI video management has been solved. This has enabled accurate mapping and personalized storage of video content, improving the intelligence and user experience of teaching video management.
Patent Information
- Application Number
- CN202511807517.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-03-06
AI Technical Summary
Existing AI video management technologies struggle to effectively link teaching behaviors with knowledge points, resulting in inefficient video content analysis, an inability to support personalized learning needs, and a lack of scientific evaluation mechanisms for video segment value, leading to inefficient allocation of storage resources and preventing high-value content from reaching users first.
By collecting teaching video data, labeling key points of teaching behavior, training knowledge point association models, extracting video frame features and generating key point motion trajectories, constructing knowledge point graphs, calculating association factors and dynamic parameters, automatically isolating abnormal video segments by combining anomaly detection algorithms, collecting user behavior data to construct user profiles, calculating matching scores between video segments and learning paths, and storing video segments according to storage levels.
It achieves accurate mapping from low-level visual features to high-level semantics, efficient calculation of dynamic parameters, and personalized classification and storage, which improves the intelligence and user experience of teaching video management and ensures data quality and system reliability.
Smart Images

Figure CN121616888A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of AI video management technology, and in particular to an intelligent management method for AI big data. Background Technology
[0002] In the field of AI video management technology, existing methods struggle to effectively link teaching behaviors with knowledge points, resulting in low efficiency in video content analysis and an inability to support personalized learning needs. Teaching videos typically originate from heterogeneous data across multiple platforms and formats, and their processing environments face complex constraints such as high real-time requirements and large load fluctuations. Traditional technologies lack the ability to fine-grained modeling of dynamic features in teaching scenarios, failing to accurately capture teaching semantics. Teachers require efficient content distribution and influence enhancement tools, while students expect intelligent recommendations tailored to their individual learning paths, creating an inherent contradiction between resource allocation and effectiveness balance. Existing solutions often address one aspect while neglecting the other, failing to fully explore the semantic value of video content or deeply integrate user behavior data with teaching content, leading to a disconnect between video management and user needs. The lack of a scientific evaluation mechanism for video segment value results in inefficient allocation of storage resources, preventing high-value content from reaching users first, while the inefficiency in identifying and processing abnormal video segments further reduces the overall reliability of the system. These issues collectively restrict the large-scale application of AI video management technology in the education sector. To overcome the above limitations, a new generation of video management methods is needed that can integrate multi-source data, dynamically perceive the value of teaching content, and support intelligent decision-making and resource optimization. Summary of the Invention
[0003] This application provides an intelligent management method for AI big data, which solves the problems of inaccurate semantic mapping of teaching video content, insufficient dynamic feature quantification, low anomaly detection efficiency, and lack of personalized storage in the existing technology. It achieves the technical effects of accurate mapping from low-level visual features to high-level semantics, efficient calculation of dynamic parameters, automatic anomaly isolation, and personalized classification and storage.
[0004] This application provides an intelligent management method for AI big data, including: S1: Collect teaching video data, annotate key points guiding teaching behavior, including the starting point of teacher gestures, the area where teaching aids are used, and key writing actions; train a knowledge point association model based on the annotated data, extract video frame features to identify key points, and generate key point motion trajectories; S2: Extract core knowledge points of the subject to construct a knowledge point graph, associate key points with knowledge points to calculate the correlation factor; calculate dynamic parameters based on key point trajectory data, and incorporate the correlation factor to optimize parameter values; S3: Define abnormal video segments, mark video segments that deviate from the threshold using an anomaly detection algorithm, and automatically isolate and delete them; S4: Collect user behavior data to build user profiles, which include learning preferences and knowledge level dimensions; construct user learning paths based on knowledge point graphs and user profiles; calculate the basic matching score between video segments and learning paths, and obtain the final matching score through dynamic parameter weighting; divide video segments into three storage levels and store them in the corresponding storage layers based on the final matching score.
[0005] Furthermore, the teaching video data includes publicly available educational platforms, courses recorded within schools, and resources from partner institutions, all in high-definition MP4 format with a frame rate of no less than 30fps. The labeled data is stored in JSON and XML formats and includes frame indexes, key point coordinates, and behavior types. The knowledge point association model is a convolutional neural network, which includes convolutional layers, pooling layers, and fully connected layers. The activation function is a linear rectified function, the loss function is mean squared error, and the training parameters are set to a batch size of 32, an initial learning rate of 0.001, and a training cycle of at least 100 rounds. Negative samples are introduced during the training process, and the accuracy and recall are evaluated through a validation set.
[0006] Furthermore, the formula for calculating the correlation factor is as follows: , in, This is the correlation factor, with a value range of [0,1], where a higher value indicates a stronger correlation. The key point association score represents the degree of matching between a single key point and a specific knowledge point, with a value range of [0,1]. The weight of each knowledge point reflects its importance, and its value ranges from [0,1]. The total number of key points represents the number of key points identified in the video segment. The formula for calculating the optimized parameter value is as follows: , in, It is a dynamic parameter. Based on dynamic parameters, , The average displacement velocity is calculated over time intervals by removing keypoints between frames. The maximum possible speed.
[0007] Furthermore, the formula for calculating the final matching score is as follows: , in, The final matching score, with a value in the range [0,1]. Adjustment factors for user preferences are calculated based on the preference dimensions in the user profile. The value is set using continuous values from the user profile. , Preference weights are calculated based on user behavior. This is the scaling factor, with a value of 0.05.
[0008] Furthermore, the method also includes: Simultaneously extract audio and video data from the video stream, analyze speech rate, gesture changes, and interaction frequency, and normalize them to construct a style vector containing narration activity, interaction density, and emotional saturation. An event detection algorithm is used to identify the start and end points of interactive segments. After converting audio to text, noise is removed and words are segmented to generate standardized semantic tags. A tag library is constructed with hierarchical indexes based on discipline and interaction type. The video stream is segmented according to knowledge points and the order of explanation is recorded. A causal graph is constructed by combining style vectors and tag libraries, and the causal influence scores of teaching factors and learning outcomes are calculated. Calculate a priority score for each video segment: , in, Priority score, For causal influence on scores, The real-time access popularity score is calculated based on recent access data, including access frequency and number of interactions, and is normalized to a value in the range [0,1].
[0009] Furthermore, the normalization formula for speech rate is: , in, To standardize speaking speed, At the current speaking speed, At maximum speaking speed, The minimum speaking speed is set to the maximum, while the minimum and maximum speaking speeds are adjusted based on historical data.
[0010] Furthermore, the method also includes the following storage optimization steps: Establish a video segment value assessment system and calculate content importance scores: , in, As a score based on the importance of the content, Knowledge point density score, Score for teaching style For relative interaction scores, The corresponding weights are initially set to 0.4, 0.3, and 0.3, and their sum is 1. Calculate the dynamic factor score: , in, For dynamic factor scores, For access scores, For time-sensitive scores, For the scoring score, , and The corresponding weights are initially set to 0.5, 0.3, and 0.2, and their sum is 1. The content importance score and the dynamic factor score are combined to form the final value score: , in, The value is a fraction, ranging from [0,1]. As a weight for content importance, For dynamic factor weights, .
[0011] Furthermore, the value score includes: allocating resource layers according to the final value score: a value score ≥ 0.8 is a high resource layer, 0.5 ≤ value score < 0.8 is a medium resource layer, and a value score < 0.5 is a low resource layer; priority queues are used to manage coding tasks, and a maximum limit of 50% of the total resources occupied by the high resource layer is set.
[0012] Furthermore, the method also includes calculating a dynamic priority score for each video segment: , in, Priority score, Value score, The access frequency is normalized based on the number of accesses within 7 days; when When the value is ≥0.75, the video segment is migrated to the higher storage layer. When the value is ≤0.3, it is migrated to a lower storage layer; the migration is carried out gradually, large files are processed in blocks, and the capacity of the target layer is checked and cleanup is triggered before migration.
[0013] Furthermore, the storage level includes: High-priority storage level: Video segments with a final matching score ≥ 0.7 are stored in the high-speed storage layer, using solid-state drive media to ensure fast access and are suitable for video content that is highly relevant to user needs, such as core knowledge point explanations and highly interactive content. Medium priority storage level: Video segments with a final matching score of 0.4 ≤ final matching score < 0.7 are stored in the standard storage tier, using high-performance mechanical hard drives to balance access performance and storage cost, meeting the access needs of regular teaching videos; Low-priority storage tier: Video segments with a final matching score < 0.4 are stored in the archive tier using cloud storage media. With storage economy as the core, they are used to archive historical video content with low access frequency, reducing the consumption of core storage resources.
[0014] One or more technical solutions provided in this application have at least the following technical effects or advantages: By collecting multi-source teaching video data and annotating key points of teaching behavior, a neural network model was trained to accurately extract features, construct a knowledge point map, and realize the dynamic association and quantification of key points and knowledge points, thus completing the accurate mapping of video content from low-level visual features to high-level semantics. By integrating correlation factors to optimize dynamic parameters, a scientific evaluation of the knowledge transfer efficiency of video segments was achieved, and combined with a multi-dimensional anomaly detection mechanism, abnormal segments were automatically isolated, ensuring data quality and system reliability. Attached Figure Description
[0015] Figure 1 This is a flowchart of an AI big data intelligent management method according to an embodiment of the present invention. Detailed Implementation
[0016] To facilitate understanding of the present invention, a more complete description of this application will be given below with reference to the accompanying drawings, which illustrate preferred embodiments of the invention. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to enable a more thorough and complete understanding of the disclosure of the present invention.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to limit the invention; the term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0018] Example 1: As Figure 1 As shown, this is an intelligent management method for AI big data.
[0019] S1: Collect teaching video data, annotate key points guiding teaching behavior, including the starting point of teacher gestures, the area where teaching aids are used, and key writing actions; train a knowledge point association model based on the annotated data, extract video frame features to identify key points, and generate key point motion trajectories; The teaching video data includes publicly available educational platforms, courses recorded within schools, and resources from partner institutions. It is uniformly in high-definition MP4 format with a frame rate of no less than 30fps. The labeled data is stored in JSON and XML formats and includes frame index, key point coordinates, and behavior type. Specifically, teaching video data covering different disciplines and teaching styles is collected. Video sources include open education platforms, in-school recorded courses, and resources from partner institutions to ensure data diversity. The video format is uniformly high-definition MP4 with a frame rate of no less than 30fps to ensure clarity for key point recognition. The annotation focuses on key points oriented towards teaching behaviors, including: the starting point of teacher gestures, such as the position of the arm raised when writing on the blackboard, or the area pointed to by the finger when demonstrating teaching aids; the gesture type is recorded, and the start and end frames of the continuous trajectory are marked; the area where teaching aids are used, such as the operating range of experimental instruments and the clickable area of multimedia courseware; the type of teaching aid and its intended use are marked, and it is linked to specific knowledge points; and key writing actions, such as the trajectory of formulas written on the blackboard and the marked areas of key words and phrases; the dynamic writing path is marked, including the start and end points of the strokes. The annotated data is stored in JSON and XML formats, including frame indexes, key point coordinates, behavior types, and contextual tags.
[0020] The knowledge point association model is a convolutional neural network, which includes convolutional layers, pooling layers, and fully connected layers. The activation function is a linear rectified function, the loss function is mean squared error, and the training parameters are set to a batch size of 32, an initial learning rate of 0.001, and a training cycle of at least 100 rounds. Negative samples are introduced during the training process, and the accuracy and recall are evaluated through a validation set.
[0021] Specifically, a knowledge point association model (convolutional neural network model) is trained based on labeled data to identify key points of teaching behaviors. During training, the model learns to extract features from video frames and outputs the coordinate information of key points. The model includes convolutional layers, pooling layers, and fully connected layers. The activation function is ReLU, and the loss function uses mean squared error to optimize coordinate prediction. The training data uses a labeled dataset with a batch size of 32 and an initial learning rate of 0.001. The training cycle is at least 100 epochs, and accuracy and recall are evaluated on the validation set after each epoch to avoid overfitting. Key point recognition focuses on the characteristics of teaching behaviors; for example, the model learns to extract the contour features of gestures and the shape features of teaching aids, and focuses on high dynamic areas through an attention mechanism. Negative samples are introduced during training to enhance discriminative power. After the model outputs the key point coordinates for each frame, optical flow is used for inter-frame matching to ensure that the same key point corresponds one-to-one in consecutive frames. The identified key points are matched between different frames to generate key point motion trajectories.
[0022] S2: Extract core knowledge points of the subject to construct a knowledge point graph, associate key points with knowledge points to calculate the correlation factor; calculate dynamic parameters based on key point trajectory data, and incorporate the correlation factor to optimize parameter values; Specifically, for the target subject, curriculum standards, textbook content, and teaching syllabus are collected to extract core knowledge points. For example, in a mathematics course, knowledge points might include algebraic foundations, geometric proofs, etc., and each knowledge point is defined as a node. A knowledge graph structure is designed based on these nodes: nodes represent knowledge points, and edges represent relationship types. Each node is accompanied by metadata, including a difficulty coefficient (0-1 scale, e.g., 0.2 for basic knowledge points and 0.8 for advanced knowledge points), importance weight, and estimated explanation time.
[0023] This approach links key teaching behaviors with a knowledge point graph, achieving a mapping from low-level visual features to high-level semantics. Input data includes the movement trajectories of key points and temporal information from the video. Utilizing the semantic information of the key points, the explanatory text in the video is extracted using speech recognition technology, and the gesture positions are matched with the knowledge points mentioned in the text. The synchronicity between the timing of key point appearances and the progress of knowledge point explanation is calculated. For example, if a gesture appears continuously during the explanation of a core part of a knowledge point, the association strength is high; conversely, if the gesture is unrelated to the knowledge point, the association strength is low. An association score is calculated for each key point, based on semantic matching degree and temporal consistency. The output is a mapping table recording the association strength (range 0-1) between each key point and the knowledge point.
[0024] The mapping results are transformed into quantifiable correlation factors. The weights of knowledge points are obtained from the knowledge point graph, and combined with the dynamic characteristics of keypoints, the movement trajectory of keypoints and the weights of knowledge points are quantified. For example, a keypoint has a large contribution when it is associated with a high-weight knowledge point and its movement is significant. For each video segment, the correlation scores of all keypoints are aggregated, and the correlation factor is calculated. , in, The correlation factor is a comprehensive indicator of the degree of correlation between key points and knowledge points in the entire video segment. Its value ranges from [0,1], with a higher value indicating a stronger correlation. The key point association score is the degree of matching between a single key point and a specific knowledge point, with a value range of [0,1]. The weight of a knowledge point reflects its importance. It is derived from the metadata of the knowledge point graph and has a value range of [0,1]. The total number of keypoints represents the number of keypoints identified in the video segment. This value is used for normalization to avoid bias caused by differences in the number of keypoints. Factor values are updated in real time based on the video timestamp. For example, during knowledge point transition periods, factor values may fluctuate. The algorithm uses a sliding window averaging process to ensure stability.
[0025] Based on the key point trajectory data output by the knowledge point association model, the basic dynamic parameters of the video segment are obtained, and the relevance factor is incorporated into the basic dynamic parameters to amplify the parameter values of periods with high knowledge relevance: , in, It is a dynamic parameter. As a basic dynamic parameter, it is the normalized dynamic intensity of the key point's motion trajectory. , The average displacement velocity is calculated over time intervals by removing keypoints between frames. To achieve the maximum possible speed, the settings are based on training set statistics. The maximum average displacement velocity should be calculated using all video segments in the training set and updated periodically to accommodate new data, ensuring... Always normalize to [0,1].
[0026] A segment-by-segment processing approach is adopted, with dynamic parameters calculated independently for each segment. The calculation process is optimized to incremental updates, triggering recalculation only when key points and knowledge points change, thus reducing computational overhead. The optimized parameters are smoothed using a sliding window averaging method to reduce instantaneous fluctuations and ensure a smooth parameter curve.
[0027] S3: Define abnormal video segments, mark video segments that deviate from the threshold using an anomaly detection algorithm, and automatically isolate and delete them; Specifically, based on the characteristics of instructional videos, video segments that may affect teaching effectiveness and the reliability of data analysis are defined as anomalies, categorized into technical anomalies, content anomalies, and temporal anomalies. Technical anomalies stem from hardware or transmission problems, such as lost video frames, audio synchronization failures, or sudden drops in resolution. These anomalies are determined using quantitative indicators (e.g., frame rate fluctuations greater than 20%). Content anomalies are segments unrelated to the teaching topic, such as suddenly inserted advertisements, environmental interference, or failures to identify key points. Determination criteria include a correlation with the knowledge point graph below a threshold (e.g., correlation factor < 0.2). Temporal anomalies are caused by disordered teaching logic, such as reversed order of knowledge point explanations or abnormal video segment lengths. Standard anomaly detection algorithms are applied to perform real-time analysis of the video stream. The algorithm compares dynamic parameters with thresholds, marking video segments with significant deviations as anomalies. Video segments marked as anomalies are automatically isolated or deleted to prevent them from affecting subsequent analysis.
[0028] S4: Collect user behavior data to build user profiles, which include learning preferences and knowledge level dimensions; construct user learning paths based on knowledge point graphs and user profiles; calculate the basic matching score between video segments and learning paths, and obtain the final matching score through dynamic parameter weighting; divide video segments into three storage levels and store them in the corresponding storage layers based on the final matching score.
[0029] Specifically, user behavior data is collected, including user viewing history, interaction records, performance data, and explicit feedback, to build dynamic user profiles. These profiles contain multi-dimensional features, such as learning preferences, knowledge levels, and behavioral patterns.
[0030] A learning path-video content matching algorithm aligns video content with the user's learning path. The path is organized in a graph format, with nodes representing knowledge points and edges indicating the learning order. Each knowledge point node includes metadata, including difficulty level, estimated learning time, and importance weight. The path structure is dynamically adjusted based on user profiles. For example, for users with low knowledge levels, the path prioritizes basic knowledge point nodes and shortens the intervals between advanced stages; while for users who prefer practical application, the path increases the weight of demonstration-type knowledge points. Path generation uses a graph algorithm and is updated in real-time based on user behavior. For example, when a user completes a knowledge point quiz, the path automatically adjusts the next target node and recalculates the priority. The update cycle is set to trigger daily to ensure the timeliness of the path.
[0031] A matching score is output by comparing the similarity between video segment content and the user's learning path. The matching score is calculated using cosine similarity, comparing the knowledge point vectors of the video segment and the user's learning path. Before calculating the cosine similarity, the vectors are projected onto a unified knowledge point space, and the union of all knowledge points from the user path and the video segment is extracted, filling in missing dimensions with 0 to ensure consistent vector dimensions. , in, The basic matching score ranges from [0,1]. A value closer to 1 indicates a greater similarity between the video content and the user's path. This is a vector of knowledge points from a video segment. The knowledge point vector of the user's learning path. The dot product of vectors is calculated using the following formula: ,in It represents the correlation strength of the video segment with the i-th knowledge point. It is the priority weight of the user path on the i-th knowledge point. This represents the total number of knowledge points. and Let represent the magnitudes of the vectors, respectively, and the calculation formula is: and .
[0032] The base score is weighted by dynamic parameters to reflect the efficiency of knowledge transfer. , in, For weighted scores, The adjustment coefficient, with a value of 0.1, was obtained through training based on historical data and is used to balance the influence of dynamic parameters.
[0033] By integrating user preference data and fine-tuning the weighted scores, the final matching score is: , in, The final matching score, with a value in the range [0,1]. Adjustment factors for user preferences are calculated based on the preference dimensions in the user profile. The value is set using continuous values from the user profile. , Preference weights are calculated based on user behavior. This is the scaling factor, with a value of 0.05.
[0034] The overall formula for calculating the matching degree is: , After calculation, the scores are normalized to ensure comparability between different video segments.
[0035] Based on the final matching score, video segments are divided into three storage tiers: Video segments with a matching score ≥ 0.7 are assigned to the high-priority storage tier, indicating high relevance to user needs, such as explanations of core knowledge points or highly interactive content; they are allocated to the high-speed storage tier to ensure fast access. Video segments with a matching score between 0.4 and 0.7 are assigned to the medium-priority storage tier and allocated to the standard storage tier. Video segments with a matching score < 0.4 are assigned to the low-priority storage tier and stored in the archive tier to reduce resource consumption.
[0036] The technical solutions described in the embodiments of this application have at least the following technical effects or advantages: This application collects multidisciplinary teaching video data and annotates key points such as gesture start points and teaching aid usage areas. It then trains a convolutional neural network model to accurately extract teaching behavior features, constructs a knowledge point graph, and dynamically associates key points with knowledge points. By calculating the correlation factor, it achieves a precise mapping of video content from low-level visual features to high-level semantics. By integrating the correlation factor into basic dynamic parameters, it optimizes the generation of dynamic parameters, quantifies the dynamism of video segments and the efficiency of knowledge transfer, and automatically isolates technical, content-related, and time-series abnormal video segments using an anomaly detection mechanism, ensuring data quality. Based on user profiles and learning path matching algorithms, it calculates the final matching score and stores video segments according to priority in high-speed, standard, or archive layers, achieving personalized management and efficient access to video content. This comprehensively improves the accuracy, processing efficiency, and user experience of intelligent analysis of teaching videos.
[0037] Example 2: Example 1 achieved semantic mapping and personalized storage of video content, but it still has shortcomings in dynamic analysis of teaching behavior sequences and causal relationship evaluation. This example further supplements the content of Example 1.
[0038] The method also includes: synchronously extracting audio and video data from the video stream, analyzing speech rate, gesture changes, and interaction frequency and normalizing them, and constructing a style vector containing explanation activity, interaction density, and emotional saturation. An event detection algorithm is used to identify the start and end points of interactive segments. After converting audio to text, noise is removed and words are segmented to generate standardized semantic tags. A tag library is constructed with hierarchical indexes based on discipline and interaction type. The video stream is segmented according to knowledge points and the order of explanation is recorded. A causal graph is constructed by combining style vectors and tag libraries, and the causal influence scores of teaching factors and learning outcomes are calculated. Specifically, audio and video data are extracted synchronously from the teaching video stream for behavioral sequence analysis. Audio data is used to analyze speech rate, calculating the average speech rate and its trend through speech recognition technology; video data is used to track gesture changes, extracting gesture trajectories based on a knowledge point association model; and interaction frequency is counted by detecting the number of interactions per unit time through teacher-student verbal alternation or visual interaction.
[0039] Normalize speech rate characteristics: , in, To standardize speaking speed, At the current speaking speed, At maximum speaking speed, The minimum speaking speed is set to the maximum, while the minimum and maximum speaking speeds are adjusted based on historical data.
[0040] Normalize gesture features: Taking gesture amplitude as an example, set a baseline amplitude according to the teacher's habits, and scale the excess portion proportionally.
[0041] A style vector is constructed, comprising explanation activity, interaction density, and emotional saturation. Explanation activity is a weighted sum of speech rate and gesture amplitude (weight ratio set to 6:4), with high scores indicating a passionate explanation style; interaction density is the product of interaction frequency and eye contact changes, reflecting the teacher's attention to students; and emotional saturation is a fusion value of audio emotional intensity and facial expression analysis, used to distinguish between rational explanations and emotional teaching.
[0042] The video stream is scanned using an event detection algorithm, and the start and end points of interactive segments are identified based on audio and visual features. For example, when a student asks a question or a teacher turns to a student, the start of the segment is marked. Semantic analysis is performed on the extracted interactive segments, converting the audio into text. The converted text is then denoised, removing filler words such as "um" and "ah," and segmented into semantic units using a word segmentation tool. For example, the sentence "Please think about the solution to this equation" is segmented into ["think", "equation", "solution"].
[0043] The extracted interaction fragments are transformed into structured descriptions. Natural language processing and contextual analysis are used to standardize and semanticize the tags, generating semantic tags. A combination of rule bases and machine learning is used to identify question and feedback types. Question types are distinguished as open-ended or closed-ended through syntactic analysis, and feedback types are determined as positive or corrective feedback based on a sentiment lexicon. A knowledge point graph is then used to map interactive text to knowledge points. For example, if the text contains the term "quadratic function," it is automatically associated with the knowledge point "quadratic function in mathematical algebra," and the association strength is calculated. Sentiment analysis outputs sentiment polarity, including positive, neutral, and negative, while also estimating the interaction difficulty based on text complexity. Semantic tags for all interaction fragments are stored in a central tag library. The tag library is hierarchically indexed by subject and interaction type, and each tag in the library includes a timestamp and confidence score. A tag library is constructed to store and manage all semantic tags. A NoSQL database is used, with hierarchical storage by subject, interaction type, and time. Each document contains a fragment ID, tag set, time metadata, and access statistics.
[0044] Based on a knowledge point graph, the video stream is segmented by knowledge point. Each knowledge point segment corresponds to a time interval, and its explanation order is recorded. The sequence model uses a directed graph representation, where nodes represent knowledge points and edges represent the logical flow of explanation. Natural language processing techniques are used to analyze the video-transcribed text and automatically identify knowledge point boundaries. For example, keyword matching and context analysis are used to determine the start and end points of knowledge points. Simultaneously, metadata is added to each knowledge point, including difficulty level, estimated explanation duration, and associated interactive tags, forming structured knowledge point sequence data.
[0045] Based on the order of knowledge point explanations, a causal graph is constructed using style vectors and interactive tag libraries. Learning effect data for each knowledge point segment is aligned with its corresponding style vector and interactive tag by timestamp. For each teaching factor, a causal strength coefficient between it and the learning effect is calculated, ranging from [0,1], with higher values indicating stronger causal relationships. A graph database is used to store causal relationships, where knowledge point nodes and teaching factor nodes are connected by directed edges, with edge weights dynamically updated based on the causal coefficient. A periodic retraining mechanism is implemented, such as monthly, to update the causal graph with new data. For example, when new video data indicates a change in the effect of a certain teaching factor, the edge weights are automatically adjusted to ensure the graph's timeliness. Knowledge point segments with high causal strength are extracted from the causal graph and marked as high-value segments.
[0046] Based on the causal analysis results, a video storage order is set, and the value of video segments is visualized as a calculable indicator. A priority score is calculated for each video segment: , in, Priority score, The causal influence score is obtained from the causal graph and represents the influence value of the segment on the learning effect, with a value range of [0,1]. The real-time access popularity score is calculated based on recent access data, including access frequency and number of interactions, and is normalized to a value in the range [0,1].
[0047] Video clips are stored in descending order of priority score. High-scoring clips ( ≥0.7) is stored in the high-speed cache layer to ensure fast access. This layer mainly stores key content such as core knowledge point explanations and frequently interacting segments, ensuring millisecond-level access speed; medium-sized segments (0.4≤) are stored in the cache layer to ensure fast access. <0.7) is stored in the standard tier, using a high-performance hard disk drive, balancing performance and cost; low-segment ( <0.4) is stored in the archive tier, focusing on storage cost-effectiveness. The score is recalculated weekly to adapt to content changes.
[0048] To verify the practical effectiveness of this technical solution, a comprehensive evaluation system was established over a three-month period. This system, centered on quantitative indicators, aims to systematically verify the solution's improvements in three dimensions: storage efficiency, learning outcomes, and system stability. Regarding storage access efficiency, the goal is to reduce the average access time for high-priority video clips by 20%, from 200 milliseconds before optimization to less than 160 milliseconds. This data will be obtained by comparing access latency measurements before and after optimization. Regarding improvements in learning outcomes, A / B testing will be used for verification. The goal is for students in the experimental group using the optimized stored videos to achieve a 15% improvement in their average test scores compared to the control group, for example, from 80 points to 92 points, thus directly measuring the positive impact of the solution on teaching effectiveness. To ensure system stability, the error rate of the storage system is required to be controlled below 0.1%. The entire evaluation period is three months, during which data will be summarized and analyzed monthly to ensure the attainability of the evaluation standards. Necessary strategy adjustments will be made based on the interim results, thereby comprehensively and objectively demonstrating the technical advantages and practical value of the solution.
[0049] The technical solutions described in the embodiments of this application have at least the following technical effects or advantages: This application achieves precise quantification of teaching styles by synchronously extracting audio and video data from teaching video streams, analyzing speech rate, gesture changes, and interaction frequency, and constructing style vectors that include explanation activity, interaction density, and emotional saturation. It automatically identifies interactive segments using event detection algorithms and generates a standardized semantic tag library using natural language processing technology, enabling intelligent extraction and semantic management of unstructured interactive content. By constructing a causal graph based on the order of knowledge point explanations, combined with style vectors and the interaction tag library, it analyzes the causal relationship between teaching factors and learning outcomes, achieving a scientific evaluation of video segment value. Priority scores are calculated based on causal influence scores and real-time access popularity scores to optimize video storage order, storing video segments in descending order of score to the cache layer, standard layer, or archive layer, improving storage efficiency and learning outcomes. Furthermore, regular re-evaluation and dynamic migration mechanisms ensure system stability, thereby enhancing the overall intelligence and practicality of teaching video management.
[0050] Example 3: Example 2 implemented teaching behavior sequence analysis and causal-driven storage optimization, but still suffers from drawbacks such as insufficient dynamic allocation of storage resources and inadequate locality optimization. This example further supplements and explains the content of Example 2.
[0051] A video segment value assessment system was established, and natural language processing was performed on the transcribed text of the videos based on a knowledge point association model. A keyword extraction algorithm was used to identify the knowledge points involved in each video segment, and the number of knowledge points per unit time was counted to form a knowledge point density score. Based on the knowledge graph, a difficulty coefficient (range [0,1]) was preset for each knowledge point. Simple knowledge points were assigned a low score (0.2), and complex knowledge points were assigned a high score (0.8). The difficulty score of a video segment was the average difficulty value of the knowledge points involved. Logical coherence was checked, analyzing whether the order of knowledge point explanations conformed to teaching logic. Deviations were detected using a sequence matching algorithm; if the order was disordered, the score was appropriately reduced. Based on the behavioral sequence analysis results, such as speech rate stability, gesture amplitude, and interaction frequency, these features were normalized to a teaching style score of [0,1]. User behavior data, such as video playback times and note marking density, was collected from the learning platform and normalized to obtain a relative interaction score. Content importance scores were calculated. , in, As a score based on the importance of the content, Knowledge point density score, Score for teaching style For relative interaction scores, The corresponding weights are initially 0.4, 0.3, and 0.3, and their sum is 1.
[0052] Real-time monitoring of daily video segment access volume, using time series analysis to identify access trends, and calculating access scores: , in, For access scores, For page views, This refers to the visit trend, specifically the growth rate of visit volume. The average completion rate of users watching video segments. , , The corresponding weights are initially set to 0.4, 0.3, and 0.3, and their sum is 1.
[0053] A decay curve is set based on video upload time. New videos have a timeliness score of 1 in the first month, decreasing by 0.1 each month thereafter. If no update is made for more than six months, the score drops below 0.3. Update thresholds are set based on subject characteristics (e.g., IT technology videos are updated every six months, basic theory videos every two years). Expired content has its score automatically reduced. An event gain coefficient is set to dynamically adjust the score according to the semester schedule. For example, exam review videos receive a 20% increase in timeliness score one month before the exam. Calculate the timeliness score: , in, For time-sensitive scores, Based on timeliness score, This is the event gain coefficient, with a value range of [1.0, 1.5].
[0054] The five-star rating is converted into a numerical value of [0,1], such as 1 star = 0.2, 5 stars = 1.0; and natural language processing technology is used to analyze the sentiment polarity (positive, negative, neutral) of user reviews and convert it into a numerical score (positive = 1, negative = 0, neutral = 0.5). The rating score is calculated as follows: , in, For the scoring score, The mean of the numerical scores. The credibility coefficient , The actual number of people who rated it. For the prior sample size, For emotional score, The text weight is determined using experimental data, with an initial value of 0.3 and a range of [0.2, 0.5].
[0055] Dynamic factor scores are calculated based on access score, timeliness score, and rating score: , in, For dynamic factor scores, For access scores, For time-sensitive scores, For the scoring score, , , The corresponding weights are initially set to 0.5, 0.3, and 0.2, and their sum is 1.
[0056] The content importance score and the dynamic factor score are combined to form the final value score: , in, The value is a fraction, ranging from [0,1]. As a weight for content importance, For dynamic factor weights, It emphasizes the essential value of content.
[0057] The value score includes: allocating resource layers based on the final value score: a value score ≥ 0.8 is a high resource layer, 0.5 ≤ value score < 0.8 is a medium resource layer, and a value score < 0.5 is a low resource layer; priority queues are used to manage coding tasks, with a maximum limit of 50% for high resource layers to occupy the total resources.
[0058] Specifically, based on the value score, resources are divided into three tiers: The high-resource tier is for video segments with a value score ≥ 0.8, allocating the highest priority resources, including high-performance computing nodes and advanced coding algorithms, significantly compressing file size while maintaining clarity. The medium-resource tier is for video segments with a value score between 0.5 and 0.8, allocating standard resources to balance quality and efficiency, with a compression ratio controlled at approximately 30:1. The low-resource tier is for video segments with a value score < 0.5, allocating basic resources and employing lightweight coding for fast processing, with a compression ratio of approximately 20:1, saving computing resources.
[0059] A priority queue management system is implemented, with video segments entering the encoding queue in descending order of their value score. High-value segments are processed first to reduce waiting latency; low-value segments are processed in batches when resources are available. Resource limits are set for each level, such as ensuring that high-resource levels do not consume more than 50% of the total resources, preventing excessive consumption by a single video segment and ensuring system fairness.
[0060] The storage levels include: Specifically, based on the initial position of video segment value allocation, the storage system is divided into three layers: The high-speed cache layer uses solid-state drives (SSDs) to store high-value video segments, characterized by low access latency but higher cost and limited capacity, suitable for core teaching content requiring frequent access. The standard storage layer uses hard disk drives (HDDs) to store medium-value video segments, balancing performance and cost, with access latency of 50-100 milliseconds, meeting regular access needs. The archive layer uses cloud storage to store low-value video segments, with higher access latency but lower cost, suitable for archiving historical videos that are rarely accessed. A capacity ratio is set for each layer, such as 20% for the high-speed layer, 50% for the standard layer, and 30% for the archive layer. This ratio can be dynamically adjusted according to the overall data volume to avoid overloading any single layer. After a video segment is uploaded, it is directly assigned a layer based on its calculated value score, and a storage layer tag is added to each video segment.
[0061] Calculate the dynamic priority score for each video segment in real time: , in, Priority score, Value score, The access frequency is normalized based on the number of accesses within 7 days.
[0062] When the priority score is ≥0.75, the video segment is migrated from a lower-level layer to a higher-level layer; when the priority score is ≤0.3, the video segment is migrated from a higher-level layer to a lower-level layer. The initial threshold value is set based on historical data and can be dynamically fine-tuned according to system load. Changes in value score and access frequency are monitored; when the change exceeds 10%, a migration check is immediately triggered. The migration process uses a gradual operation: a copy is created in the new layer, and after verifying its integrity, the old data is deleted to avoid access interruption. For large files, chunked migration is supported to reduce the impact on system performance. The target layer capacity is checked during migration; if insufficient, layer cleanup is automatically triggered to ensure a smooth migration. Simultaneously, migration queue priorities are set, with high-value segments processed first.
[0063] Calculate the co-occurrence access probability among video segments, for example, by identifying frequently accessed video segment groups through association rule mining. Spatial locality strength is represented by co-occurrence probability; higher probability indicates stronger spatial association. Analyze access time series to identify recurring access patterns. Use sliding window statistics (e.g., 7 days) to calculate the recent access frequency of video segments and calculate the variance of access intervals; smaller variance indicates stronger temporal locality (e.g., access at fixed times each day). Use clustering algorithms to group video segments according to access patterns, marking high-frequency access groups as hotspot areas. Simultaneously detect abnormal access, distinguishing between normal and temporary locality. Store the quantification results as feature vectors and create an index for fast retrieval. The feature library is updated regularly to ensure data freshness.
[0064] The data block size is adaptively adjusted based on the intensity of spatial locality. For video groups with high spatial locality, the data block size is increased to reduce disk seek times; for videos with low spatial locality, a smaller block size is maintained for flexible management. The adjustment algorithm uses heuristic rules, such as the block size being proportional to the co-occurrence probability. After adjustment, access latency is monitored; if latency increases, the size is rolled back to the original size. The differences in video segment sizes are also considered, with larger files using larger blocks and smaller files using smaller blocks.
[0065] Based on spatial locality analysis, highly correlated video segments are physically stored in the same disk sector or adjacent nodes, and hot video segments are preferentially placed in adjacent positions in the high-speed storage layer. Based on temporal locality analysis, the next video segment that a user is likely to access is predicted and preloaded into the cache. The preloading trigger conditions include historical access sequences and the current access context.
[0066] The technical solutions described in the embodiments of this application have at least the following technical effects or advantages: This application achieves quantitative evaluation of video segment value by calculating content importance scores and dynamic factor scores and integrating them into a final value score. Based on the value score, video segments are allocated to high, medium, and low resource layers, and encoding resources are dynamically adjusted, achieving efficient resource allocation. By calculating priority scores in real time and triggering storage layer migration based on thresholds, combined with a monitoring mechanism to ensure timely dynamic adjustments, adaptive optimization of the storage structure is achieved. By analyzing spatiotemporal locality characteristics, data block size and storage location are adaptively set, and highly correlated video segments are stored nearby or pre-loaded into the cache, improving data locality and reducing access latency. Through value-driven and locality optimization, the inefficiency of static storage strategies is solved, improving storage resource utilization, system response speed, and management intelligence.
[0067] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. An intelligent management method for AI big data, characterized in that, Comprise: S1: Collect teaching video data, label teaching behavior guide key points, including teacher gesture starting point, teaching aid use area and key writing action; Train knowledge point association model based on labeled data, extract video frame features to identify key points, and generate key point motion trajectories; S2: Extract subject core knowledge points to construct knowledge point graph, associate key points with knowledge points to calculate association factor; Calculate dynamic parameters based on key point trajectory data and integrate association factor to optimize parameter value; S3: Define abnormal video segments, mark video segments deviating from threshold value and automatically isolate and delete them through abnormal detection algorithm; S4: Collect user behavior data to construct user portrait, which includes learning preference and knowledge level dimensions; construct user learning path based on knowledge point graph and user portrait; Calculate basic matching score of video segment and learning path, and get final matching score through dynamic parameter weighting; According to the final matching score, the video segment is divided into three storage levels and stored in the corresponding storage layer. 2.The AI big data intelligent management method of claim 1, wherein, The teaching video data includes public education platform, on-campus recorded courses and cooperative agency resources, which are unified into high-definition MP4 format with frame rate not less than 30fps, and the labeled data is stored in JSON and XML format and contains frame index, key point coordinates and behavior type; The knowledge point association model is a convolutional neural network, which includes convolutional layer, pooling layer and fully connected layer, the activation function uses linear rectifier function, the loss function is mean square error, the training parameters are set as batch size 32, initial learning rate 0.001 and training period at least 100 rounds, and negative samples are introduced in the training process and the accuracy and recall rate are evaluated by validation set. 3.The AI big data intelligent management method of claim 1, wherein, The association factor calculation formula is: , wherein, is a correlation factor, with a value range of [0, 1], and the higher the value, the stronger the correlation; is a key point correlation score, which is the matching degree of a single key point with a specific knowledge point, with a value range of [0, 1]; is a knowledge point weight, which reflects the importance of the knowledge point, with a value range of [0, 1]; is the total number of key points, which is the number of key points identified in the video segment; The optimized parameter value calculation formula is: , wherein, is a dynamic quantity, is a base dynamic quantity, , is an average displacement velocity, calculated by dividing the inter-frame displacement of the keypoint by the time interval; is a maximum possible velocity. 4.The AI big data intelligent management method of claim 1, wherein, The final matching score calculation formula is: , where, is the final match score, taking values in the range [0, 1], is the user preference adjustment factor, computed based on the preference dimensions in the user profile; is set by the continuous values of the user profile, , is the preference weight, computed by the user behavior, is the scaling factor, taking the value 0.
05. 5.The AI big data intelligent management method of claim 1, wherein, The method further comprises: Synchronously extract audio and video data of video stream, analyze speech rate, gesture change and interaction frequency and normalize, construct style vector containing explanation activity, interaction density and emotional saturation; Use event detection algorithm to identify interaction segment start and end points, denoise, segment and generate standardized semantic labels after converting audio to text, and construct label library indexed by subject and interaction type; Segment video stream by knowledge points and record explanation order, construct causal diagram combining style vector and label library, and calculate causal influence score of teaching factors and learning effect; Calculate priority score for each video segment: , wherein, is a priority score, is a causal impact score, is a real-time access heat score, calculated from recent access data, including access frequency, interaction times indicators, normalized to [0, 1] values. 6.The AI big data intelligent management method of claim 5, wherein, The speech rate normalization formula is: , wherein, is the standard speech rate, is the current speech rate, is the maximum speech rate, is the minimum speech rate, the maximum and minimum speech rates are adjusted by historical data. 7.The AI big data intelligent management method of claim 1, wherein, The method further comprises the following storage optimization steps: Establish video segment value evaluation system to calculate content importance score: , wherein, is a content importance score, is a knowledge point density score, is a teaching style score, is a relative interaction score, is a corresponding weight, the initial weights are 0.4, 0.3 and 0.3, and the sum is 1. Calculate dynamic factor score: , wherein, is a dynamic factor score, is an access score, is an age score, is a rating score, , and are corresponding weights, with initial values of 0.5, 0.3 and 0.2, and satisfying and = 1. Fuse content importance score and dynamic factor score into final value score: , wherein, is a value score, with a value range [0, 1], is a content importance weight, is a dynamic factor weight, . 8.The AI big data intelligent management method of claim 7, wherein, The value score includes: according to the final value score, the resource layer is divided into high resource layer, medium resource layer and low resource layer according to the value score: value score≥0.8, 0.5≤value score<0.8, value score<0.5; use priority queue to manage coding tasks, and set the upper limit of total resources occupied by high resource layer to not more than 50%. 9.The AI big data intelligent management method of claim 7, wherein, The method further comprises calculating dynamic priority score of each video segment: , wherein, is a priority score, is a value score, is a visit frequency, normalized according to the number of visits within 7 days; When ≥ 0.75 video segment migrates to high tier storage, ≤ 0.3 migrates to low tier storage; migration uses gradual operation, large file chunking, checks target tier capacity before migration and triggers cleanup. 10.The AI big data intelligent management method of claim 1, wherein, The storage level includes: High priority storage level: video segments with final matching score ≥ 0.7 are stored in high-speed storage layer, using solid state disk medium, to ensure fast access, and are suitable for core knowledge point explanation, high interactivity and other video content highly related to user needs. Medium priority storage level: video segments with 0.4 ≤ final matching score < 0.7 are stored in standard storage layer, using high-performance mechanical hard disk, to balance access performance and storage cost, and to meet the access needs of regular teaching videos. Low priority storage level: video segments with final matching score < 0.4 are stored in archive layer, using cloud storage medium, with storage economy as the core, for archiving historical video content with low access frequency, to reduce the occupation of core storage resources.
Citation Information
Cited By
Pathological full-slice image storage method and data supply method oriented to AI processing
CN122117208A