Classical music cross-modal high-reliability experimental data set construction method for full-scene teaching application
Through the standardized processing of multimodal data, cross-modal attention mechanism and reinforcement learning algorithm, personalized classical music teaching content is generated, solving the shortcomings of existing data sets in multimodal fusion and teaching adaptation, and achieving a highly reliable immersive teaching experience.
Patent Information
- Application Number
- CN202510509129.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The existing classical music teaching data set lacks the deep integration of multimodal resources, making it difficult to achieve organic integration of cross-modal data and dynamic adaptation of teaching scenarios, resulting in insufficient logic and coherence of teaching content, and it is difficult to meet the diverse and personalized teaching needs.
By collecting multimodal data such as audio, video, music scores, etc., standardizing using preprocessing technology, using cross-modal attention mechanisms to fusion, building a knowledge graph, combining graph convolution networks and reinforcement learning algorithms to generate a sequence of personalized teaching content, and presenting immersive teaching scenarios through virtual reality technology to optimize the teaching sequence in real time.
It realizes efficient integration of multimodal data and intelligent generation of personalized teaching content, improves the interactiveness and learning experience of teaching, and meets the high reliability and universal needs of full-scene teaching.
Smart Images

Figure CN120561840A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a method for constructing a cross-modal, highly reliable experimental dataset of classical music for full-scenario teaching applications. Background Art
[0002] Problem background:
[0003] As an integral part of cultural heritage, classical music holds irreplaceable value in music education and cultural inheritance. With the advancement of digital technology, the application of cross-modal data in classical music instruction has become a research hotspot. By integrating multimodal resources such as audio, video, and sheet music, it can significantly enhance the immersiveness of instruction and the efficiency of knowledge transfer. However, current research and application still face numerous challenges, and a systematic dataset construction method is urgently needed to support full-scenario instructional needs.
[0004] Existing methods for constructing classical music instructional datasets are often limited to single-modal data, such as relying solely on audio or musical notation, and lack the deep integration of multimodal resources. Furthermore, the dataset construction process is often not optimized for teaching scenarios, and the data organization method struggles to adapt to learners' individual needs, resulting in fragmented teaching content and a lack of coherence in learning paths. These limitations make it difficult for existing datasets to achieve high reliability and universal applicability in complex teaching scenarios.
[0005] The core challenge lies in how to achieve the organic integration of cross-modal data and dynamic adaptation to teaching scenarios. First, the heterogeneity of cross-modal data makes it difficult to accurately establish semantic associations between data, affecting the logic and coherence of teaching content. Second, the lack of intelligent teaching arrangement makes it difficult to dynamically adjust the order of content presentation based on learners' cognitive level and interests. These unresolved technical factors have limited the application of datasets in full-scenario teaching, making it difficult to meet diverse and personalized teaching needs.
[0006] Therefore, how to design a highly reliable dataset construction method that can integrate multimodal data and dynamically generate personalized teaching sequences has become a key issue for full-scenario teaching applications. Summary of the Invention
[0007] The present invention provides a method for constructing a cross-modal, highly reliable experimental dataset of classical music for full-scenario teaching applications, which mainly includes:
[0008] By collecting multimodal data such as audio, video, and music scores, we construct an original data set, and use preprocessing technology to standardize the data of each modality to obtain a multimodal data set in a unified format;
[0009] Extract features from multimodal data sets. For audio data, Fourier transform is used to obtain spectral features. For video data, convolutional neural networks are used to extract visual features. For music score data, symbol parsing is used to obtain note sequences. The feature representation of each modality is obtained.
[0010] A cross-modal attention mechanism is used to perform weighted fusion of the feature representations of each modality and calculate the semantic similarity between features. If the similarity is greater than the preset threshold S (S is 0.8), a semantic association is established to obtain the fused semantic feature vector.
[0011] Based on the semantic feature vectors, a knowledge graph is constructed, where nodes represent music elements and edges represent semantic relationships between elements. A graph convolutional network is used to encode the knowledge graph to obtain a semantic representation of the teaching content.
[0012] Extract subgraphs from the semantic representation of teaching content, and use reinforcement learning algorithms to adjust the weights of subgraph nodes based on the cognitive level parameter C (C is an integer from 1 to 5) input by the learner to obtain a personalized teaching content sequence.
[0013] Through the sequence generation model, the personalized teaching content sequence is mapped into multimodal output, generating accompaniment for the audio modality, teaching animation for the video modality, and annotation for the music score modality, thus obtaining multimodal teaching resources;
[0014] Using virtual reality rendering technology, multimodal teaching resources are integrated into immersive teaching scenes, presenting content sequences in real time to achieve a dynamic teaching experience;
[0015] Based on learners' real-time feedback data, an online learning algorithm is used to update the parameters of the reinforcement learning model. If the feedback score is lower than the preset threshold T (T is 0.7), the content sequence is readjusted to obtain the optimized teaching sequence.
[0016] By continuously collecting learner interaction data, updating the knowledge graph and semantic feature vector, and using incremental learning algorithms to optimize the cross-modal fusion model, an adaptive teaching dataset is obtained.
[0017] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:
[0018] The present invention discloses an intelligent music teaching method based on multimodal data. The method collects multimodal data such as audio, video, and musical scores to construct an original data set and perform preprocessing. A cross-modal attention mechanism is used to fuse the features of each modality to construct a music knowledge graph. Based on the learner's cognitive level, a reinforcement learning algorithm is used to generate a personalized teaching content sequence and map it into a multimodal teaching resource. Virtual reality technology is used to present an immersive teaching scene, and the teaching sequence is optimized in real time based on learner feedback. The present invention realizes the intelligent generation and adaptive adjustment of teaching content, improves the personalization and interactivity of music teaching, and effectively enhances the teaching effect and learning experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a flowchart of a method for constructing a cross-modal, highly reliable experimental dataset of classical music for full-scenario teaching applications.
[0020] Figure 2 This is a schematic diagram of a method for constructing a cross-modal, highly reliable experimental dataset of classical music for full-scenario teaching applications according to the present invention.
[0021] Figure 3 This is another schematic diagram of the method for constructing a cross-modal, highly reliable experimental dataset of classical music for full-scenario teaching applications according to the present invention. DETAILED DESCRIPTION
[0022] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0023] like Figure 1-3 In this embodiment, a method for constructing a classical music cross-modal high-reliability experimental dataset for full-scenario teaching applications may specifically include:
[0024] In step S101, an original data set is constructed by collecting multimodal data such as audio, video, and music scores, and preprocessing technology is used to standardize the data of each modality to obtain a multimodal data set in a unified format.
[0025] Audio, video, and musical score data are acquired from the environment or musical instruments using sensors and recording devices to construct a raw dataset containing multimodal data. A format conversion tool is used to unify the sampling rate of the audio data to obtain standardized audio data; the frame rate and resolution of the video data are adjusted to obtain standardized video data; and the musical score data is symbolically serialized to obtain standardized musical score data. If the sampling rate of the standardized audio data matches a preset threshold, it is stored in a unified format dataset; if the frame rate and resolution of the standardized video data match a preset threshold, it is also stored in a unified format dataset; and if the symbol sequence of the standardized musical score data is complete, it is also stored in a unified format dataset, resulting in a preliminary integrated multimodal dataset. A feature extraction algorithm is used to extract the spectral features of the audio, the inter-frame features of the video, and the note sequence features of the musical score from the preliminary integrated multimodal dataset to obtain a multimodal feature set. Principal component analysis is used to reduce the dimensionality of the multimodal feature set to obtain a compressed feature dataset. If the dimensionality of the compressed feature dataset is below a preset threshold, data alignment techniques are used to synchronize the time axes of the audio, video, and musical score features to obtain a time-aligned fused dataset. The inter-modal correlation analysis of the time-aligned fusion dataset is performed using a clustering algorithm to obtain a unified representation of the multimodal data.
[0026] For example, when constructing a multimodal dataset, a musical performance was first recorded at a 44.1kHz sampling rate using an audio capture device, such as a microphone array. A 4K camera was used to simultaneously capture the performance at 60fps, and a standard MIDI-formatted musical score file was scanned. The audio data was spectrally analyzed using the FFT algorithm to extract 128-dimensional MFCC features. Keyframe extraction and optical flow calculation were performed on the video data using the OpenCV library. The musical score data was then parsed into MusicXML format using OMR technology. In the preprocessing phase, the audio was normalized to a range of -1dB to 1dB in amplitude. The video data was uniformly resized to a 1920x1080 resolution and converted to H.264 encoding. The musical score data was transformed using XSLT to ensure that all symbols conformed to the MusicXML 3.0 standard. Feature fusion employed a cross-modal attention mechanism, aligning the audio Mel-spectrogram with the video I3D features in the temporal dimension. Inter-modal correlation weights were calculated using a Transformer architecture with 8 attention heads and 512 hidden layer dimensions. During data augmentation, the audio is subjected to ±10% speed change and -6dB to 6dB gain perturbation, and the video data is subjected to random horizontal flipping and
[0027] The music data was rotated by ±15 degrees, and the music score data was augmented by randomly transposing it by ±2 semitones. The final dataset contained 1000 hours of synchronized multimodal data. Model performance was evaluated using 5-fold cross-validation, with a bidirectional LSTM network achieving 87.3% classification accuracy on the fused features.
[0028] Step S102: extract features from the multimodal data set, use Fourier transform to obtain spectral features for audio data, use convolutional neural network to extract visual features for video data, and use symbol analysis to obtain note sequences for music score data to obtain feature representations of each modality.
[0029] Audio data, video data, and musical score data are obtained from a multimodal dataset and stored as first audio data, first video data, and first musical score data, respectively, to obtain an initial data set. A Fourier transform is applied to the first audio data to calculate the frequency components of each time window and generate a first spectral feature. A convolutional neural network is applied to the first video data to extract the spatial features of the frame sequence and generate a first visual feature. Symbol parsing is applied to the first musical score data to identify note and rhythm information and generate a first note sequence. If the dimensions of the first spectral feature, the first visual feature, and the first note sequence are inconsistent, the high-dimensional features are downsampled using linear interpolation to obtain a second spectral feature, a second visual feature, and a second note sequence. Based on the second spectral feature, the second visual feature, and the second note sequence, a principal component analysis algorithm is applied to extract the common patterns of the features of each modality and obtain a fused feature representation. Using the fused feature representation, a clustering algorithm is used to group the multimodal data and obtain a classification result.
[0030] For example, in audio data processing, the audio signal is first sampled at a sampling rate of 44100 Hz, and then the sampled signal is Fourier transformed to convert the time domain signal into a frequency domain signal to obtain a spectrum feature.
[0031] For example, a 2-second audio signal can be decomposed into 1024 frequency components using the Fast Fourier Transform (FFT) algorithm, with each component corresponding to a specific frequency range, thereby obtaining the audio's spectral characteristics. In video data processing, convolutional neural networks (CNNs) are used to extract visual features.
[0032] For example, a video frame with a resolution of 1920x1080 is first resized to 224x224 and then fed into a pre-trained ResNet-50 model. Through convolutional and pooling layers, a 2048-dimensional feature vector is ultimately obtained, representing the visual features of the frame. In music score data processing, symbol parsing methods are used to obtain the note sequence.
[0033] For example, for a 16-bar musical score, we first convert it to MIDI format. Then, by parsing the MIDI file, we extract the pitch, duration, and start time of each note, ultimately obtaining a note sequence that represents the symbolic features of the score. This approach allows us to extract modal features from audio, video, and musical score data, providing a foundation for subsequent multimodal fusion analysis.
[0034] In step S103, a cross-modal attention mechanism is used to perform weighted fusion on the feature representations of each modality and calculate the semantic similarity between the features. If the similarity is greater than a preset threshold S (S is 0.8), a semantic association is established to obtain a fused semantic feature vector.
[0035] Initial feature representations are obtained from the audio modality, video modality, and musical score modality and stored as first audio features, first video features, and first musical score features, respectively, to obtain an initial feature set. A cross-modal attention mechanism is used to weight the first audio features, first video features, and first musical score features, and the attention weights between the features of each modality are calculated to obtain second audio features, second video features, and second musical score features. A semantic similarity matrix is calculated for the second audio features, second video features, and second musical score features. If an element in the matrix is greater than a preset threshold S, it is determined that the corresponding feature pairs have a semantic association, and a semantic association set is obtained. Based on the semantic association set, a linear transformation is used to fuse the second audio features, second video features, and second musical score features to generate a first fused feature vector. The first fused feature vector is subjected to dimensionality reduction processing using a principal component analysis algorithm to extract the main semantic components, resulting in a second fused feature vector. If the dimension of the second fused feature vector is greater than the preset dimension D, it is downsampled using linear interpolation to obtain a third fused feature vector. According to the third fused feature vector, the cosine similarity is used to calculate its matching degree with the preset semantic template, and the semantic category with the highest matching degree is determined to obtain the final semantic classification result.
[0036] For example, in the cross-modal attention mechanism, the feature representations of the two modalities, text and image, are first extracted. The text features are generated into 768-dimensional vectors through the BERT model, and the image features are generated into 2048-dimensional vectors through the ResNet-50 model. Then, the attention mechanism is used to calculate the semantic similarity between the text and image features. Specifically, the cosine similarity algorithm is used, and the formula is sim(T,I) = (T·I) /
[0037] (||T||*||I||), where T is the text feature vector and I is the image feature vector. Assuming the calculated similarity is 0.85, which is greater than the preset threshold of 0.8, a semantic association is established. Then, the text and image features are fused using a weighted fusion method, with the weights dynamically adjusted based on the similarity. The fused semantic feature vector is F = 0.6*T + 0.4*I. Finally, the fused feature vector is input into a classifier for classification. The classifier uses a Softmax function and outputs a probability distribution for each category, thus completing the cross-modal semantic feature fusion and classification task.
[0038] Step S104: construct a knowledge graph based on the semantic feature vector, where nodes represent music elements and edges represent semantic relationships between elements. A graph convolutional network is used to encode the knowledge graph to obtain a semantic representation of the teaching content.
[0039] Semantic feature vectors of musical elements are obtained, and initial vector representations are generated using a preset semantic analysis model to obtain a semantic feature set for the musical elements. Based on this semantic feature set, a knowledge graph is constructed, with nodes consisting of musical elements and edges representing semantic association weights between elements, generating an initial knowledge graph structure. A graph convolutional network is used to encode the initial knowledge graph, and node features and edge weights are iteratively updated to obtain an encoded graph feature representation. If the encoded graph feature representation meets the preset convergence criteria, the semantic representations of each node in the graph are extracted to generate a preliminary semantic representation of the teaching content. If not, the graph convolutional network parameters are adjusted and the encoding is re-encoded. The preliminary semantic representation is optimized using a preset semantic mapping model to generate a final semantic representation that matches the teaching content. Based on the final semantic representation, a clustering algorithm is used to classify the teaching content, obtaining a classified teaching content set. For this classified teaching content set, a structured teaching content output is generated to determine the final teaching content representation.
[0040] For example, when constructing a knowledge graph of music elements, the semantic feature vectors of the music elements are first extracted. For example, "melody" is represented as [0.85, 0.12, 0.34], and "rhythm" is represented as [0.45, 0.78, 0.23]. These vectors are calculated from the music text data using the TF-IDF algorithm. Then, a knowledge graph is constructed based on these feature vectors. The nodes represent the music elements, and the edges represent the semantic relationships between the elements. For example, the edge weight between "melody" and "rhythm" is 0.67, indicating that there is a strong semantic association between them. In order to encode the knowledge graph, a graph convolutional network (GCN) is used. The input layer receives the feature vector of the node, the hidden layer uses the ReLU activation function, and the output layer generates a semantic representation of each node.
[0041] For example, after two layers of GCN processing, the semantic representation of the "melody" node is updated to [0.92, 0.15, 0.38], while the semantic representation of the "rhythm" node is updated to [0.48, 0.81, 0.25]. Finally, these semantic representations are used for semantic analysis of teaching content. For example, by calculating cosine similarity, it is determined that the similarity between "melody" and teaching content A is 0.89, and the similarity between "melody" and teaching content B is 0.76, thus providing a basis for recommending teaching content.
[0042] Step S105: extract a subgraph from the semantic representation of the teaching content, and use a reinforcement learning algorithm to adjust the weights of the subgraph nodes according to the cognitive level parameter C (C is an integer from 1 to 5) input by the learner to obtain a personalized teaching content sequence.
[0043] Semantic representations are obtained from the teaching content database. Subgraphs are extracted using a preset semantic analysis model to obtain an initial subgraph structure. Based on the learner's cognitive level parameter C, a reinforcement learning algorithm is used to calculate the initial weights of each node in the subgraph and determine the weight distribution. If the weight distribution deviates from the preset personalized target by more than a threshold, the reinforcement learning algorithm is used to iteratively adjust the node weights to obtain an optimized subgraph. Based on the optimized subgraph, a sequence generation model is used to generate candidate teaching content sequences to obtain a sequence set. Based on the learner's cognitive level parameter C, the sequence set is screened for teaching content sequences that match parameter C to determine the optimal sequence. If the optimal sequence's matching score falls below a preset threshold, the sequence generation model is used to regenerate the sequence, resulting in the final personalized teaching content sequence. Based on the final personalized teaching content sequence, structured teaching content data is output and the delivery format is determined. The seven steps described above, including semantic representation, subgraph extraction, weight optimization, sequence generation, and screening, form a logical chain. The output of each step serves as the input for the next step, ensuring the achievement of business objectives.
[0044] For example, when extracting a subgraph from the semantic representation of teaching content, it is first necessary to construct a knowledge graph, with the knowledge points in the teaching content as nodes and the relationships between the knowledge points as edges.
[0045] For example, in an article about "machine learning," the knowledge points "supervised learning" and "unsupervised learning" can serve as nodes, and the relationship between them, "classification method," can serve as an edge. Natural language processing technology is used to extract keywords and semantic relationships from the article to construct a knowledge graph. Next, a reinforcement learning algorithm is used to adjust the subgraph node weights based on the cognitive level parameter C (assuming C is 3) input by the learner. The reinforcement learning algorithm can use Q-learning to optimize node weights by defining the state (current subgraph node), action (selecting the next node), and reward (learner's level of understanding).
[0046] For example, if a learner has a high level of understanding of "supervised learning," the algorithm increases the weight of the "classification method" node connected to that node, prioritizing related content recommendations in subsequent learning. Through multiple iterations, the algorithm gradually adjusts the weights of subgraph nodes to generate a personalized sequence of teaching content.
[0047] For example, the resulting sequence might be "supervised learning" → "classification method" → "logistic regression," ensuring that learners can gradually master knowledge based on their own cognitive level. The entire process is automated through information technology, eliminating the need for human intervention and enabling intelligent and personalized recommendations for instructional content.
[0048] Step S106: Map the personalized teaching content sequence into a multimodal output through a sequence generation model, generate accompaniment for the audio modality, generate teaching animation for the video modality, and generate annotations for the music score modality to obtain multimodal teaching resources.
[0049] A sequence generation model is used to obtain a personalized teaching content sequence. An encoder is used to extract features from the sequence to obtain a content feature vector. A variational autoencoder is used to perform modal decomposition on the content feature vector, generating audio, video, and musical score modal features, resulting in a multimodal feature set. For the audio modal features, if the feature vector matches a preset accompaniment style template, a recurrent neural network is used to generate an accompaniment sequence, resulting in audio accompaniment data. For the video modal features, a generative adversarial network is used to map the feature vector to an animation frame sequence, synthesizing the teaching animation to obtain video animation data. For the musical score modal features, if the feature vector contains a note sequence, a sequence labeling algorithm is used to generate musical score annotations, resulting in musical score annotation data. The audio accompaniment data, video animation data, and musical score annotation data are obtained and timeline synchronized using a multimodal fusion algorithm to obtain a multimodal teaching resource. The content sequence verification module then performs consistency testing on the multimodal teaching resource. If the test results meet a preset consistency threshold, the final multimodal teaching resource is output.
[0050] For example, in the sequence generation model, BERT is first used to extract the semantic features of the personalized teaching text. For example, when the input is "periodic explanation of trigonometric functions", a 768-dimensional vector representation is output, which is encoded into a 128-dimensional time series feature by LSTM. For the audio modality, the DiffWave model is used to generate accompaniment with a frame length of 50ms, and the Mel spectrum constraint is used to ensure synchronization with the explanation rhythm. For example, when the keyword "y=sin(x)" is detected, a string timbre with a fundamental frequency of 440Hz is generated. The video modality uses StyleGAN-V to generate 1280×720 resolution animations, and the content consistency is controlled based on the CLIP text alignment loss. When the phase of the explanation changes, the dynamic visualization of the sine wave is automatically inserted, and the inter-frame optical flow loss is kept below 0.03 to ensure a smooth transition. The music score annotation is generated by MusicTransformer, and the explanation nodes are aligned with the minimum unit of 0.25 beats. It is automatically inserted when analyzing the concept of "amplitude change". <forte>Dynamics marking uses an HMM algorithm to ensure that musical note timing errors are less than 5%. A cross-modal attention mechanism is employed in the multimodal fusion stage to control audio-video synchronization errors within ±80ms. A trimodal contrast loss function is used to achieve a cosine similarity in the feature space exceeding 0.85. The final output is packaged in an MP4 container, with a video bitrate of 5000kbps and audio encoded using AAC-LC. The musical notation is embedded in the metadata area in SVG vector format.
[0051] Step S107 , using virtual reality rendering technology, integrates multimodal teaching resources into an immersive teaching scene, presents content sequences in real time, and obtains a dynamic teaching experience.
[0052] Multimodal teaching resource data is acquired through virtual reality technology, and initial immersive scene data is generated through integration processing to obtain a basic scene model. If the resolution of the initial immersive scene data is lower than a preset threshold, the basic scene model is optimized through rendering technology to obtain a high-resolution scene model. Based on the high-resolution scene model, a dynamic content sequence is generated using a real-time rendering algorithm to obtain a content sequence dataset. If the interactive response time of the content sequence dataset exceeds a preset threshold, the dynamic content sequence is adjusted using an interactive experience optimization algorithm to obtain an optimized content sequence. User interaction data is obtained from the optimized content sequence, and the immersive scene is updated using scene generation technology to obtain a dynamic interactive scene. The dynamic interactive scene is analyzed in real time using experience optimization technology to determine whether the user experience data meets the preset threshold, thereby obtaining the final interactive teaching experience.
[0053] For example, in virtual reality rendering technology, a geometric model of the teaching scene is first constructed using 3D modeling tools such as Blender or Unity. The model accuracy is controlled within 1 mm to ensure realism, and PBR material maps are used to enhance surface details. Then, a ray tracing algorithm (such as path tracing) is used for real-time rendering. The number of samples per pixel is set to 64 to reduce noise, and the rendering resolution is oversampled from 1080p to 4K using DLSS technology to improve performance. In the multimodal resource integration stage, text, audio, video and other resources are encoded into a unified format. For example, H.265 is used to compress the video stream with a bit rate set to 8Mbps, and audio is encoded using AAC with a sampling rate of 48kHz. Multimodal data synchronization is ensured through spatiotemporal alignment algorithms (such as dynamic time warping (DTW)), with an error control within ±50 milliseconds. When the content sequence is presented in real time, an LSTM-based prediction model is used to preload the next 3 seconds of teaching content, and the buffer size is set to 500MB to avoid lag. The dynamic teaching experience relies on eye tracking technology, which collects user gaze data at a frequency of 120Hz. Combined with Foveated Rendering, it dynamically adjusts rendering quality, maintaining 4K resolution in the gaze area while reducing it to 720p in surrounding areas, reducing GPU load by 40%. Finally, through user behavior data analysis (such as the K-means clustering algorithm with k=5), the scene layout is continuously optimized to ensure that more than 90% of users can locate the core teaching elements within 2 seconds.
[0054] In step S108, based on the learner's real-time feedback data, an online learning algorithm is used to update the reinforcement learning model parameters. If the feedback score is lower than the preset threshold T (T is 0.7), the content sequence is readjusted to obtain an optimized teaching sequence.
[0055] Real-time feedback data from learners is obtained, and a real-time feedback score is generated through data cleaning and feature extraction. If the real-time feedback score is lower than the preset threshold, an online update of the reinforcement learning model is triggered, and the model parameters are adjusted using the online learning algorithm to obtain the updated model parameters. Based on the updated model parameters, the reinforcement learning model predicts the score of the content sequence to generate a set of candidate content sequences. For the set of candidate content sequences, the sequence scoring function is used to calculate the expected score of each sequence to obtain the optimized teaching sequence with the highest score. By optimizing the teaching sequence, the presentation order of the learning content is adjusted to generate an adjusted content sequence. The real-time feedback score of the adjusted content sequence is obtained and compared with the preset threshold to determine whether to continue adjusting the sequence to obtain the final teaching sequence. If the feedback score of the final teaching sequence is still lower than the preset threshold, new features are extracted through feedback data processing, and the online learning algorithm is updated to obtain new model parameters.
[0056] For example, in a real-time learning scenario, the system updates the weight parameters of the reinforcement learning model using an online gradient descent algorithm, such as the Adam optimizer with a learning rate of η = 0.01. For each learner interaction data entry (e.g., a correct answer rate of 0.65), the system immediately calculates the time-delay error (TD) δ = 0.12 between the action value Q(s, a) output by the current policy network and the target value. When the average of five consecutive feedback scores of 0.68 falls below a threshold of T = 0.7, the content sequence reorganization mechanism is triggered. First, the K-means clustering algorithm (k = 3) is used to group the knowledge point mastery feature vectors [0.4, 0.7, 0.3] in the historical interaction data. Item response theory is then used to calculate the difficulty parameter b = 1.2 and the discrimination a = 0 for each knowledge point.
[0057] 8. Then, based on a Markov decision process, a state transition matrix P is constructed. The transition probability from state s_t to s_{t+1} is calculated using the Bayesian update formula P(s'|s,a)=N(μ,σ²). The mean μ is dynamically adjusted based on the most recent 20 pieces of feedback data. The effectiveness of the restructured instructional sequence was verified through A / B testing. The experimental group using the new sequence saw an improvement in NDCG@3 to 0.82, significantly higher than the control group's original sequence of 0.
[0058] 71. A sliding window mechanism is used to maintain a 100-byte priority experience replay buffer, prioritizing samples with a TD error δ > 0.1 for model training, ensuring a 35% increase in key data utilization. For special cases, such as isolated nodes with a cosine similarity below 0.3 in the knowledge point relevance matrix, a content importance reordering module based on the PageRank algorithm is activated, advancing the teaching order of core knowledge points by 2-3 positions.
[0059] Step S109: By continuously collecting learner interaction data, updating the knowledge graph and semantic feature vector, and using the incremental learning algorithm to optimize the cross-modal fusion model, an adaptive teaching data set is obtained.
[0060] An initial interaction dataset is generated by collecting learner interaction data. Structured information is extracted from the initial interaction dataset, and the knowledge graph is updated to obtain an updated knowledge graph. Based on the updated knowledge graph, semantic features are generated and converted into feature vectors to obtain a semantic feature vector. If the dimension of the semantic feature vector meets the preset threshold, the semantic feature vector is processed using a cross-modal fusion algorithm to generate a fusion model. If not, the feature vector dimension is adjusted and reprocessed to obtain a fusion model. The fusion model is optimized using an incremental learning algorithm, and the model parameters are updated to obtain an optimized fusion model. Based on the optimized fusion model, an adaptive teaching dataset is generated to obtain an adaptive teaching dataset. Feedback information is extracted from the adaptive teaching dataset, and the learner interaction data is updated to obtain an updated interaction dataset.
[0061] For example, during the continuous collection of learner interaction data, tracking technology is used to record learners' clickstreams, dwell time (e.g., video viewing time exceeding 90 seconds triggers knowledge point marking), and question accuracy (e.g., math problem accuracy below 60% is marked as a weak point) on the teaching platform. This data is collected in real time and stored in HDFS using the Flume framework at a throughput of 200 records per second. The knowledge graph update phase utilizes the Neo4j graph database. When it detects that more than 30% of learners associate "quadratic function" with "extreme value problem" in their queries, an edge relationship between these two types of nodes is automatically added with a weight of 0.85. The TransR algorithm is also used to expand the node vector dimension from 128 to 256 to improve representational capabilities. Semantic feature vector optimization utilizes incremental training of the BERT model. For every 5,000 new study notes, the original model is fine-tuned for two epochs at a learning rate of 0.0001, resulting in an improvement in the cosine similarity of "Newton's Law" in vector space from 0.72 to 0.89. The cross-modal fusion model uses the MMoE architecture. When the ratio of video viewing data to exercise data reaches 1:3, a dynamic weight adjustment algorithm is used to reduce the visual modality gateway weight from 0.6 to 0.45, and the text modality gateway weight is correspondingly increased by 0.15. Incremental learning uses the EWC algorithm, calculating the diagonal value of the Fisher information matrix when the model parameters are updated (for example, the fully connected layer parameter importance score is locked when it is greater than 0.7), ensuring that the historical accuracy rate does not drop by more than 2%. The resulting adaptive teaching dataset is divided into knowledge unit difficulty levels through K-means clustering (k=8, silhouette coefficient 0.65), and dynamically recommends learning paths based on Markov chain prediction (state transition probability matrix dimension 20×20), achieving a recommendation hit rate of 83% for highly relevant knowledge points. The entire system is updated in a closed loop on a daily basis, and the knowledge graph reconstruction process is automatically triggered when the seven-day retention rate of a knowledge point falls below 40%.
[0062] The above only lists some preferred embodiments of the present invention, but the present invention is not limited thereto, and many improvements and modifications can be made. As long as the improvements and modifications are made on the basis of the basic principles of the present invention, they should be considered to fall within the scope of protection of the present invention.< / forte>
Claims
1. A method for constructing a cross-modal, highly reliable experimental dataset of classical music for full-scenario teaching applications, characterized by: The method comprises: By collecting multimodal data of audio, video and music scores, the original data set is constructed, and preprocessing technology is used to standardize the data of each modality to obtain a multimodal data set in a unified format; Extract features from multimodal data sets. For audio data, Fourier transform is used to obtain spectral features. For video data, convolutional neural networks are used to extract visual features. For music score data, symbol parsing is used to obtain note sequences. The feature representation of each modality is obtained. A cross-modal attention mechanism is used to perform weighted fusion of the feature representations of each modality and calculate the semantic similarity between features. If the similarity is greater than the preset threshold S (S is 0.8), a semantic association is established to obtain the fused semantic feature vector. Based on the semantic feature vectors, a knowledge graph is constructed, where nodes represent music elements and edges represent semantic relationships between elements. A graph convolutional network is used to encode the knowledge graph to obtain a semantic representation of the teaching content. Extract subgraphs from the semantic representation of teaching content, and use reinforcement learning algorithms to adjust the weights of subgraph nodes based on the cognitive level parameter C (C is an integer from 1 to 5) input by the learner to obtain a personalized teaching content sequence. Through the sequence generation model, the personalized teaching content sequence is mapped into multimodal output, generating accompaniment for the audio modality, teaching animation for the video modality, and annotation for the music score modality, thus obtaining multimodal teaching resources; Using virtual reality rendering technology, multimodal teaching resources are integrated into immersive teaching scenes, presenting content sequences in real time to achieve a dynamic teaching experience; Based on learners' real-time feedback data, an online learning algorithm is used to update the parameters of the reinforcement learning model. If the feedback score is lower than the preset threshold T (T is 0.7), the content sequence is readjusted to obtain the optimized teaching sequence. By continuously collecting learner interaction data, updating the knowledge graph and semantic feature vector, and using incremental learning algorithms to optimize the cross-modal fusion model, an adaptive teaching dataset is obtained.
2. The method according to claim 1, characterized in that The method involves collecting multimodal data such as audio, video, and music scores to construct an original data set, and using preprocessing technology to standardize each modal data to obtain a multimodal data set in a unified format, including: Acquire audio data, video data, and musical score data from the environment or musical instruments through sensors and recording devices to construct a raw dataset containing multimodal data; Use format conversion tools to unify the sampling rate of audio data to obtain standardized audio data; Adjusting the frame rate and resolution of the video data to obtain standardized video data; Perform symbol serialization on the music score data to obtain standardized music score data; If the sampling rate of the normalized audio data is consistent with the preset threshold, it is stored in a unified format dataset; If the frame rate and resolution of the normalized video data match a preset threshold, it is stored in a unified format dataset; If the symbol sequence of the standardized music score data is complete, it is stored in a unified format dataset to obtain a preliminary integrated multimodal dataset; Using a feature extraction algorithm, the audio spectrum features, video frame features, and musical note sequence features are obtained from the preliminarily integrated multimodal dataset to obtain a multimodal feature set. The principal component analysis algorithm is used to reduce the dimensionality of the multimodal feature set to obtain a compressed feature data set; If the dimension of the compressed feature dataset is lower than a preset threshold, the audio features, video features, and music score features are synchronized on the time axis using data alignment technology to obtain a time-aligned fused dataset. The inter-modal correlation analysis of the time-aligned fusion dataset is performed using a clustering algorithm to obtain a unified representation of the multimodal data.
3. The method according to claim 1, characterized in that The method extracts features from a multimodal data set, uses Fourier transform to obtain spectral features for audio data, uses convolutional neural networks to extract visual features for video data, and uses symbol parsing to obtain note sequences for musical score data, to obtain feature representations of each modality, including: Acquiring audio data, video data, and music score data from a multimodal dataset, storing the data as first audio data, first video data, and first music score data, respectively, to obtain an initial data set; Applying Fourier transform to the first audio data to calculate the frequency components of each time window and generate a first frequency spectrum feature; A convolutional neural network is used for the first video data to extract spatial features of the frame sequence and generate a first visual feature; Using symbol analysis on the first musical score data to identify note and rhythm information and generate a first note sequence; If the dimensions of the first spectral feature, the first visual feature, and the first note sequence are inconsistent, downsampling the high-dimensional features by linear interpolation to obtain a second spectral feature, a second visual feature, and a second note sequence; According to the second spectrum feature, the second visual feature and the second note sequence, a principal component analysis algorithm is used to extract the common pattern of each modal feature to obtain a fusion feature representation; By fusing feature representations, clustering algorithms are used to group multimodal data and obtain classification results.
4. The method according to claim 1, wherein The cross-modal attention mechanism is used to perform weighted fusion of the feature representations of each modality and calculate the semantic similarity between features. If the similarity is greater than the preset threshold S (S is 0.8), a semantic association is established to obtain the fused semantic feature vector, including: Obtaining initial feature representations from the audio modality, the video modality, and the music score modality, storing them as a first audio feature, a first video feature, and a first music score feature, respectively, to obtain an initial feature set; A cross-modal attention mechanism is used to perform weighted processing on the first audio feature, the first video feature, and the first music score feature, and the attention weights between the modal features are calculated to obtain the second audio feature, the second video feature, and the second music score feature. For the second audio feature, the second video feature, and the second musical score feature, a semantic similarity matrix is calculated between the features. If an element in the matrix is greater than a preset threshold S, it is determined that the corresponding feature pair has a semantic association, and a semantic association set is obtained. According to the semantic association set, a linear transformation is used to fuse the second audio feature, the second video feature, and the second music score feature to generate a first fused feature vector; Using the principal component analysis algorithm, the first fused feature vector is subjected to dimensionality reduction processing to extract the main semantic components and obtain the second fused feature vector; If the dimension of the second fused feature vector is greater than the preset dimension D, downsampling it by linear interpolation to obtain a third fused feature vector; According to the third fused feature vector, the cosine similarity is used to calculate its matching degree with the preset semantic template, and the semantic category with the highest matching degree is determined to obtain the final semantic classification result.
5. The method according to claim 1, characterized in that The method constructs a knowledge graph based on the semantic feature vector, where nodes represent music elements and edges represent semantic relationships between elements. The knowledge graph is encoded using a graph convolutional network to obtain a semantic representation of the teaching content, including: Obtaining the semantic feature vector of the music element, generating an initial vector representation through a preset semantic analysis model, and obtaining a semantic feature set of the music element; Based on the semantic feature set of music elements, a knowledge graph is constructed. The nodes are composed of music elements, and the edges are represented by the semantic association weights between elements, thus generating the initial knowledge graph structure. A graph convolutional network is used to encode the initial knowledge graph, and node features and edge weights are iteratively updated to obtain the encoded graph feature representation. If the encoded graph feature representation meets the preset convergence conditions, the semantic representation of each node in the graph is extracted to generate a preliminary semantic representation of the teaching content; If not, adjust the graph convolutional network parameters and re-encode; Through the preset semantic mapping model, the preliminary semantic representation is optimized to generate the final semantic representation that matches the teaching content; According to the final semantic representation, clustering algorithm is used to classify the teaching content to obtain the classified teaching content set; For the classified teaching content set, structured teaching content output is generated and the final teaching content representation is determined.
6. The method according to claim 1, characterized in that The method extracts a subgraph from the semantic representation of the teaching content, and uses a reinforcement learning algorithm to adjust the weights of the subgraph nodes according to the cognitive level parameter C (C is an integer between 1 and 5) input by the learner to obtain a personalized teaching content sequence, including: Obtain semantic representation from the teaching content database, extract subgraphs through a preset semantic analysis model, and obtain the initial subgraph structure; According to the cognitive level parameter C input by the learner, the reinforcement learning algorithm is used to calculate the initial weight of each node in the subgraph and determine the weight distribution; If the deviation between the weight distribution and the preset personalized target exceeds a threshold, the node weights are iteratively adjusted through the reinforcement learning algorithm to obtain the optimized subgraph; Through the optimized subgraph, a sequence generation model is used to generate candidate teaching content sequences to obtain a sequence set; According to the cognitive level parameter C input by the learner, the teaching content sequence that matches the parameter C is selected from the sequence set to determine the optimal sequence; If the matching degree of the optimal sequence is lower than the preset threshold, the sequence generation model is returned to regenerate the sequence to obtain the final personalized teaching content sequence; Output structured teaching content data and determine the delivery format through the final personalized teaching content sequence; The above seven steps form a logical chain through semantic representation, subgraph extraction, weight optimization, and sequence generation and screening. The output of the previous step serves as the input of the next step to ensure the realization of business goals.
7. The method according to claim 1, characterized in that The personalized teaching content sequence is mapped into a multimodal output through the sequence generation model, accompaniment is generated for the audio modality, teaching animation is generated for the video modality, and annotation is generated for the music score modality, thereby obtaining multimodal teaching resources, including: The personalized teaching content sequence is obtained through the sequence generation model, and the encoder is used to extract features from the sequence to obtain the content feature vector; A variational autoencoder is used to perform modal decomposition on the content feature vector to generate audio modal features, video modal features, and music score modal features, thereby obtaining a multimodal feature set. For audio modal features, if the feature vector matches the preset accompaniment style template, the accompaniment sequence is generated through a recurrent neural network to obtain audio accompaniment data; Based on the video modality features, the feature vector is mapped into an animation frame sequence through a generative adversarial network, and the teaching animation is synthesized to obtain video animation data; For the music score modal features, if the feature vector contains a note sequence, the music score annotation is generated through the sequence annotation algorithm to obtain the music score annotation data; Acquire audio accompaniment data, video animation data, and music notation data, use a multimodal fusion algorithm to synchronize the time axis, and obtain multimodal teaching resources; The multimodal teaching resources are tested for consistency through the content sequence verification module. If the test result meets the preset consistency threshold, the final multimodal teaching resources are output.
8. The method according to claim 1, characterized in that The virtual reality rendering technology is used to integrate multimodal teaching resources into an immersive teaching scene, presenting content sequences in real time to obtain a dynamic teaching experience, including: Using virtual reality technology to obtain multimodal teaching resource data, the initial immersive scene data is generated by integration processing to obtain the scene basic model; If the resolution of the initial immersive scene data is lower than a preset threshold, the scene base model is optimized through rendering technology to obtain a high-resolution scene model; Based on the high-resolution scene model, a real-time rendering algorithm is used to generate dynamic content sequences to obtain a content sequence dataset; If the interactive response time of the content sequence data set exceeds a preset threshold, the dynamic content sequence is adjusted using an interactive experience optimization algorithm to obtain an optimized content sequence; Obtain user interaction data from the optimized content sequence, use scene generation technology to update the immersive scene, and obtain a dynamic interactive scene; Through experience optimization technology, dynamic interactive scenarios are analyzed in real time to determine whether the user experience data meets the preset threshold and obtain the final interactive teaching experience.
9. The method according to claim 1, characterized in that The online learning algorithm is used to update the reinforcement learning model parameters based on the learner's real-time feedback data. If the feedback score is lower than the preset threshold T (T is 0.7), the content sequence is readjusted to obtain an optimized teaching sequence, including: Obtain learners' real-time feedback data, and generate real-time feedback scores through data cleaning and feature extraction; If the real-time feedback score is lower than the preset threshold, the online update of the reinforcement learning model is triggered, and the model parameters are adjusted using the online learning algorithm to obtain the updated model parameters; Based on the updated model parameters, the reinforcement learning model is used to predict the ratings of the content sequences and generate a set of candidate content sequences. For the candidate content sequence set, the sequence scoring function is used to calculate the expected score of each sequence, and the optimized teaching sequence with the highest score is obtained; By optimizing the teaching sequence, the presentation order of learning content is adjusted to generate an adjusted content sequence; Obtain real-time feedback scores of the adjusted content sequence, and determine whether to continue adjusting the sequence by comparing it with a preset threshold to obtain the final teaching sequence; If the feedback score of the final teaching sequence is still lower than the preset threshold, new features are extracted through feedback data processing, and the online learning algorithm is updated to obtain new model parameters.
10. The method according to claim 1, characterized in that The method continuously collects learner interaction data, updates the knowledge graph and semantic feature vectors, and uses an incremental learning algorithm to optimize the cross-modal fusion model to obtain an adaptive teaching dataset, including: Generate an initial interaction dataset by collecting learner interaction data; Extract structured information from the initial interaction dataset, update the knowledge graph, and obtain an updated knowledge graph; Generate semantic features based on the updated knowledge graph, convert them into feature vectors, and obtain semantic feature vectors; If the dimension of the semantic feature vector meets the preset threshold, the cross-modal fusion algorithm is used to process the semantic feature vector to generate a fusion model; If it is not satisfied, the feature vector dimension is adjusted and reprocessed to obtain the fusion model; The fusion model is optimized through the incremental learning algorithm, the model parameters are updated, and the optimized fusion model is obtained; Generate an adaptive teaching data set according to the optimized fusion model to obtain an adaptive teaching data set; Feedback information is extracted from the adaptive teaching dataset, and the learner interaction data is updated to obtain an updated interaction dataset.
Citation Information
Patent Citations
Audio synthesis method and device, computer equipment and storage medium
CN114360492A
Zheng audio playing method and device, electronic equipment and computer readable medium
CN114758639A
Data processing method and device, electronic equipment and storage medium
CN115115913A
Conversation method and device for service robot and service robot conversation system
CN118778818A
VR interaction method and device based on meta-universe virtual reality technology
CN118860156A
Cited By
Knowledge graph-driven textbook automatic generation method and system
CN120832870A
A knowledge graph driven textbook automatic generation method and system
CN120832870B
College music education creative course visualization method and system
CN121070228A
Teaching three-dimensional target optimization system and method
CN121685851A
Knowledge graph and multi-mode intelligent interaction integrated oral medicine teaching system
CN121999666A