An educational resource management method and system based on multimodal semantic analysis
By using multimodal semantic analysis methods in the education resource management system, the education resource data and user-related data are processed and resource representation vectors are generated, which solves the problem that existing systems cannot accurately search and recommend multimedia resources, and achieves efficient and accurate resource management.
Patent Information
- Application Number
- CN202510328193.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The existing educational resource management system cannot conduct accurate resource searches and recommendations, especially when dealing with multimedia resources, the recognition accuracy is unstable and it is difficult to deal with complex or inconsistent information.
Using a method based on multimodal semantic analysis, by obtaining educational resource data and user-related data, it is input into the text feature model, image feature model, audio feature model and video feature model, and each modal feature vector is calculated, and attention weight is calculated based on the multimodal semantic fusion model, resource representation vector is generated, and finally clustering and scoring is performed to output the educational resources retrieved by the user.
It realizes accurate resource search and recommendation of the educational resource management system, can effectively process multimedia resources, improve identification accuracy and system efficiency.
Smart Images

Figure CN119848326B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of educational resource management, and in particular, to an educational resource management method and system based on multimodal semantic analysis. Background Art
[0002] Currently, with the development and popularization of Internet technology, more and more learning resources have flooded into the educational resource system. Most traditional educational resource management methods are based on means such as classification and keyword search. However, this method can only retrieve text information and cannot meet the retrieval requirements for multimedia resources such as videos and audios.
[0003] In an existing technology, image recognition technology based on deep learning is used to automatically process image resources; text analysis technology is used to identify and segment long text resources; audio recognition technology is used to process audio resources, etc. These methods have improved the accuracy and efficiency of resource management to a certain extent, but there are also some drawbacks: Image recognition has good effects for high-quality images, but due to differences in image quality and resolution, the recognition accuracy is not stable. Text analysis technology has certain requirements for the quality and format of text and is difficult to process complex or inconsistent formats of information. Audio recognition technology faces similar problems to image recognition, and the quality and format of audio affect the accuracy of recognition.
[0004] There is a problem in the existing technology that the educational resource management system cannot perform accurate resource search and recommendation. Summary of the Invention
[0005] The present invention provides an educational resource management method and system based on multimodal semantic analysis to solve the problem that the educational resource management system cannot perform accurate resource search and recommendation.
[0006] In a first aspect, to solve the above technical problem, the present invention provides an educational resource management method based on multimodal semantic analysis, including:
[0007] Obtain educational resource data and user-related data; wherein, the user-related data is the resource name clicked by the user, the knowledge points corresponding to the resource, and the tag data corresponding to the resource;
[0008] Input the educational resource data and the user-related data into a preset text feature model, image feature model, audio feature model, and video feature model respectively to obtain text features, image features, audio features, and video features;
[0009] Calculate text feature vectors, image feature vectors, audio feature vectors, and video feature vectors respectively according to the text features, the image features, the audio features, and the video features;
[0010] Based on the multi-modal semantic fusion model, calculate the attention weights among the text feature vector, the image feature vector, the audio feature vector, and the video feature vector;
[0011] According to the attention weights, after weighted averaging the text feature vector, the image feature vector, the audio feature vector, and the video feature vector, obtain a resource representation vector;
[0012] Input the resource representation vector into the clustering library, and output all resources whose distance from the resource representation vector is less than a preset target distance as the educational resources retrieved by the user;
[0013] Score all the educational resources retrieved by the user, and perform an output operation on the educational resources whose scores exceed a pre-configured optimal score threshold.
[0014] In an alternative embodiment, the training process of the text feature model, the image feature model, the audio feature model, and the video feature model includes:
[0015] Obtain historical text data, historical image data, historical audio data, and historical video data;
[0016] Perform data preprocessing on the historical text data to obtain text normalized data;
[0017] Use the text normalized data as input data, construct an initial text feature model based on the Glove algorithm, and train the initial text feature model according to the sample text data to obtain a text feature model;
[0018] Perform data preprocessing on the historical image data to obtain image normalized data;
[0019] Use the image normalized data as input data, construct an initial text feature model using the VGG network of deep learning, and train the initial text feature model according to the sample image data to obtain an image feature model;
[0020] Perform data preprocessing on the historical audio data to obtain audio normalized data;
[0021] Use the audio normalized data as input data, construct an initial audio feature model based on a convolutional neural network, and train the initial audio feature model according to the sample audio data to obtain an audio feature model;
[0022] Perform data preprocessing on the historical video data to obtain video normalized data;
[0023] Use the standardized video data as input data, construct an initial video feature model using the VGG network of deep learning, and train the initial video feature model according to the sample video data to obtain a video feature model.
[0024] In an alternative embodiment, for the text feature, the image feature, the audio feature, and the video feature, calculating a text feature vector, an image feature vector, an audio feature vector, and a video feature vector respectively includes:
[0025] Calculate the text feature vector through the following formula:
[0026]
[0027] Calculate the image feature vector through the following formula:
[0028]
[0029] Calculate the audio feature vector through the following formula:
[0030]
[0031] Calculate the video feature vector through the following formula:
[0032]
[0033] Wherein, represent the text feature vector, the image feature vector, the audio feature vector, and the video feature vector respectively; represent the resource data respectively at the layer of text feature, image feature, audio feature, video feature; represent the modality numbers of text, image, audio, and video respectively.
[0034] In an alternative embodiment, for the multi-modal semantic fusion model, calculating the attention weights between the text feature vector, the image feature vector, the audio feature vector, and the video feature vector includes:
[0035] Construct a multi-modal semantic fusion model through the following formula:
[0036]
[0037] Calculate the attention weight between the text feature vector and the image feature vector through the following formula:
[0038]
[0039] Calculate the attention weight of the video feature vector through the following formula:
[0040]
[0041] The attention weights of the audio feature vectors are calculated by the following formula:
[0042]
[0043] where respectively represent the attention weights of text data, image data, video data, and audio data; represents the activation function; represents the learnable weight matrix for bilinear transformation; respectively represent the text feature vector, image feature vector, audio feature vector, and video feature vector; represents the resource the attention weight between the corresponding text feature vector and image feature vector, represents the resource the attention weight of the corresponding video feature vector, represents the resource the attention weight of the corresponding audio feature vector; respectively represent the dimensions of the text feature vector and image feature vector; represents the dimension of the video feature vector; represents the dimension of the audio feature vector; and represent the learnable weight matrix for multi-head attention transformation.
[0044] In an alternative embodiment, after weighting and averaging the text feature vector, the image feature vector, the audio feature vector, and the video feature vector according to the attention weights, obtaining a resource representation vector includes:
[0045] The resource representation vector is calculated by the following formula:
[0046]
[0047] where represents the resource representation vector; represents the resource the attention weight between the corresponding text feature vector and image feature vector, represents the resource the attention weight of the corresponding video feature vector, represents the resource the attention weight of the corresponding audio feature vector; respectively represent the text feature vector, image feature vector, audio feature vector, and video feature vector.
[0048] In an alternative embodiment, inputting the resource representation vector into a clustering library and outputting all resources whose distance from the resource representation vector is less than a preset target distance as the educational resources retrieved by the user includes:
[0049] Using a clustering algorithm to cluster the resource representation vector and randomly selecting data points as the initial clustering centers;
[0050] Initializing the parameters of the clustering algorithm and setting the initial clustering centers as the first clustering features;
[0051] Iterating on the first clustering features to obtain second clustering features and recording the number of iterations;
[0052] Judging whether the second clustering features are superior to the first clustering features. If so, updating the second clustering features as the first clustering features; if not, discarding the second clustering features;
[0053] Continuing the next iteration. When the number of iterations is greater than or equal to a preset number of iterations, outputting the current clustering center values as the optimal clustering center values;
[0054] Calculating the distance between the resource representation vector and the optimal clustering center values as the target distance;
[0055] Outputting all resources whose distance from the resource representation vector is less than the target distance as the educational resources retrieved by the user.
[0056] In an alternative embodiment, scoring all the educational resources retrieved by the user and performing an output operation on the educational resources whose scores exceed a preconfigured optimal score threshold includes:
[0057] Scoring the educational resources retrieved by the user through the following formula:
[0058]
[0059] If , then output the educational resource;
[0060] If , then do not output the educational resource;
[0061] Where represents the score; represents the activation function; and represent the parameters of the scoring model; represents a constant; represents the optimal score threshold.
[0062] In an alternative embodiment, the configuration process of the optimal scoring threshold includes:
[0063] Obtain historical scoring data;
[0064] Preprocess the historical scoring data and split it into non - overlapping subsets;
[0065] Divide all the subsets into training set data and validation set data;
[0066] Use the training set data as input data, construct an initial threshold model using a threshold regression algorithm, train the initial threshold model according to the validation set data, and use the output of the trained initial threshold model as the optimal scoring threshold.
[0067] In a second aspect, the present invention provides an educational resource management system based on multimodal semantic analysis, including:
[0068] A data acquisition module for acquiring educational resource data and user - related data; wherein, the user - related data is the resource name clicked by the user, the knowledge points corresponding to the resource, and the label data corresponding to the resource;
[0069] A feature acquisition module for inputting the educational resource data and the user - related data into a preset text feature model, image feature model, audio feature model, and video feature model respectively to obtain text features, image features, audio features, and video features;
[0070] A feature vector calculation module for calculating text feature vectors, image feature vectors, audio feature vectors, and video feature vectors respectively according to the text features, the image features, the audio features, and the video features;
[0071] An attention weight calculation module for calculating the attention weights between the text feature vectors, the image feature vectors, the audio feature vectors, and the video feature vectors based on a multimodal semantic fusion model;
[0072] A resource representation vector calculation module for weighted - averaging the text feature vectors, the image feature vectors, the audio feature vectors, and the video feature vectors according to the attention weights to obtain a resource representation vector;
[0073] An educational resource retrieval module for inputting the resource representation vector into a clustering library and outputting all resources whose distance from the resource representation vector is less than a preset target distance as the educational resources retrieved by the user;
[0074] A scoring module, configured to score the educational resources retrieved by all the users, and output the educational resources whose scores exceed a pre-configured optimal score threshold.
[0075] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the method for managing educational resources based on multimodal semantic analysis described in any one of the above is implemented.
[0076] In a fourth aspect, the present invention further provides a computer-readable storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method for managing educational resources based on multimodal semantic analysis described in any one of the above.
[0077] Compared with the prior art, the present invention has the following beneficial effects:
[0078] The present invention discloses a method for managing educational resources based on multimodal semantic analysis, including obtaining educational resource data and user-related data; respectively inputting the educational resource data and the user-related data into a preset text feature model, image feature model, audio feature model, and video feature model to obtain text features, image features, audio features, and video features; respectively calculating text feature vectors, image feature vectors, audio feature vectors, and video feature vectors according to the text features, the image features, the audio features, and the video features; calculating the attention weights between the text feature vectors, the image feature vectors, the audio feature vectors, and the video feature vectors based on a multimodal semantic fusion model; according to the attention weights, performing weighted averaging on the text feature vectors, the image feature vectors, the audio feature vectors, and the video feature vectors to obtain a resource representation vector; inputting the resource representation vector into a clustering library, and outputting all the resources whose distance from the resource representation vector is less than a preset target distance as the educational resources retrieved by the user; scoring all the educational resources retrieved by the user, and outputting the educational resources whose scores exceed a pre-configured optimal score threshold. In summary, the present method has the following effects: it can enable the educational resource management system to perform accurate resource search and recommendation. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 is a schematic flowchart of the method for managing educational resources based on multimodal semantic analysis provided by the first embodiment of the present invention;
[0080] Figure 2It is a schematic structural diagram of an educational resource management system based on multimodal semantic analysis provided by the second embodiment of the present invention. Specific embodiments
[0081] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0082] Refer to Figure 1 , the first embodiment of the present invention provides an educational resource management method based on multimodal semantic analysis, including the following steps:
[0083] S11, obtain educational resource data and user-related data; wherein, the user-related data is the resource name clicked by the user, the knowledge points corresponding to the resource, and the tag data corresponding to the resource;
[0084] S12, input the educational resource data and the user-related data into a preset text feature model, image feature model, audio feature model, and video feature model respectively to obtain text features, image features, audio features, and video features;
[0085] S13, calculate text feature vectors, image feature vectors, audio feature vectors, and video feature vectors respectively according to the text features, the image features, the audio features, and the video features;
[0086] S14, based on the multimodal semantic fusion model, calculate the attention weights among the text feature vector, the image feature vector, the audio feature vector, and the video feature vector;
[0087] S15, according to the attention weights, after weighted averaging the text feature vector, the image feature vector, the audio feature vector, and the video feature vector, obtain a resource representation vector;
[0088] S16, input the resource representation vector into a clustering library, and output all resources whose distance from the resource representation vector is less than a preset target distance as the educational resources retrieved by the user;
[0089] S17, score all the educational resources retrieved by the user, and perform an output operation on the educational resources whose scores exceed a pre-configured optimal score threshold.
[0090] In step S11, educational resource data and user-related data are obtained; among them, the user-related data are the resource names clicked by the user, the knowledge points corresponding to the resources, and the tag data corresponding to the resources.
[0091] It should be noted that the educational resource data refer to the titles, knowledge points, tags, and text information of the titles corresponding to the educational resources, and the educational resource data can be obtained by accessing the Education Source educational database and periodicals of EBSCO and using search engines such as Google Scholar and Baidu Scholar.
[0092] The user-related data can be obtained by using Python crawlers to simulate browser behaviors, automatically scraping, analyzing, and collecting data from the Internet. HTTP requests are sent to obtain web page content by using libraries such as requests and selenium, and the Beautiful Soup library is used to parse the web pages to extract the required data, that is, the resource names clicked by the user, the knowledge points corresponding to the resources, and the tag data corresponding to the resources.
[0093] In step S12, the educational resource data and the user-related data are respectively input into a preset text feature model, image feature model, audio feature model, and video feature model to obtain text features, image features, audio features, and video features.
[0094] Historical text data, historical image data, historical audio data, and historical video data are obtained;
[0095] For the historical text data, data preprocessing is performed to obtain text standardized data;
[0096] Taking the text standardized data as input data, an initial text feature model is constructed based on the Glove algorithm, and the initial text feature model is trained according to the sample text data to obtain a text feature model;
[0097] For the historical image data, data preprocessing is performed to obtain image standardized data;
[0098] Taking the image standardized data as input data, the VGG network of deep learning is used to construct an initial text feature model, and the initial text feature model is trained according to the sample image data to obtain an image feature model;
[0099] For the historical audio data, data preprocessing is performed to obtain audio standardized data;
[0100] Taking the audio standardized data as input data, an initial audio feature model is constructed based on a convolutional neural network, and the initial audio feature model is trained according to the sample audio data to obtain an audio feature model;
[0101] Perform data preprocessing on the historical video data to obtain video standardized data;
[0102] Use the video standardized data as input data, construct an initial video feature model using the VGG network of deep learning, and train the initial video feature model according to the sample video data to obtain a video feature model.
[0103] It should be noted that the Glove algorithm is an unsupervised learning method for learning word embeddings. Its core idea is to capture the semantic relationships between words by modeling the co-occurrence frequencies of words. The Glove algorithm combines methods based on global statistical information and methods based on local windows, and directly uses the statistical information in the entire corpus to train word vectors, rather than by predicting the surrounding words. The Glove algorithm can count the co-occurrence times of all word pairs in the corpus to form a word co-occurrence matrix. The co-occurrence times reflect the symbiotic relationship between words and are an important basis for constructing word vectors.
[0104] The significance of the Glove algorithm is that it can capture the semantic relationships between words and obtain word vectors through global statistical information, which makes it more effective when dealing with large-scale datasets. Compared with Word2Vec, Glove not only relies on local statistical information, but combines global statistical information to obtain word vectors, which helps to better capture the semantic relationships between words.
[0105] The VGG network is a classic convolutional neural network architecture. The VGG network is very deep, and the basic models have two types: 16 layers (VGG16) and 19 layers (VGG19). This deep structure enables the network to learn more features and improve the model's representation ability for images.
[0106] In step S13, according to the text features, the image features, the audio features, and the video features, calculate text feature vectors, image feature vectors, audio feature vectors, and video feature vectors respectively.
[0107] Calculate the text feature vector through the following formula:
[0108]
[0109] Calculate the image feature vector through the following formula:
[0110]
[0111] Calculate the audio feature vector through the following formula:
[0112]
[0113] The video feature vector is calculated by the following formula:
[0114]
[0115] Wherein, respectively represent the text feature vector, the image feature vector, the audio feature vector, and the video feature vector; respectively represent the resource data at the layer's text features, image features, audio features, video features; respectively represent the modality numbers of text, image, audio, and video.
[0116] It should be noted that the modality number refers to the number of the inherent vibration characteristics of the structural system. Each order of modality contains a modal frequency, a modal shape, a modal damping, a modal stiffness, and a modal mass. The acquisition of the modality number can be achieved through modal analysis. Modal analysis is a method for studying the dynamic characteristics of structures, which includes analytical (theoretical) modal analysis, experimental modal analysis, and operational modal analysis. The modality number can also be obtained through computational methods such as finite element analysis, by solving the eigenvectors or eigenvalues corresponding to the vibration characteristic equation of the structure, so as to obtain the modal parameters.
[0117] In step S14, based on the multi-modal semantic fusion model, calculate the attention weights among the text feature vector, the image feature vector, the audio feature vector, and the video feature vector.
[0118] The multi-modal semantic fusion model is constructed by the following formula:
[0119]
[0120] The attention weight between the text feature vector and the image feature vector is calculated by the following formula:
[0121]
[0122] The attention weight of the video feature vector is calculated by the following formula:
[0123]
[0124] The attention weight of the audio feature vector is calculated by the following formula:
[0125]
[0126] Wherein, respectively represent the attention weights of the text data, the image data, the video data, and the audio data; represents the activation function; denotes a learnable weight matrix for bilinear transformation; respectively denote the text feature vector, the image feature vector, the audio feature vector, and the video feature vector; denotes the resource the attention weight between the corresponding text feature vector and the image feature vector, denotes the resource the attention weight of the corresponding video feature vector, denotes the resource the attention weight of the corresponding audio feature vector; respectively denote the dimensions of the text feature vector and the image feature vector; denotes the dimension of the video feature vector; denotes the dimension of the audio feature vector; and denotes a learnable weight matrix for multi-head attention transformation.
[0127] It should be noted that the multi-modal semantic fusion model is a model that integrates information from different modalities, such as text, image, audio, video, etc., to obtain a richer and more accurate semantic understanding. This model can make full use of the complementarity between different modalities and integrate the information extracted from different modalities into a stable multi-modal representation. The multi-modal semantic fusion model can improve the decision-making accuracy. It combines data from different modalities, thus providing more comprehensive information, enabling the multi-modal semantic fusion model to make more accurate decisions.
[0128] The attention weight is the core concept in the attention mechanism. It represents the degree of importance that the model attaches to different input information when processing a specific task. The attention weight is obtained by calculating the similarity or correlation between a query ( ) and a set of keys ( ). It determines the weight coefficients of each value when aggregating information, and these weight coefficients reflect the importance of different input information for the current task. The attention weight can be obtained by an attention score function, such as the dot product, scaled dot product, etc. and the similarity or correlation; The attention weight can also be obtained by applying a function for normalization: the calculated scores are processed by the function to obtain a normalized weight distribution, making the sum of all weights equal to 1, forming a valid probability distribution.
[0129] The calculated attention weights enable the model to focus on the information most important for the current task and ignore the less important information. The attention weight mechanism also allows the model to dynamically focus on different parts when processing sequential data, thereby improving the model's performance and accuracy.
[0130] The bilinear transformation is a special conformal mapping, which can also be regarded as a Möbius transformation. It can convert the transfer function of a linear time-invariant system filter in the continuous time domain into the transfer function of a linear and shift-invariant filter in the discrete time domain. This transformation is achieved by mapping the points on the axis in the plane to the unit circle in the complex plane, thus realizing the conversion between continuous-time and discrete-time systems. The bilinear transformation can preserve the stability of the continuous-time filter in the discrete-time filter, that is, if the continuous-time filter is stable, then the discrete-time filter obtained by the bilinear transformation is also stable.
[0131] The bilinear transformation can preserve the stability of the continuous-time filter in the discrete-time filter, that is, if the continuous-time filter is stable, then the discrete-time filter obtained by the bilinear transformation is also stable.
[0132] The multi-head attention transformation extends the traditional attention mechanism. Its core idea is to learn the correlations of different subspaces of the data in parallel through multiple different "heads", thereby improving the model's expressive ability. To obtain the multi-head attention transformation, the data can be first input and then mapped to multiple subspaces through different weight matrices. Each subspace calculates independent attention scores, and then the scaled dot-product attention operation is performed.
[0133] The multi-head attention projects the data through multiple heads, enabling the model to focus on different aspects of the data. In a specific application scenario, for example, in a sentence, one head focuses on the relationship between the subject and the predicate, and another head focuses on the temporal consistency of the context.
[0134] In step S15, according to the attention weights, the text feature vector, the image feature vector, the audio feature vector, and the video feature vector are weighted and averaged to obtain a resource representation vector.
[0135] The resource representation vector is calculated by the following formula:
[0136]
[0137] where, represents the resource representation vector; represents the resource The attention weights between the corresponding text feature vector and image feature vector represent resources The attention weights of the corresponding video feature vector represent resources The attention weights of the corresponding audio feature vector; respectively represent the text feature vector, image feature vector, audio feature vector and video feature vector.
[0138] It should be noted that the resource representation vector is a vector used in the fields of machine learning and information retrieval. It is used to represent the semantic features of text, images, audio, video, or other types of resources. These vectors are extracted by deep learning models, which can capture the meaning of the resources and represent the semantic relationships between the resources in numerical form. The resource representation vector can convert abstract semantic features into numerical forms that can be processed by computers. Or in a high-dimensional space, resources with similar semantics will be mapped to adjacent regions, enabling the computer to recognize and process the similarities between them.
[0139] In step S16, input the resource representation vector into the clustering library, and output all resources whose distance from the resource representation vector is less than the preset target distance as the educational resources retrieved by the user.
[0140] Use a clustering algorithm to cluster the resource representation vector, and randomly select a data point as the initial clustering center;
[0141] Initialize the parameters of the clustering algorithm, and set the initial clustering center as the first clustering feature;
[0142] Iterate on the first clustering feature to obtain the second clustering feature, and record the number of iterations;
[0143] Judge whether the second clustering feature is better than the first clustering feature. If so, update the second clustering feature to the first clustering feature; if not, discard the second clustering feature;
[0144] Continue the next iteration. When the number of iterations is greater than or equal to the preset number of iterations, output the current clustering center value as the optimal clustering center value;
[0145] Calculate the distance between the resource representation vector and the optimal clustering center value as the target distance;
[0146] Output all resources whose distance from the resource representation vector is less than the target distance as the educational resources retrieved by the user.
[0147] It should be noted that the clustering algorithm is an unsupervised learning method used to divide a set of data points into several clusters, such that the data points within the same cluster are similar to each other, while the data points in different clusters are quite different. The clustering algorithm can be obtained through the K-means algorithm. Randomly select k central points, assign each data point to the nearest central point, and then calculate the average value of each cluster as the new central point. Repeat this process until the central points no longer change.
[0148] The clustering algorithm can divide the objects in the dataset into multiple groups or clusters according to similarity, realizing the organization and classification of data; and reducing the dimension of data through clustering to simplify the problem.
[0149] The resource representation vector is a vector used in the fields of machine learning and information retrieval. It is used to represent the semantic features of text, images, audio, video, or other types of resources. These vectors are extracted by deep learning models, which can capture the meaning of the resources and represent the semantic relationships between resources in numerical form. The resource representation vector can convert abstract semantic features into numerical forms that can be processed by computers. Or in a high-dimensional space, semantically similar resources will be mapped to adjacent regions, enabling the computer to recognize and process the similarities between them.
[0150] In step S17, score all the educational resources retrieved by the user, and perform an output operation on the educational resources whose scores exceed the pre-configured optimal score threshold. Score the educational resources retrieved by the user through the following formula:
[0151]
[0152] If , then output the educational resource;
[0153] If , then do not output the educational resource;
[0154] Among them, represents the score; represents the activation function; and represent the parameters of the scoring model; represents a constant; represents the optimal score threshold.
[0155] It should be noted that the activation function is a key component in artificial neural networks, which is used to introduce non - linear factors into each neuron of the neural network. Without an activation function, no matter how many layers the neural network has, it can actually only perform linear transformations and cannot solve non - linear problems. The activation function enables the neural network to learn and simulate non - linear relationships and complex functions, and determines whether a neuron should be activated, thus enhancing the selectivity of the network.
[0156] Furthermore, the configuration process of the optimal scoring threshold includes:
[0157] Obtain historical scoring data;
[0158] Pre - process the historical scoring data and split it into non - overlapping subsets;
[0159] Divide all the subsets into training set data and validation set data;
[0160] Use the training set data as input data, construct an initial threshold model using the threshold regression algorithm, train the initial threshold model according to the validation set data, and take the output of the trained initial threshold model as the optimal scoring threshold.
[0161] It should be noted that the training set data and the validation set data are two parts of the dataset used to evaluate the model performance: the training set is the dataset used to train the model, that is, the model learns and adjusts parameters through these data to minimize the prediction error; the validation set is the dataset used to verify the generalization ability of the model, that is, to evaluate the performance of the model on unseen data to prevent overfitting of the model. The training set is used to train the model to help the model learn the mapping relationship from input to output; the validation set is used to evaluate the generalization ability of the model, that is, the prediction ability of the model for new data.
[0162] The following takes a relatively common scenario as an example to describe the working process of the present invention. Please also refer to Figure 1 the schematic diagram of the working scenario of the method of
[0163] In a language learning center of a university, the staff uses an intelligent recommendation system to help students find educational resources suitable for them by obtaining educational resource data and student user-related data; respectively inputting the educational resource data and the student user-related data into a preset text feature model, image feature model, audio feature model, and video feature model to obtain text features, image features, audio features, and video features; respectively calculating text feature vectors, image feature vectors, audio feature vectors, and video feature vectors according to the text features, the image features, the audio features, and the video features; calculating attention weights among the text feature vector, the image feature vector, the audio feature vector, and the video feature vector based on a multimodal semantic fusion model; after weighted averaging the text feature vector, the image feature vector, the audio feature vector, and the video feature vector according to the attention weights, obtaining a resource representation vector; inputting the resource representation vector into a clustering library, and outputting all resources with a distance less than a target distance from the resource representation vector as the educational resources retrieved by the student user; scoring all the educational resources retrieved by the student user, and outputting the educational resources with a score greater than an optimal score threshold, so as to help students find educational resources suitable for them.
[0164] In summary, the present invention discloses an educational resource management method based on multimodal semantic analysis, including obtaining educational resource data and user-related data; respectively inputting the educational resource data and the user-related data into a preset text feature model, image feature model, audio feature model, and video feature model to obtain text features, image features, audio features, and video features; respectively calculating text feature vectors, image feature vectors, audio feature vectors, and video feature vectors according to the text features, the image features, the audio features, and the video features; calculating attention weights among the text feature vector, the image feature vector, the audio feature vector, and the video feature vector based on a multimodal semantic fusion model; after weighted averaging the text feature vector, the image feature vector, the audio feature vector, and the video feature vector according to the attention weights, obtaining a resource representation vector; inputting the resource representation vector into a clustering library, and outputting all resources with a distance less than a preset target distance from the resource representation vector as the educational resources retrieved by the user; scoring all the educational resources retrieved by the user, and performing an output operation on the educational resources with a score exceeding a pre-configured optimal score threshold. In summary, the present method has the following effects: it can enable the educational resource management system to perform accurate resource search and recommendation.
[0165] Referring to Figure 2 , the second embodiment of the present invention provides an educational resource management system based on multimodal semantic analysis, including:
[0166] A data acquisition module, configured to acquire educational resource data and user-related data; wherein, the user-related data is the resource name clicked by the user, the knowledge points corresponding to the resource, and the label data corresponding to the resource;
[0167] A feature acquisition module, configured to input the educational resource data and the user-related data into a preset text feature model, image feature model, audio feature model, and video feature model respectively, to obtain text features, image features, audio features, and video features;
[0168] A feature vector calculation module, configured to calculate a text feature vector, an image feature vector, an audio feature vector, and a video feature vector respectively according to the text features, the image features, the audio features, and the video features;
[0169] An attention weight calculation module, configured to calculate the attention weights among the text feature vector, the image feature vector, the audio feature vector, and the video feature vector based on a multi-modal semantic fusion model;
[0170] A resource representation vector calculation module, configured to perform weighted averaging on the text feature vector, the image feature vector, the audio feature vector, and the video feature vector according to the attention weights to obtain a resource representation vector;
[0171] An educational resource retrieval module, configured to input the resource representation vector into a clustering library, and output all resources whose distance from the resource representation vector is less than a preset target distance as the educational resources retrieved by the user;
[0172] A scoring module, configured to score all the educational resources retrieved by the user, and perform an output operation on the educational resources whose scores exceed a pre-configured optimal score threshold.
[0173] It should be noted that an educational resource management system based on multi-modal semantic analysis provided in an embodiment of the present invention is used to execute all the process steps of an educational resource management method based on multi-modal semantic analysis in the above embodiment, and the working principles and beneficial effects of the two correspond one by one, so they will not be elaborated here.
[0174] An embodiment of the present invention further provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a data processing terminal program. When the processor executes the computer program, the steps in the above-mentioned method embodiments of the educational resource management method based on multi-modal semantic analysis are implemented, such as Figure 1The step S11 shown. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in the above-described device embodiments, such as the scoring module.
[0175] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device.
[0176] The electronic device may be a computing device such as a desktop computer, a notebook, a palm computer, and a smart tablet. The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above components are only examples of the electronic device and do not constitute a limitation on the electronic device. It may include more or fewer components than the above, or combine some components, or different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.
[0177] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device and connects various parts of the entire electronic device using various interfaces and lines.
[0178] The memory can be used to store the computer program and / or modules. By running or executing the computer program and / or modules stored in the memory, and invoking the data stored in the memory, the processor can implement various functions of the electronic device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0179] Among them, if the modules / units integrated in the electronic device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0180] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0181] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An educational resource management method based on multimodal semantic analysis, characterized in that: Executed by a computer, including: Acquire educational resource data and user-related data; wherein the user-related data is the resource name clicked by the user, the knowledge point corresponding to the resource, and the tag data corresponding to the resource; According to the educational resource data and the user-related data, the educational resource data and the user-related data are respectively input into a preset text feature model, an image feature model, an audio feature model and a video feature model to obtain text features, image features, audio features and video features; According to the text features, the image features, the audio features and the video features, respectively calculating a text feature vector, an image feature vector, an audio feature vector and a video feature vector; Based on a multimodal semantic fusion model, calculating the attention weights among the text feature vector, the image feature vector, the audio feature vector, and the video feature vector; According to the attention weight, weighted average the text feature vector, the image feature vector, the audio feature vector, and the video feature vector is obtained to obtain a resource representation vector; Inputting the resource representation vector into a clustering library, and outputting all resources whose distances to the resource representation vector are less than a preset target distance as educational resources retrieved by the user; Scoring all educational resources retrieved by the user, and outputting educational resources whose scores exceed a pre-configured optimal scoring threshold; The educational resource management method based on multimodal semantic analysis is characterized in that the attention weights among the text feature vector, the image feature vector, the audio feature vector, and the video feature vector are calculated based on the multimodal semantic fusion model, including: The multimodal semantic fusion model is constructed by the following formula: The attention weight between the text feature vector and the image feature vector is calculated by the following formula: The attention weight of the video feature vector is calculated by the following formula: The attention weight of the audio feature vector is calculated by the following formula: in, Respectively represent the attention weight between text data and image data, the attention weight of video data, and the attention weight of audio data; represents the activation function; Represents a learnable weight matrix for bilinear transformation; Represent text feature vector, image feature vector, audio feature vector and video feature vector respectively; Representation Resources The attention weights between the corresponding text feature vector and image feature vector, Representation Resources The attention weight of the corresponding video feature vector, Representation Resources The attention weight of the corresponding audio feature vector; Represent the dimensions of text feature vector and image feature vector respectively; Represents the dimension of the video feature vector; Represents the dimension of the audio feature vector; and represents a learnable weight matrix for multi-head attention transformation; The educational resource management method based on multimodal semantic analysis is characterized in that, according to the attention weight, the text feature vector, the image feature vector, the audio feature vector and the video feature vector are weighted averaged to obtain a resource representation vector, including: The resource representation vector is calculated using the following formula: in, Represents a resource representation vector; Representation Resources The attention weights between the corresponding text feature vector and image feature vector, Representation Resources The attention weight of the corresponding video feature vector, Representation Resources The attention weight of the corresponding audio feature vector; Represent text feature vector, image feature vector, audio feature vector and video feature vector respectively.
2. The educational resource management method based on multimodal semantic analysis according to claim 1 is characterized in that: The training process of the text feature model, the image feature model, the audio feature model and the video feature model includes: Acquire historical text data, historical image data, historical audio data, and historical video data; Performing data preprocessing on the historical text data to obtain text standardization data; Taking the text normalization data as input data, constructing an initial text feature model based on the Glove algorithm, and training the initial text feature model according to the sample text data to obtain a text feature model; Performing data preprocessing on the historical image data to obtain image standardized data; Using the image normalization data as input data, using a deep learning VGG network to build an initial text feature model, and training the initial text feature model according to sample image data to obtain an image feature model; Performing data preprocessing on the historical audio data to obtain audio standardization data; Using the audio normalization data as input data, constructing an initial audio feature model based on a convolutional neural network, and training the initial audio feature model according to the sample audio data to obtain an audio feature model; Performing data preprocessing on the historical video data to obtain video standardized data; The video normalization data is used as input data, an initial video feature model is constructed using a deep learning VGG network, and the initial video feature model is trained according to sample video data to obtain a video feature model.
3. The educational resource management method based on multimodal semantic analysis according to claim 2 is characterized in that: According to the text features, the image features, the audio features and the video features, a text feature vector, an image feature vector, an audio feature vector and a video feature vector are calculated respectively, including: The text feature vector is calculated using the following formula: The image feature vector is calculated by the following formula: The audio feature vector is calculated by the following formula: The video feature vector is calculated using the following formula: in, Represent text feature vector, image feature vector, audio feature vector and video feature vector respectively; Respectively represent resource data In the Text features, image features, audio features, and video features of layer modalities; Represents the number of modalities for text, image, audio, and video respectively.
4. The educational resource management method based on multimodal semantic analysis according to claim 1 is characterized in that: The step of inputting the resource representation vector into a clustering library and outputting all resources whose distances from the resource representation vector are less than a preset target distance as educational resources retrieved by the user includes: Clustering the resource representation vectors using a clustering algorithm, and randomly selecting data points as initial cluster centers; Initialize the parameters of the clustering algorithm and set the initial cluster center as the first clustering feature; Iterate the first clustering feature to obtain a second clustering feature, and record the number of iterations; Determine whether the second clustering feature is better than the first clustering feature, if so, update the second clustering feature to the first clustering feature; if not, discard the second clustering feature; Continue to the next iteration, when the number of iterations is greater than or equal to the preset number of iterations, output the current cluster center value as the optimal cluster center value; Calculating the distance between the resource representation vector and the optimal cluster center value as the target distance; All resources whose distance to the resource representation vector is less than the target distance are output as educational resources retrieved by the user.
5. The educational resource management method based on multimodal semantic analysis according to claim 4 is characterized in that: The step of scoring all educational resources retrieved by the user and outputting educational resources with scores exceeding a pre-configured optimal score threshold comprises: The educational resources retrieved by the user are scored using the following formula: like , then output the educational resources; like , then the educational resources are not output; in, Indicates the rating; represents the activation function; represents the parameters of the scoring model; Indicates the preset weight of the corresponding educational resources; represents a constant; Represents the optimal scoring threshold.
6. The educational resource management method based on multimodal semantic analysis according to claim 5 is characterized in that: The configuration process of the optimal scoring threshold includes: Get historical rating data; Preprocessing the historical rating data to obtain non-overlapping subsets; Dividing all the subsets into training set data and validation set data; The training set data is used as input data, an initial threshold model is constructed using a threshold regression algorithm, the initial threshold model is trained according to the validation set data, and the output of the trained initial threshold model is used as the optimal scoring threshold.
7. An educational resource management system based on multimodal semantic analysis, characterized in that: The method for managing educational resources based on multimodal semantic analysis as claimed in any one of claims 1 to 6 comprises: The data acquisition module is used to acquire educational resource data and user-related data; wherein the user-related data is the resource name clicked by the user, the knowledge point corresponding to the resource, and the tag data corresponding to the resource; A feature acquisition module, used to input the educational resource data and the user-related data into a preset text feature model, an image feature model, an audio feature model, and a video feature model, respectively, to obtain text features, image features, audio features, and video features; A feature vector calculation module, used to calculate a text feature vector, an image feature vector, an audio feature vector and a video feature vector according to the text feature, the image feature, the audio feature and the video feature; An attention weight calculation module, used to calculate the attention weights among the text feature vector, the image feature vector, the audio feature vector, and the video feature vector based on a multimodal semantic fusion model; A resource representation vector calculation module, configured to obtain a resource representation vector by weighted averaging the text feature vector, the image feature vector, the audio feature vector, and the video feature vector according to the attention weight; An educational resource retrieval module, used for inputting the resource representation vector into a clustering library, and outputting all resources whose distance to the resource representation vector is less than a preset target distance as educational resources retrieved by the user; The scoring module is used to score all educational resources retrieved by the user and output the educational resources whose scores exceed a pre-configured optimal scoring threshold.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the educational resource management method based on multimodal semantic analysis as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Search method, device and system and storage medium
CN110858232A
Man-machine collaborative health case matching method based on chronic disease big data and system thereof
CN113345587A
Video cover extraction method and device, equipment and computer readable storage medium
CN113762052A
Test question recommendation method and device, electronic equipment and storage medium
CN113837910A
Emotion analysis method based on multi-modal cross attention mechanism image-text fusion
CN116844179A