A multi-modal education resource automatic labeling method based on time domain causal information
By constructing a multi-level descriptive framework and processing temporal causal information, and combining it with deep neural networks for multimodal educational resource annotation, the problem of low accuracy in multimodal educational resource annotation is solved, achieving higher annotation accuracy and adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies have low accuracy and poor fault tolerance in the annotation of multimodal educational resources, making it difficult to effectively handle the causal essential feature learning of multi-source heterogeneous multimodal educational resources, thus limiting the performance of automatic annotation.
A descriptive framework for multi-level educational resource usage records is constructed. Multimodal educational resource usage records are collected and intervened through temporal causal information. Deep neural networks are used for data fusion representation. An online-offline incremental learning annotation model is constructed to output educational resource annotation results.
It improves the accuracy and robustness of automatic annotation of multimodal educational resources, enabling it to better characterize the causal essential features of educational resources and adapt to diverse educational resource usage scenarios.
Smart Images

Figure CN115170360B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to an automatic annotation method for multimodal educational resources based on temporal causal information. Background Technology
[0002] Educational resource annotation is fundamental to important tasks such as resource organization and recommendation. Currently, most methods are based on manual rules or simple learning algorithms, with few designed based on actual resource usage records. This results in low fault tolerance, limited data modality processing, narrow application scenarios, and low accuracy. This is particularly challenging in the internet age with the explosive growth of multimodal educational resources, posing a significant challenge to efficient automatic annotation. Deep neural network technology has made considerable progress in multimodal data analysis and processing. Compared to traditional manual rules or shallow learning algorithms, these methods offer advantages such as better general modeling and fitting of complex mechanisms. However, deep learning technology is driven by correlations, which limits its ability to learn the causal essence of complex, multi-source, heterogeneous, and multimodal educational resources, severely restricting the performance of automatic annotation of massive amounts of multimodal educational resources. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a highly accurate and robust method for automatic annotation of multimodal educational resources based on temporal causal information. This method starts from the perspective of the educational resource usage process and utilizes both data from the educational resource usage process and data from the multimodal educational resources themselves to enhance the causal nature representation learning of multimodal educational resources.
[0004] One aspect of this invention provides an automatic annotation method for multimodal educational resources based on temporal causal information, comprising:
[0005] Construct a descriptive framework for multi-level educational resource usage records;
[0006] Multimodal educational resource usage records are collected based on the described framework, and temporal causal intervention is performed on the collected educational resource usage records.
[0007] Data on the multimodal use of educational resources is integrated and represented.
[0008] Based on the fusion representation of the multimodal educational resources, an online-offline incremental learning annotation model is constructed;
[0009] Input the multimodal educational resources to be labeled into the online-offline incremental learning labeling model, and output the labeling results of the educational resources.
[0010] Optionally, the construction of the multi-level educational resource usage record description framework specifically involves: establishing a three-level semantic description mechanism for resource usage based on the characteristics of online education scenarios, namely "behavior-goal-state". This step includes:
[0011] Constructing a behavioral description layer: Based on the characteristics of resource usage in learning scenarios, and using the temporal predicate logic expression method, a semantic model of interaction relationships between "user-resource-action" is constructed. The semantic model of interaction relationships is used to describe the interaction relationships between users and resources in learning scenarios, and to establish a vocabulary and instance relationship set for each element. The elements include, but are not limited to, behavior, goal, and educational resource usage status.
[0012] Constructing a target description layer: Based on the BDI model and incorporating knowledge from the learning resource domain, construct an action-target-state operation semantic representation with action antecedent, action consequent, and usage state. Introduce the target and resource usage state into the operation semantics of user actions to achieve the association between the three levels of action, target, and resource usage state.
[0013] Construct a layer describing resource usage status: extract a set of resource usage performance evaluation metrics directly from log records or indirectly using pre-trained machine learning models.
[0014] Optionally, the step of collecting multimodal educational resource usage records according to the description framework includes:
[0015] Based on online streaming log data and multimedia sensor data, and with the help of multi-type deep neural network learning algorithms, combined with a three-level resource usage semantic description mechanism, we construct a method for user and resource identification and user-resource-action interaction relationship extraction in the learning scenario, complete the semantic description of online cross-time domain original resource usage process data, and realize the semantic perception and organization of multimodal time-series process data.
[0016] The semantic process data is represented as {(d raw ,L t ,t link )},d raw L represents the raw collected data. t This indicates that the semantic tag set is obtained from the three-layer semantic description mechanism, t link This field represents a record of time information.
[0017] The online streaming log data and multimedia sensor data include, but are not limited to, online structured learning log data, audio data, text data, and video data.
[0018] Optionally, the temporal causal intervention on the collected educational resource usage records includes:
[0019] To address the temporal causal characteristics of multimodal educational resource use in the learning process, a deconfusion process model is developed from the perspective of temporal distributed causal relationships.
[0020] Based on the perspective of causal relationships of time-series events in the learning process, and with the help of a set of semantically perceptible labels for the learning process data, an intervention mechanism for pre-training constrained by time-series causal relationships is constructed, and the process representation is deconfused in the continuous time domain by fitting equations from the perspective of discrete events.
[0021] We construct a multimodal x-Transformer-based intervention learning mechanism using discrete temporal Mask operators, where x in x-Transformer represents different modal data.
[0022] Optionally, the intervention mechanism for pre-training temporal causal relationship constraints, based on the causal relationship of time-series events in the learning process and utilizing the semantically perceptible tag set of the learning process data, and fitting the equation in the continuous time domain from the perspective of discrete events to deconfound the process representation, includes:
[0023] Based on the set of semantically perceptible labels of learned data, we construct the event-event, event-time relationship sets, time series label sets, and causal relationship label sets for the learning behavior process.
[0024] In the temporal dimension of learning process data, a method for calculating temporal label constraints of learning behavior events is constructed.
[0025] A complete temporal causal relationship constraint method is constructed to solve the regularization constraint for the selection calculation of Mask position for learning process data.
[0026] Optionally, the fusion representation of data on the multimodal use of educational resources includes:
[0027] Based on semantically organized learning process data, the flexible x-Transformer is used as the underlying general backbone network. Combined with the characteristics of different data types, the x-Transformer architecture is optimized and adjusted using domain-general optimization methods; where x represents different modal data.
[0028] Feature associations are extracted from data of different modalities, and an association degree calculation function is constructed for cross-modal attention fusion enhancement.
[0029] To construct a cross-modal temporal attention fusion model, the data is first combined at preset intervals to collect data at each temporally aligned moment. Then, the data is input into a recurrent network consisting of stacked cross-modal attention fusion modules. The data at the same moment are fused first, and then feature diffusion is performed in the temporal dimension to obtain modality-invariant semantic representations.
[0030] The data on the use of multimodal and cross-temporal educational resources are fused to obtain a set of features for specific educational resources.
[0031] Optionally, constructing an online-offline incremental learning annotation model based on the fused representation of the multimodal educational resources includes:
[0032] For the semantic incremental class of tags or words, semantic ontology technology is first used to perform semantic analysis on the sampled text words. For newly added text words, multimodal information of the newly added text words is collected, and multimodal features of the newly added text words are obtained. The multimodal features and the corresponding text words are used to construct double tuple features. Based on the double tuple features, a corresponding binary classification learning task is constructed to identify newly added text words and non-newly added text words, thus completing the first type of incremental learning.
[0033] To identify salient representation features of the resource usage process, the semantic feature representations of the resource usage process data are first obtained. These obtained semantic feature representations are then compared with existing semantic representations related to the resource usage process corresponding to the educational resource. The salientity is used to determine whether a new salient representation feature is added. For newly identified salient features, they are added to the corresponding educational resource usage process feature library, and the relevant text semantics are retained to prepare for the next incremental learning, thus completing the second type of incremental learning.
[0034] Optionally, the step of inputting the multimodal educational resources to be labeled into the online-offline incremental learning labeling model and outputting the labeling results of the educational resources includes:
[0035] Construct a feature representation of the educational resources that need to be labeled;
[0036] Obtain the deep semantic representation of tags in the tag library;
[0037] Construct an automatic labeling model for educational resources;
[0038] The automatic labeling model for educational resources outputs multimodal labeling results.
[0039] The educational resource labels are updated based on the similarity between sample data and label data in the automatic labeling model for educational resources, thus completing the update and optimization of the automatic labeling model for educational resources.
[0040] Another aspect of the present invention provides an electronic device, including a processor and a memory;
[0041] The memory is used to store programs;
[0042] The processor executes the program to implement the method described above.
[0043] Another aspect of this invention provides a computer-readable storage medium storing a program that is executed by a processor to implement the methods described above.
[0044] This invention also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned method.
[0045] Embodiments of this invention construct a descriptive framework for multi-level educational resource usage records; collect multimodal educational resource usage records according to the descriptive framework, and perform temporal causal intervention on the collected educational resource usage records; perform fusion representation on the multimodal usage process data of educational resources; construct an online-offline incremental learning annotation model based on the fusion representation of the multimodal educational resources; input the multimodal educational resources to be annotated into the online-offline incremental learning annotation model, and output the educational resource annotation results. Embodiments of this invention exhibit high accuracy and robustness. This invention can enhance the causal essence representation learning of multimodal educational resources by starting from the perspective of the educational resource usage process and simultaneously utilizing educational resource usage process data and multimodal educational resource data itself. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 The overall process flowchart provided for embodiments of the present invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0049] To address the problems existing in the prior art, one aspect of this invention provides a method for automatic annotation of multimodal educational resources based on temporal causal information, comprising:
[0050] Construct a descriptive framework for multi-level educational resource usage records;
[0051] Multimodal educational resource usage records are collected based on the described framework, and temporal causal intervention is performed on the collected educational resource usage records.
[0052] Data on the multimodal use of educational resources is integrated and represented.
[0053] Based on the fusion representation of the multimodal educational resources, an online-offline incremental learning annotation model is constructed;
[0054] Input the multimodal educational resources to be labeled into the online-offline incremental learning labeling model, and output the labeling results of the educational resources.
[0055] Optionally, the construction of the multi-level educational resource usage record description framework specifically involves: establishing a three-level semantic description mechanism for resource usage based on the characteristics of online education scenarios, namely "behavior-goal-state". This step includes:
[0056] Constructing a behavioral description layer: Based on the characteristics of resource usage in learning scenarios, and using the temporal predicate logic expression method, a semantic model of interaction relationships between "user-resource-action" is constructed. The semantic model of interaction relationships is used to describe the interaction relationships between users and resources in learning scenarios, and to establish a vocabulary and instance relationship set for each element. The elements include, but are not limited to, behavior, goal, and educational resource usage status.
[0057] Constructing a target description layer: Based on the BDI model and incorporating knowledge from the learning resource domain, construct an action-target-state operation semantic representation with action antecedent, action consequent, and usage state. Introduce the target and resource usage state into the operation semantics of user actions to achieve the association between the three levels of action, target, and resource usage state.
[0058] Construct a layer describing resource usage status: extract a set of resource usage performance evaluation metrics directly from log records or indirectly using pre-trained machine learning models.
[0059] Optionally, the step of collecting multimodal educational resource usage records according to the description framework includes:
[0060] Based on online streaming log data and multimedia sensor data, and with the help of multi-type deep neural network learning algorithms, combined with a three-level resource usage semantic description mechanism, we construct a method for user and resource identification and user-resource-action interaction relationship extraction in the learning scenario, complete the semantic description of online cross-time domain original resource usage process data, and realize the semantic perception and organization of multimodal time-series process data.
[0061] The semantic process data is represented as {(d raw ,L t ,t link )},d rawL represents the raw collected data. t This indicates that the semantic tag set is obtained from the three-layer semantic description mechanism, t link This field represents a record of time information.
[0062] The online streaming log data and multimedia sensor data include, but are not limited to, online structured learning log data, audio data, text data, and video data.
[0063] Optionally, the temporal causal intervention on the collected educational resource usage records includes:
[0064] To address the temporal causal characteristics of multimodal educational resource use in the learning process, a deconfusion process model is developed from the perspective of temporal distributed causal relationships.
[0065] Based on the perspective of causal relationships of time-series events in the learning process, and with the help of a set of semantically perceptible labels for the learning process data, an intervention mechanism for pre-training constrained by time-series causal relationships is constructed, and the process representation is deconfused in the continuous time domain by fitting equations from the perspective of discrete events.
[0066] We construct a multimodal x-Transformer-based intervention learning mechanism using discrete temporal Mask operators, where x in x-Transformer represents different modal data.
[0067] Optionally, the intervention mechanism for pre-training temporal causal relationship constraints, based on the causal relationship of time-series events in the learning process and utilizing the semantically perceptible tag set of the learning process data, and fitting the equation in the continuous time domain from the perspective of discrete events to deconfound the process representation, includes:
[0068] Based on the set of semantically perceptible labels of learned data, we construct the event-event, event-time relationship sets, time series label sets, and causal relationship label sets for the learning behavior process.
[0069] In the temporal dimension of learning process data, a method for calculating temporal label constraints of learning behavior events is constructed.
[0070] A complete temporal causal relationship constraint method is constructed to solve the regularization constraint for the selection calculation of Mask position for learning process data.
[0071] Optionally, the fusion representation of data on the multimodal use of educational resources includes:
[0072] Based on semantically organized learning process data, the flexible x-Transformer is used as the underlying general backbone network. Combined with the characteristics of different data types, the x-Transformer architecture is optimized and adjusted using domain-general optimization methods; where x represents different modal data.
[0073] Feature associations are extracted from data of different modalities, and an association degree calculation function is constructed for cross-modal attention fusion enhancement.
[0074] To construct a cross-modal temporal attention fusion model, the data is first combined at preset intervals to collect data at each temporally aligned moment. Then, the data is input into a recurrent network consisting of stacked cross-modal attention fusion modules. The data at the same moment are fused first, and then feature diffusion is performed in the temporal dimension to obtain modality-invariant semantic representations.
[0075] The data on the use of multimodal and cross-temporal educational resources are fused to obtain a set of features for specific educational resources.
[0076] Optionally, constructing an online-offline incremental learning annotation model based on the fused representation of the multimodal educational resources includes:
[0077] For the semantic incremental class of tags or words, semantic ontology technology is first used to perform semantic analysis on the sampled text words. For newly added text words, multimodal information of the newly added text words is collected, and multimodal features of the newly added text words are obtained. The multimodal features and the corresponding text words are used to construct double tuple features. Based on the double tuple features, a corresponding binary classification learning task is constructed to identify newly added text words and non-newly added text words, thus completing the first type of incremental learning.
[0078] To identify salient representation features of the resource usage process, the semantic feature representations of the resource usage process data are first obtained. These obtained semantic feature representations are then compared with existing semantic representations related to the resource usage process corresponding to the educational resource. The salientity is used to determine whether a new salient representation feature is added. For newly identified salient features, they are added to the corresponding educational resource usage process feature library, and the relevant text semantics are retained to prepare for the next incremental learning, thus completing the second type of incremental learning.
[0079] Optionally, the step of inputting the multimodal educational resources to be labeled into the online-offline incremental learning labeling model and outputting the labeling results of the educational resources includes:
[0080] Construct a feature representation of the educational resources that need to be labeled;
[0081] Obtain the deep semantic representation of tags in the tag library;
[0082] Construct an automatic labeling model for educational resources;
[0083] The automatic labeling model for educational resources outputs multimodal labeling results.
[0084] The educational resource labels are updated based on the similarity between sample data and label data in the automatic labeling model for educational resources, thus completing the update and optimization of the automatic labeling model for educational resources.
[0085] Another aspect of the present invention provides an electronic device, including a processor and a memory;
[0086] The memory is used to store programs;
[0087] The processor executes the program to implement the method described above.
[0088] Another aspect of this invention provides a computer-readable storage medium storing a program that is executed by a processor to implement the methods described above.
[0089] This invention also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned method.
[0090] The specific implementation process of the present invention will now be described in detail with reference to the accompanying drawings:
[0091] To address the technical problems existing in the prior art, this invention provides a method for automatic annotation of multimodal educational resources using temporal causal heuristic pre-training. This model utilizes temporal causal heuristic information applied to large-scale multimodal educational resources, starting from the perspective of the educational resource usage process. It simultaneously utilizes data from the educational resource usage process and data from the multimodal educational resources themselves to enhance the causal essence representation learning of multimodal educational resources, thereby effectively improving the accuracy and robustness of automatic annotation of multimodal educational resources.
[0092] Figure 1 A flowchart illustrating the overall steps of an embodiment of the present invention; as shown below. Figure 1 As shown, the overall method of the present invention includes: constructing a descriptive framework for multi-level educational resource usage records; collecting multimodal educational resource usage records according to the descriptive framework, and performing temporal causal intervention on the collected educational resource usage records; performing fusion representation on the multimodal usage process data of educational resources; constructing an online-offline incremental learning annotation model based on the fusion representation of the multimodal educational resources; inputting the multimodal educational resources to be annotated into the online-offline incremental learning annotation model, and outputting the educational resource annotation results.
[0093] The implementation process of each step is described in detail below:
[0094] 1. Description of the three-tiered educational resource usage record:
[0095] The usage records of multimodal educational resources in online learning environments have complex, cross-temporal and spatial characteristics, and are difficult to perceive due to changing contexts. This invention utilizes a semantic description mechanism to perceive and semantically organize the data of the educational resource usage process, providing a data source for subsequent educational resource representation pre-training and automatic annotation tasks.
[0096] A three-tiered resource usage record description mechanism: Based on the characteristics of online education, a three-tiered semantic description mechanism for resource usage—"behavior-goal-state"—is established. This three-tiered mechanism can be formally described as follows:
[0097] action(id,id,t),egread(studentA,BookAch01,t), (1.1.1)
[0098] Op(X,Y),pre:cond1,…,post:{bool,list(Op),S i},state:S i (1.1.2)
[0099] S = {S1,S2,…,S} n},S i =(s1,s2,…,s N ), (1.1.3)
[0100] Equation (1.1.1) describes the behavior layer (first layer), where `action` represents a generalized action (e.g., reading), and `id` represents an identifier that can identify a specific educational user type or resource. Equation (1.1.2) describes the target layer (second layer), where `Op` represents an operation, `pre` represents the antecedent of the operation, `post` represents the consequent of the operation, and `state` represents the set of resource usage effectiveness evaluation states. Equation (1.1.3) describes the resource usage state layer (third layer), where S... i This represents a set of evaluation indicators for the effectiveness of resource utilization, such as rating scores, sentiment indicators, etc., and can also be empty.
[0101] Construction of a three-tier resource usage record description mechanism:
[0102] The first layer of construction: The construction of this description mechanism needs to be based on the characteristics of resource usage in the learning scenario. Using the temporal predicate logic expression method, a semantic model of the interaction relationship between "user-resource-action" is constructed to describe the interaction relationship between users and resources in the learning scenario. A vocabulary of elements such as behavior, goal, and educational resource usage status and their instance relationship set are established.
[0103] The second layer of construction: Next, drawing on the BDI model and introducing knowledge from the learning resource domain, we construct an action-goal-state operation semantic representation with action antecedent-consequence-usage state. We introduce the goal and resource usage state into the operation semantics of user actions, realizing the association between the three levels of action, goal, and resource usage state.
[0104] The third layer of construction: as described in equation (1.1.3), s i Let {s} represent the set of evaluation indicators for the specific resource utilization effect at time i. i,grade ,s i,em ,…},as s i,grade Resource rating score, s i,em These evaluation metrics, which indicate sentiment tendencies, can be partially extracted directly from the logs, partially extracted indirectly using pre-trained machine learning models, or simply left blank. The absence of some data does not affect subsequent resource labeling tasks.
[0105] 2. Perception of data during the use of educational resources:
[0106] Based on online streaming log data and multimedia sensor data (including online structured learning log data, audio, text, video, and other unstructured data), and leveraging multiple types of deep neural network learning algorithms, combined with a three-level resource usage record description mechanism, a method for user and resource identification and user-resource-action interaction relationship extraction in the learning scenario is constructed. This completes the semantic description of the online cross-temporal original resource usage process data. The semantic process data can be represented as {(d raw ,L t ,t link )}, where d raw L represents the raw collected data. t This indicates that a semantic tag set is obtained from a three-layer semantic description mechanism. The value t can be empty. link It represents a time information record field, thereby enabling semantic perception and organization of multimodal time-series process data.
[0107] 3. Applying temporal causal intervention to represent learning through educational resources:
[0108] In order to learn unbiased cross-temporal deep semantic feature representations of educational resources and overcome the inherent defects of association-driven deep learning, this invention improves the unbiased representation of educational resource features by performing temporal causal intervention-based pre-training under a large-scale multimodal pre-training framework by calculating the probability of the Mask operator.
[0109] Continuous temporal causal deconfusion description: Considering the temporal causal characteristics of multimodal educational resource use in the learning process, the deconfusion process is modeled from the perspective of temporal distributed causal relationships, and can be described as follows:
[0110]
[0111] in, This indicates that the temporal causal confusion operator is considered, and cond(t) represents the temporal causal distribution.
[0112] Discrete event fitting to continuous temporal causality for deconfusion:
[0113] Since the de-obfuscation process in equation (1.3.1) is described in the continuous time domain, distribution estimation is difficult and computationally complex. This invention, based on the causal relationship of time-series events in the learning process, utilizes the semantically perceptible label {L} from the learning process data. t The set is used to construct a temporal causal relationship constraint pre-training intervention mechanism, and to fit the continuous time domain decontamination process representation of equation (1.3.1) from the perspective of discrete events. The process is as follows, including three steps:
[0114] 1) First, based on learned data, semantically aware labels {L} are created. t} sets, constructing sets of event-event (EE) and event-time (ET) relationships and a set of time series labels for the learning behavior process. T and the causal relationship label set R C , where R T To label the temporal relationships of behavioral events, such as "before", "after", "included", "simultaneous", etc., R... c To mark the causal relationship on an event, it can have three values: "triggered", "triggered", and "no connection".
[0115] 2) Then, on the time-series dimension of the learning process data, a method for calculating the time-series label constraints of learning behavior events is constructed, and its optimization solution process is as follows:
[0116]
[0117] Among them, rate ee rate et For scoring functions, Let Y be the constraint space, and the scoring function contains the prior probability distribution, which can be implemented using the SoftMax function;
[0118] 3) Next, combining equation (1.3.2), a complete temporal causal relationship constraint method is constructed. Regularization constraints are then calculated for the Mask position selection of the learning process data. The optimization process is as follows:
[0119]
[0120]
[0121] in, and Both are scoring functions. and Let ω be the indicator function. Y For the search constraint space.
[0122] Intervention based on discrete-time Mask operator pre-training:
[0123] Combining equation (1.3.3) for optimization search, a multimodal x-Transformer-based discrete-time Mask operator-based intervention learning mechanism is constructed. In x-Transformer, x represents different modalities of data (such as text, video, image, physiological data, etc.). This architecture will be described in detail later. This intervention learning mechanism mainly introduces the Mask operator into the multimodal input unit, i.e., {d...} x,i,t Mask i,t}, where d x,i,t Mask represents the i-th input unit of the modal data x of resource usage at time t. i,t The value of is calculated by combining the Mask operator with equation (1.3.3), and the data type and value range can be set according to actual needs. Using the Mask operator to intervene in the training process helps the model learn general unbiased deep semantic features for the automatic annotation task of educational resources.
[0124] 4. Fusion representation of multimodal educational resource usage data
[0125] Considering that the cross-temporal data in the learning process contains rich modal complementary information, we can perform modal and temporal synchronous fusion on them, and at the same time, provide causal heuristic information by means of the above-mentioned temporal causal intervention mechanism, which can effectively promote the learning of stable representations of multimodal educational resources.
[0126] Multimodal feature extraction backbone network architecture design:
[0127] First, based on semantically organized learning process data, this invention adopts the flexible x-Transformer as the underlying general backbone network, where x represents different modal data. Combining the characteristics of different data types (such as text, video, image, physiological data, etc.), the x-Transformer architecture is optimized and adjusted using domain-standard optimization methods, including setting the number of multi-heads and the size of the model's Q, K, and V.
[0128] Calculation of modal correlation of data during resource usage process:
[0129] To enhance the consistency and stability of multimodal data information in the process of educational resource use, feature associations of different modalities are further extracted, and an association degree calculation function is constructed for cross-modal attention fusion enhancement. The construction process is as follows: Define the model input set as... in Let m be a pair of tuples, representing the m-th tuple. j The i-th time step of the data modality includes the original modality data, semantic labels (which can be empty), and the Mask operator bits, i.e., {{(d raw ,L t ,t link Mask i,t M represents the set of data modal types, and the feature embedding parameter matrix set is... The correlation between any two modal data at time i can be expressed as:
[0130]
[0131] Where, m j ,m k ∈M, SoftMax is the normalization function, and C(·) represents the confidence function, which can be implemented using a random gated neural network.
[0132] Cross-modal temporal attention fusion:
[0133] To address the potential information asymmetry or mismatch in temporal data across different resource usage modalities, a cross-modal temporal attention fusion model is constructed. First, the data is combined at intervals of ΔT to collect data at t time points with temporal alignment. Then, this data is input into a recurrent network consisting of stacked cross-modal attention fusion modules. Data from the same time point are first fused, and then feature diffusion is performed along the temporal dimension to obtain modality-invariant semantic representations. This process can be modeled as follows:
[0134]
[0135] Where [·] represents feature concatenation, U is the traversal symbol, σ is the activation function, CMA(·) represents the cross-modal attention fusion module, and η i and Let m represent the diffusion factor and mode m at time i, respectively. j The state variables of the data, in particular,
[0136] Specifically, the cross-modal attention fusion module consists of multiple cross-modal attention mechanisms. It consists of a correlation-weighted fusion module, defined as follows:
[0137]
[0138] in, This indicates that for mode m j The implicit weights of the data are Q, K, and V, which represent the query, key, and value input at each time step, respectively.
[0139] By combining the above modules and intervention mechanisms, the synchronous and efficient fusion of data on the use of multimodal and cross-temporal educational resources is achieved, and the resulting set of features for a specific educational resource is denoted as {Z}. id,if,t}, where id represents the unified identifier of the educational resource that needs to be labeled, if represents the feature identifier of the usage process of the resource, and t represents the time information.
[0140] 5. Automatic labeling of online and offline incremental learning resources:
[0141] 1) Offline Incremental Learning: Resource semantic annotation requires a tag library, an element vocabulary and instance relationship set from the three-layer semantic description, and representational features of resource usage processes. This information is dynamically changing and growing, necessitating a dynamic incremental learning method to incorporate this changed or newly added information into the resource annotation task. This enhances the accuracy and adaptability of educational resource annotation to changing contexts. This invention requires incremental learning in two main categories: semantic learning, including the tag library, element vocabulary and instance relationship set from the three-layer description, and significant representational features of resource usage processes, i.e., features that significantly differ from existing representational features. These two incremental learning methods are described below:
[0142] The first approach, targeting semantic increments such as tags or words, firstly employs domain-specific semantic ontology techniques to perform semantic analysis (including synonyms and near-synonyms) on the sampled text. Tags or words deemed non-new are directly ignored. For the remaining text words, which cannot be directly classified as new or non-new, multimodal information related to these words is collected, and the aforementioned pre-trained model is used to obtain relevant multimodal features. Using these multimodal features and corresponding text words as binary features, neural network techniques are used to construct a corresponding binary classification learning task, namely, new and non-new. For tags or words determined to be new, they are added to the corresponding vocabulary, and the relevant multimodal features are retained to prepare for the next incremental learning.
[0143] The second approach involves incremental learning of salient representation features of resource usage processes. First, the pre-trained model described earlier is used to acquire semantic feature representations of these resource usage process data. These acquired semantic feature representations are then compared with existing semantic representations related to the resource usage process of the corresponding educational resource. The saliency is used to determine whether a feature is newly added. The determination of saliency can be achieved using neural network technology to construct a corresponding saliency regression task, setting a threshold for saliency to define a feature as salient. For features determined to be newly added, they are added to the corresponding educational resource usage process feature library, and the relevant textual semantics are retained to prepare for the next incremental learning iteration.
[0144] 2) Online Annotation Application: Utilizing the feature acquisition method for time-domain causal heuristic intervention training proposed in this invention, the problem of weak causal information learning ability in deep learning can be effectively overcome. The acquired feature representations related to the educational resource usage process can better reflect the usage scenarios of educational resources, making automatic annotation of educational resources more accurate. The resource annotation proposed in this invention mainly adopts an online continuous hot annotation method, meaning that educational resource annotation is not a one-time process; the labeled tags can be continuously updated. The automatic annotation steps for educational resources proposed in this invention are as follows:
[0145] (1). Construct the feature representation of the educational resources that need to be labeled. The required data includes the deep semantic features of the educational resources themselves extracted by the feature extraction network, the usage process data representation features of the educational resources obtained by the aforementioned pre-trained model, and the corresponding related semantic labels or words, label preference data (set according to the labeling task scenario, such as subject knowledge point labeling, or security level labeling, etc.). Using these provided data, obtain the complete global representation of the resource (which can be implemented by designing an end-to-end or non-end-to-end variable input neural network).
[0146] (2). Obtain the deep semantic representation of the tags in the tag library, which includes the semantic embedding representation of the tags themselves and the related multimodal deep semantic representation features. The global representation of the tags can be obtained by position pooling.
[0147] (3). Using the resource representation and label representation obtained in steps 1 and 2, construct an automatic labeling model for educational resources. The specific implementation model can be implemented using a deep neural network. The main task is to calculate the similarity between the educational resource representation and the label representation. If the label representation and the educational resource representation to be labeled are greater than a certain threshold, the resource can be labeled with the label.
[0148] (4) If continuous updates are needed, simply repeat steps 1 to 3 and update the educational resource tags according to the similarity score.
[0149] In summary, the present invention provides a temporal causal heuristic pre-training method for automatic annotation of multimodal educational resources. This model utilizes temporal causal heuristic information applied to large-scale multimodal educational resources, starting from the perspective of the educational resource usage process. It simultaneously utilizes data from the educational resource usage process and data from the multimodal educational resources themselves to enhance the causal essence representation learning of multimodal educational resources, thereby effectively improving the accuracy and robustness of automatic annotation of multimodal educational resources.
[0150] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this invention are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and sub-operations described as part of a larger operation are executed independently.
[0151] Furthermore, although the invention has been described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the described functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the invention. Rather, given the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of conventional skill of an engineer. Therefore, those skilled in the art can implement the invention as set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of the invention, which is determined by the full scope of the appended claims and their equivalents.
[0152] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0153] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0154] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0155] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0156] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0157] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
[0158] The above is a detailed description of the preferred embodiments of the present invention, but the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention, and these equivalent modifications or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A method for automatic annotation of multimodal educational resources based on temporal causal information, characterized in that, include: Construct a descriptive framework for multi-level educational resource usage records; Multimodal educational resource usage records are collected based on the described framework, and temporal causal intervention is performed on the collected educational resource usage records. Data on the multimodal use of educational resources is integrated and represented. Based on the fusion representation of the multimodal educational resources, an online-offline incremental learning annotation model is constructed; Input the multimodal educational resources to be labeled into the online-offline incremental learning labeling model, and output the labeling results of the educational resources; The aforementioned descriptive framework for constructing multi-level educational resource usage records specifically involves: establishing a three-level semantic description mechanism for resource usage based on the characteristics of online education scenarios, namely "behavior-goal-state". This step includes: Constructing a behavioral description layer: Based on the characteristics of resource usage in learning scenarios, and using the temporal predicate logic expression method, a semantic model of interaction relationships between "user-resource-action" is constructed. The semantic model of interaction relationships is used to describe the interaction relationships between users and resources in learning scenarios, and to establish a vocabulary and instance relationship set for each element. The elements include, but are not limited to, behavior, goal, and educational resource usage status. Constructing a target description layer: Based on the BDI model and incorporating knowledge from the learning resource domain, construct an action-target-state operation semantic representation with action antecedent, action consequent, and usage state. Introduce the target and resource usage state into the operation semantics of user actions to achieve the association between the three levels of action, target, and resource usage state. Construct a layer describing resource usage status: extract a set of resource usage performance evaluation metrics directly from log records or indirectly using pre-trained machine learning models; The collection of multimodal educational resource usage records based on the described framework includes: Based on online streaming log data and multimedia sensor data, and with the help of multi-type deep neural network learning algorithms, combined with a three-level resource usage semantic description mechanism, we construct a method for user and resource identification and user-resource-action interaction relationship extraction in the learning scenario, complete the semantic description of online cross-time domain original resource usage process data, and realize the semantic perception and organization of multimodal time-series process data. The semantic process data is represented as follows: , This represents the original collected data. This indicates that a set of semantic tags is obtained from a three-layer semantic description mechanism. This field represents a record of time information. The online streaming log data and multimedia sensor data include, but are not limited to, online structured learning log data, audio data, text data, and video data. The aforementioned temporal causal intervention on the collected educational resource usage records includes: To address the temporal causal characteristics of multimodal educational resource use in the learning process, a deconfusion process model is developed from the perspective of temporal distributed causal relationships. Based on the perspective of causal relationships of time-series events in the learning process, and with the help of a set of semantically perceptible labels for the learning process data, an intervention mechanism for pre-training constrained by time-series causal relationships is constructed, and the process representation is deconfused in the continuous time domain by fitting equations from the perspective of discrete events. Construct a multimodal x-Transformer-based discrete temporal Mask operator intervention learning mechanism, where x in x-Transformer represents different modal data; The step of constructing an online-offline incremental learning annotation model based on the fusion representation of the multimodal educational resources includes: For the semantic incremental class of tags or words, semantic ontology technology is first used to perform semantic analysis on the sampled text words. For newly added text words, multimodal information of the newly added text words is collected, and multimodal features of the newly added text words are obtained. The multimodal features and the corresponding text words are used to construct double tuple features. Based on the double tuple features, a corresponding binary classification learning task is constructed to identify newly added text words and non-newly added text words, thus completing the first type of incremental learning. To identify salient representation features of the resource usage process, the semantic feature representations of the resource usage process data are first obtained. These obtained semantic feature representations are then compared with existing semantic representations related to the resource usage process corresponding to the educational resource. The salientity is used to determine whether a new salient representation feature is added. For newly identified salient features, they are added to the corresponding educational resource usage process feature library, and the relevant text semantics are retained to prepare for the next incremental learning, thus completing the second type of incremental learning.
2. The automatic annotation method for multimodal educational resources based on temporal causal information according to claim 1, characterized in that, Based on the perspective of causal relationships between temporal events in the learning process, and leveraging the semantically perceptible tag set of the learning process data, an intervention mechanism for pre-training constrained by temporal causal relationships is constructed. This mechanism fits an equation from the perspective of discrete events to deconfound process representations in the continuous time domain, including: Based on the set of semantically perceptible labels of learned data, we construct the event-event, event-time relationship sets, time series label sets, and causal relationship label sets for the learning behavior process. In the temporal dimension of learning process data, a method for calculating temporal label constraints of learning behavior events is constructed. A complete temporal causal relationship constraint method is constructed to solve the regularization constraint for the selection calculation of Mask position for learning process data.
3. The automatic annotation method for multimodal educational resources based on temporal causal information according to claim 1, characterized in that, The fusion representation of data on the multimodal use of educational resources includes: Based on semantically organized learning process data, the flexible x-Transformer is used as the underlying general backbone network. Combined with the characteristics of different data types, the x-Transformer architecture is optimized and adjusted using domain-general optimization methods; where x represents different modal data. Feature associations are extracted from data of different modalities, and an association degree calculation function is constructed for cross-modal attention fusion enhancement. To construct a cross-modal temporal attention fusion model, the data is first combined at preset intervals to collect data at each temporally aligned moment. Then, the data is input into a recurrent network consisting of stacked cross-modal attention fusion modules. The data at the same moment are fused first, and then feature diffusion is performed in the temporal dimension to obtain modality-invariant semantic representations. The data on the use of multimodal and cross-temporal educational resources are fused to obtain a set of features for specific educational resources.
4. The automatic annotation method for multimodal educational resources based on temporal causal information according to claim 1, characterized in that, The process of inputting the multimodal educational resources to be labeled into the online-offline incremental learning labeling model and outputting the labeling results of the educational resources includes: Construct a feature representation of the educational resources that need to be labeled; Obtain the deep semantic representation of tags in the tag library; Construct an automatic labeling model for educational resources; The automatic labeling model for educational resources outputs multimodal labeling results. The educational resource labels are updated based on the similarity between sample data and label data in the automatic labeling model for educational resources, thus completing the update and optimization of the automatic labeling model for educational resources.
5. An electronic device, characterized in that, Including the processor and memory; The memory is used to store programs; The processor executes the program to implement the method as described in any one of claims 1 to 4.
6. A computer-readable storage medium, characterized in that, The storage medium stores a program that is executed by a processor to implement the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Education mobile application associated search pushing method and device
CN104462297A
Method for constructing multi-label annotation model of learning resources
CN107590229A