A Method and System for Evaluating Teaching Outcomes on a Cloud-Based Education Platform Based on AIGC
Patent Information
- Application Number
- CN202610794723.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-03
- Publication Date
- 2026-09-11
AI Technical Summary
本发明解决了现有技术存在的评估模态单一,存在语义鸿沟,权重分配僵化,缺乏自适应寻优能力,诊断归因浅层,干预缺乏精准性与生成性的问题
1)消除模态孤岛,实现深层跨模态证据链融合:通过AIGC多模态大模型中的跨模态对齐与时空对齐技术,将原本孤立的文本、音频、视觉、行为模态映射至统一语义空间,利用自注意力机制自动发现隐性关联(如叹气与鼠标停滞),使得评估特征不再是简单拼接,而是蕴含时序与逻辑的深层跨模态证据链。
Smart Images

Figure CN122736387A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of teaching outcome evaluation technology, and in particular to a teaching outcome evaluation method and system based on an AIGC cloud education platform. Background Technology
[0002] With the widespread application of cloud computing and artificial intelligence technologies in education, cloud-based education platforms have become important teaching carriers. Traditional methods for assessing teaching outcomes mainly rely on summative tests (such as exam scores) or simple online behavioral statistics (such as video viewing time and quiz accuracy), which have the following technical shortcomings: 1) Single assessment modality and semantic gap: The learner's learning state is comprehensively represented by multimodal information such as text, voice, facial expression, and behavior. Traditional methods are difficult to achieve deep integration of cross-modal data and cannot capture implicit cross-modal correlation features such as "sighing accompanied by mouse pausing", resulting in one-sided assessment results.
[0003] 2) Rigid weight allocation and lack of adaptive optimization ability: When integrating multi-dimensional features such as cognition, skills and emotions, existing evaluation models usually adopt expert experience weighting or static regression algorithms, which makes it difficult to adaptively adjust the weights according to the data distribution of different learners, resulting in weak model generalization ability.
[0004] 3) Superficial attribution in diagnosis, lack of precision and generative intervention: Most existing technologies only provide cold scores or comments generated by fixed rules. They cannot conduct in-depth attribution diagnosis based on multidimensional evidence like human teachers, nor can they dynamically generate personalized learning paths that include emotional intervention and cognitive scaffolding, resulting in a serious disconnect between "assessment" and "teaching".
[0005] Therefore, there is an urgent need for a cloud-based teaching outcome evaluation method that can deeply integrate multimodal data, adaptively optimize evaluation weights, and possess generative attribution diagnosis and intervention capabilities. Summary of the Invention
[0006] This invention provides a method and system for evaluating teaching outcomes on a cloud-based education platform based on AIGC. This invention addresses the problems of existing technologies, such as a single evaluation modality, semantic gaps, rigid weight allocation, lack of adaptive optimization capabilities, superficial diagnostic attribution, and a lack of precision and generative intervention.
[0007] In a first aspect, embodiments of the present invention provide a method for evaluating teaching outcomes on a cloud-based education platform based on AIGC, the method comprising: Multimodal learning data of learners is collected in real time on the cloud education platform. The collected multimodal learning data is input into the pre-trained AIGC multimodal large model, and multidimensional evaluation feature vectors are extracted through cross-modal alignment technology. A teaching outcome evaluation model containing multi-dimensional weight vectors is constructed, and a fitness function is constructed based on historical labeled data. The fitness function is used to calculate the error between the model evaluation result and the real evaluation label. The improved Grey Goose optimization algorithm is used to iteratively optimize the multi-dimensional weight vector of the teaching outcome evaluation model, resulting in an optimized teaching outcome evaluation model containing the optimal multi-dimensional weight vector. The multidimensional assessment feature vector is input into the optimized teaching outcome assessment model to generate a precise assessment matrix that includes the final assessment score and multidimensional independent scores. Based on the precise assessment matrix, prompt words are generated. Input the prompt words into the AIGC multimodal large model, and use AIGC's generative capabilities to generate personalized attribution diagnosis reports and personalized learning paths; The system analyzes the personalized learning path, extracts the knowledge point tags and recommendation difficulty coefficients to be pushed, and inputs the personalized attribution diagnosis report, knowledge point tags, and recommendation difficulty coefficients into the adaptive learning engine of the cloud education platform to update the learner's knowledge status graph and adjust the push logic of the next stage of teaching content.
[0008] The technical solution provided in this application has at least the following beneficial effects: 1) Eliminate modal silos and achieve deep cross-modal evidence chain fusion: Through cross-modal alignment and spatiotemporal alignment technology in the AIGC multimodal large model, the originally isolated text, audio, visual and behavioral modalities are mapped to a unified semantic space. The self-attention mechanism is used to automatically discover implicit associations (such as sighing and mouse pausing), so that the evaluation features are no longer simple splicing, but a deep cross-modal evidence chain containing temporal sequence and logic.
[0009] 2) Breaking through local optima and achieving adaptive dynamic optimization of evaluation weights: An improved Grey Goose optimization algorithm is introduced, which uses Logistic chaotic mapping to enhance the ergodicity of the initial population. An exploration-development dual-mechanism switching strategy based on convergence factor is designed, which integrates PSO speed mechanism and Grey Goose cooperation idea. This effectively avoids the subjectivity of traditional weighting method and the premature convergence problem of traditional optimization algorithm, and ensures that the weights of the evaluation model are globally optimal.
[0010] 3) Generative attribution diagnosis to achieve a closed loop of "assessment-diagnosis-intervention": Breaking through the limitations of traditional technologies that only output numerical values, it combines a precise assessment matrix with a large model knowledge enhancement mechanism, uses AIGC generative capabilities to output attribution diagnosis reports that include educational theory support, and generates a multi-dimensional and personalized learning path of "emotion-cognition-skill".
[0011] 4) Construct an emotion-sensitive knowledge graph to achieve fine-grained adaptive intervention: When updating learners' knowledge state graphs, break through the traditional framework of only recording cognitive levels, and label emotional characteristics and attribution conclusions as attributes to graph nodes. Combined with difficulty coefficients and multimodal preference matching, achieve a leap from "pushing questions" to "pushing a comprehensive teaching experience that includes scaffolding strategies and emotional intervention".
[0012] In one optional implementation, the multimodal learning data includes text interaction records, voice response data, video behavior images, and operation timing logs; The multidimensional assessment feature vector includes cognitive dimension features, skill dimension features, and emotional dimension features; The AIGC multimodal large model includes a data scheduling layer, a single-modal feature encoder layer, a cross-modal alignment and semantic mapping layer, a multimodal deep fusion and inference layer, a multidimensional evaluation feature vector decoupling and extraction layer, a knowledge enhancement and memory layer, an instruction fine-tuning and prompt word understanding layer, and a multimodal generator layer. The single-modal feature encoder layer includes several single-modal feature encoders, including a visual encoder based on the ViT algorithm, an audio encoder based on the Whisper Encoder, a text encoder based on the LLaMA algorithm, and a temporal encoder based on the LSTM algorithm. The cross-modal alignment and semantic mapping layer includes a modal adapter, a spatiotemporal alignment module, and a contrastive learning refinement module; The multimodal deep fusion and inference layer includes the core Transformer engine.
[0013] In one alternative implementation, learners' multimodal learning data is collected in real time on a cloud-based education platform. This collected multimodal learning data is then input into a pre-trained AIGC multimodal large model. Multidimensional evaluation feature vectors are extracted using cross-modal alignment techniques, including: Multimodal learning data of learners is collected in real time on a cloud-based education platform. The multimodal learning data is preprocessed to obtain preprocessed multimodal learning data, which includes preprocessed video frames, preprocessed audio segments, preprocessed interactive text, and preprocessed operation time sequence segments. The preprocessed multimodal learning data is input into the data scheduling layer of the pre-trained AIGC multimodal large model, and the corresponding single-modal feature encoder in the single-modal feature encoder layer is scheduled according to the modality type of the preprocessed multimodal learning data. Using a single-modal feature encoder, feature encoding is performed on the corresponding preprocessed modal model in the preprocessed multimodal learning data to obtain several high-dimensional feature vectors. The high-dimensional feature vectors include visual feature vectors, audio feature vectors, text feature vectors, and behavioral feature vectors of different modalities. Several high-dimensional feature vectors are input into the cross-modal alignment and semantic mapping layer of the AIGC multimodal large model. Using a modal adapter, the high-dimensional feature vectors of different modalities are mapped to the projection space to obtain a projected token sequence containing all modalities. Using the spatiotemporal alignment module, the projected token sequence is sequentially timestamped, windowed, and feature-fused to obtain a multimodal joint feature sequence containing temporal context. The contrastive learning refinement module utilizes pre-built positive and negative sample pairs to refine the multimodal joint feature sequence containing temporal context by minimizing the InfoNCE loss function, thereby obtaining a multimodal feature sequence in a unified space. The multimodal feature sequence in the unified space is input into the multimodal deep fusion and inference layer of the AIGC multimodal large model. The text instructions input by the learner are concatenated with the multimodal feature sequence in the unified space to obtain an ultra-long input sequence. Based on the self-attention mechanism, the core Transformer engine is used to infer ultra-long input sequences and obtain fused multimodal hidden state sequences. The multi-modal hidden state sequence is input into the multi-dimensional evaluation feature vector decoupling extraction layer of the AIGC multi-modal large model to decouple the multi-dimensional evaluation feature vector, thereby obtaining a multi-dimensional evaluation feature vector including cognitive dimension features, skill dimension features and emotional dimension features.
[0014] In one optional implementation, a teaching outcome evaluation model containing multi-dimensional weight vectors is constructed. A fitness function is built based on historical labeled data. This fitness function is used to calculate the error between the model's evaluation results and the actual evaluation labels, including: Using a multi-dimensional weight vector as the variable to be optimized, a multi-dimensional evaluation feature vector as the input, and the final evaluation score as the output, a teaching outcome evaluation model is constructed, with the following formula: In the formula, X represents the final evaluation score corresponding to the quantity to be optimized; X is a multi-dimensional weight vector, corresponding to the position vector of an individual in the improved grey goose optimization algorithm; X is the quantity to be optimized, defined as an individual in the improved grey goose optimization algorithm. These are the multi-dimensional weights in the multi-dimensional weight vector. To evaluate the cognitive, skill, and affective dimensions of the feature vector in a multidimensional assessment; Scoring for the cognitive dimension; Scoring is based on the skills dimension; Score the emotional dimension; Based on historical labeled datasets, the root mean square error is defined as the fitness function, with the following formula: In the formula, Let X be the fitness value corresponding to the quantity to be optimized. For the historical labeled dataset, the first k The final evaluation score of historical labeled data under the quantity X to be optimized; For the historical labeled dataset, the first k The true labels of historical annotation data; k Indicator values for historical labeled data; N The total number of historical labeled data; the fitness function is used to calculate the error between the model evaluation result and the actual evaluation label.
[0015] In one alternative implementation, an improved grey goose optimization algorithm is used to iteratively optimize the multi-dimensional weight vector of the teaching outcome evaluation model, resulting in an optimized teaching outcome evaluation model containing the optimal multi-dimensional weight vector, including: The multi-dimensional weight vector of the teaching outcome evaluation model is encoded into the position vector of an individual in the improved Grey Goose optimization algorithm. Based on the fitness function, an improved grey goose optimization algorithm is used to iteratively optimize the multi-dimensional weight vector to obtain the optimal multi-dimensional weight vector. Based on the optimal multi-dimensional weight vector, the teaching outcome evaluation model is optimized to obtain the optimized teaching outcome evaluation model containing the optimal multi-dimensional weight vector.
[0016] In one alternative implementation, based on the fitness function, an improved grey goose optimization algorithm is used to iteratively optimize the multi-dimensional weight vector to obtain the optimal multi-dimensional weight vector, including: The chaotic sequence is generated using the Logistic mapping and then mapped to the solution space of individuals in the improved Grey Goose optimization algorithm to obtain an initial population including several initial individuals and the initial velocity of each initial individual. Based on the fitness function, calculate the fitness value of each individual in the initial population or the updated population of the previous iteration, and determine the optimal position of each individual and the global optimal individual of the population based on the fitness value. Based on the individual optimal position and the global optimal individual, a convergence factor and PSO mechanism are introduced, and an improved grey goose optimization algorithm is used to select the exploratory and development behaviors to perform with search control parameters, so as to obtain the updated population of the current iteration. Repeatedly update the position of the population. When the current iteration reaches the maximum number of iterations or the fitness value of the global best individual meets the requirements, terminate the iterative update of the population and output the final global best individual. The position vector of the final globally optimal individual is decoded to obtain the optimal multi-dimensional weight vector.
[0017] In one alternative implementation, a multidimensional evaluation feature vector is input into the optimized teaching outcome evaluation model to generate a precise evaluation matrix that includes the final evaluation score and multidimensional independent scores. Based on the precise evaluation matrix, prompt words are generated, including: The multidimensional assessment feature vector is input into the optimized teaching outcome assessment model to generate the final assessment score, and the corresponding multidimensional independent scores are extracted. The multidimensional independent scores include cognitive dimension scores, skill dimension scores, and affective dimension scores. The final assessment scores, cognitive dimension scores, skill dimension scores, and affective dimension scores are combined to obtain a precise assessment matrix. Extract the dimension corresponding to the lowest score in the precise evaluation matrix as the key variable, and generate the corresponding prompt words based on the prompt word template.
[0018] In one alternative implementation, the prompt words are input into the AIGC multimodal large model, and the AIGC's generative capabilities are used to generate personalized attribution diagnostic reports and personalized learning paths, including: The prompt words and their corresponding precise evaluation matrix and multidimensional evaluation feature vector are input into the instruction fine-tuning and prompt word understanding layer of the AIGC multimodal big model. The task intent in the prompt words is parsed, and the structured numerical features of the precise evaluation matrix and multidimensional evaluation feature vector are transformed into text-descriptive tokens that the AIGC multimodal big model can understand. These tokens are then concatenated with the fused multimodal hidden state sequence to obtain a guided input sequence containing complete instructions and multimodal evidence. Using the knowledge enhancement and memory layer of the AIGC multimodal large model, based on the key variables in the accurate evaluation matrix, the corresponding external knowledge is retrieved from the external database, and the external knowledge is injected into the guided input sequence in the form of additional tokens to obtain the enhanced guided input sequence; The enhanced guided input sequence is fed into the multimodal deep fusion and inference layer of the AIGC multimodal large model. Based on the self-attention mechanism, the core Transformer engine is used to infer the enhanced guided input sequence and obtain the attribution conclusion sequence. The attribution conclusion sequence is input into the multimodal generator layer of the AIGC multimodal large model, and the generative capabilities of AIGC are used to generate personalized attribution diagnostic reports and personalized learning paths.
[0019] In one optional implementation, the personalized learning path is parsed, the required knowledge point tags and recommendation difficulty coefficients are extracted, and the personalized attribution diagnostic report, knowledge point tags, and recommendation difficulty coefficients are input into the adaptive learning engine of the cloud-based education platform. This updates the learner's knowledge state graph and adjusts the push logic for the next stage of teaching content, including: Using named entity recognition algorithms or dependency parsing methods, we parse personalized learning paths, extract core knowledge point entities from the cognitive intervention part of personalized learning paths, and map them to knowledge point tags in the platform knowledge base of cloud education platforms. The recommendation difficulty coefficient is calculated by combining the action modifiers in the personalized learning path and the conclusions of the personalized attribution diagnostic report. The multidimensional evaluation feature vector, the core problem words in the personalized attribution diagnosis report, the knowledge point tags, and the recommendation difficulty coefficient are vectorized and fused to construct the input feature vector of the input state space of the adaptive learning engine. Based on the input feature vector, the learner's knowledge state graph is updated using an adaptive learning engine to obtain the updated knowledge state graph; Based on the updated knowledge state graph, an adaptive learning engine is used to recalculate the priority, path, and presentation format of teaching content, thereby adjusting the logic for pushing teaching content in the next stage.
[0020] Secondly, embodiments of the present invention provide a cloud-based education platform teaching outcome evaluation system based on AIGC, used to implement a cloud-based education platform teaching outcome evaluation method. The system includes: The AIGC feature extraction unit is used to collect learners' multimodal learning data in real time on the cloud education platform, input the collected multimodal learning data into the pre-trained AIGC multimodal large model, and extract multidimensional evaluation feature vectors through cross-modal alignment technology. The teaching outcome evaluation model construction unit is used to construct a teaching outcome evaluation model containing multi-dimensional weight vectors and to construct a fitness function based on historical labeled data. The fitness function is used to calculate the error between the model evaluation result and the actual evaluation label. The iterative optimization unit is used to iteratively optimize the multi-dimensional weight vector of the teaching outcome evaluation model using the improved Grey Goose optimization algorithm, so as to obtain the optimized teaching outcome evaluation model containing the optimal multi-dimensional weight vector. The teaching outcome assessment unit is used to input multi-dimensional assessment feature vectors into the optimized teaching outcome assessment model, generate a precise assessment matrix including the final assessment score and multi-dimensional independent scores, and generate prompt words based on the precise assessment matrix. The attribution and feedback generation unit is used to input prompt words into the AIGC multimodal large model and use AIGC's generative capabilities to generate personalized attribution diagnostic reports and personalized learning paths. The closed-loop optimization unit is used to analyze the personalized learning path, extract the knowledge point tags and recommendation difficulty coefficients to be pushed, input the personalized attribution diagnosis report, knowledge point tags and recommendation difficulty coefficients into the adaptive learning engine of the cloud education platform, update the learner's knowledge status graph, and adjust the push logic of the next stage of teaching content.
[0021] A third aspect of this invention provides an electronic device, which includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by at least one processor, such that the at least one processor can perform the method proposed in the first aspect of the present invention.
[0022] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in the first aspect of the present invention. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the electronic device structure of the hardware operating environment involved in the embodiments of the present invention; Figure 2 This is a flowchart illustrating the steps of a cloud-based education platform teaching outcome evaluation method based on AIGC, as provided in an embodiment of the present invention. Figure 3 This is a functional unit diagram of a cloud-based education platform teaching outcome evaluation system based on AIGC, provided in an embodiment of the present invention. Detailed Implementation
[0024] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0025] The present invention will be further described below with reference to the accompanying drawings.
[0026] Reference Figure 1 , Figure 1 This is a schematic diagram of the electronic device structure of the hardware operating environment involved in the embodiments of the present invention.
[0027] like Figure 1 As shown, the electronic device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0028] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0029] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and an electronic program for a cloud-based education platform teaching outcome evaluation system based on Artificial Intelligence Generated Content (AIGC).
[0030] exist Figure 1 In the illustrated electronic device, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the electronic device of the present invention can be set in the electronic device. The electronic device calls the electronic program of the AIGC-based cloud education platform teaching achievement evaluation system stored in the memory 1005 through the processor 1001, and executes the AIGC-based cloud education platform teaching achievement evaluation method provided in the embodiment of the present invention.
[0031] Reference Figure 2 The present invention provides an evaluation method for teaching outcomes on a cloud-based education platform based on AIGC, the method comprising: S201: Collect learners' multimodal learning data in real time on the cloud education platform, input the collected multimodal learning data into the pre-trained AIGC multimodal large model, and extract multidimensional evaluation feature vectors through cross-modal alignment technology; S202: Construct a teaching outcome evaluation model containing multi-dimensional weight vectors, and construct a fitness function based on historical labeled data. The fitness function is used to calculate the error between the model evaluation result and the actual evaluation label. S203: Using the improved Grey Goose optimization algorithm, the multi-dimensional weight vector of the teaching outcome evaluation model is iteratively optimized to obtain the optimized teaching outcome evaluation model containing the optimal multi-dimensional weight vector. S204: Input the multi-dimensional assessment feature vector into the optimized teaching outcome assessment model to generate a precise assessment matrix that includes the final assessment score and multi-dimensional independent scores, and generate prompt words based on the precise assessment matrix; S205: Input the prompt words into the AIGC multimodal large model, and use AIGC's generative capabilities to generate personalized attribution diagnosis reports and personalized learning paths; S206: Analyze the personalized learning path, extract the knowledge point tags and recommendation difficulty coefficients to be pushed, input the personalized attribution diagnosis report, knowledge point tags and recommendation difficulty coefficients into the adaptive learning engine of the cloud education platform, update the learner's knowledge status graph, and adjust the push logic of the next stage of teaching content.
[0032] The technical solution provided in this application has at least the following beneficial effects: 1) Eliminate modal silos and achieve deep cross-modal evidence chain fusion: Through cross-modal alignment and spatiotemporal alignment technology in the AIGC multimodal large model, the originally isolated text, audio, visual and behavioral modalities are mapped to a unified semantic space. The self-attention mechanism is used to automatically discover implicit associations (such as sighing and mouse pausing), so that the evaluation features are no longer simple splicing, but a deep cross-modal evidence chain containing temporal sequence and logic.
[0033] 2) Breaking through local optima to achieve adaptive dynamic optimization of evaluation weights: An improved grey goose optimization algorithm is introduced, which uses Logistic chaotic mapping to enhance the ergodicity of the initial population. An exploration-development dual-mechanism switching strategy based on convergence factor is designed. The Particle Swarm Optimization (PSO) mechanism and the grey goose cooperation idea are integrated, which effectively avoids the subjectivity of traditional weighting methods and the premature convergence problem of traditional optimization algorithms, and ensures that the weights of the evaluation model are globally optimal.
[0034] 3) Generative attribution diagnosis to achieve a closed loop of "assessment-diagnosis-intervention": Breaking through the limitations of traditional technologies that only output numerical values, it combines a precise assessment matrix with a large model knowledge enhancement mechanism, uses AIGC generative capabilities to output attribution diagnosis reports that include educational theory support, and generates a multi-dimensional and personalized learning path of "emotion-cognition-skill".
[0035] 4) Construct an emotion-sensitive knowledge graph to achieve fine-grained adaptive intervention: When updating learners' knowledge state graphs, break through the traditional framework of only recording cognitive levels, and label emotional characteristics and attribution conclusions as attributes to graph nodes. Combined with difficulty coefficients and multimodal preference matching, achieve a leap from "pushing questions" to "pushing a comprehensive teaching experience that includes scaffolding strategies and emotional intervention".
[0036] In one optional implementation, the multimodal learning data includes text interaction records, voice response data, video behavior images, and operation timing logs; The multidimensional assessment feature vector includes cognitive dimension features, skill dimension features, and emotional dimension features; The AIGC multimodal large model includes a data scheduling layer, a single-modal feature encoder layer, a cross-modal alignment and semantic mapping layer, a multimodal deep fusion and inference layer, a multidimensional evaluation feature vector decoupling and extraction layer, a knowledge enhancement and memory layer, an instruction fine-tuning and prompt word understanding layer, and a multimodal generator layer. The single-modal feature encoder layer includes several single-modal feature encoders, including a visual encoder built based on the VisionTransformer (ViT) algorithm, an audio encoder built based on WhisperEncoder, a text encoder built based on the Large Language Model Meta AI (LLaMA) algorithm, and a temporal encoder built based on the Long Short-Term Memory (LSTM) algorithm. The cross-modal alignment and semantic mapping layer includes a modal adapter, a spatiotemporal alignment module, and a contrastive learning refinement module; The multimodal deep fusion and inference layer includes the core Transformer engine.
[0037] In one alternative implementation, learners' multimodal learning data is collected in real time on a cloud-based education platform. This collected multimodal learning data is then input into a pre-trained AIGC multimodal large model. Multidimensional evaluation feature vectors are extracted using cross-modal alignment techniques, including: S2011: Collect learners' multimodal learning data in real time on the cloud education platform, preprocess the multimodal learning data to obtain preprocessed multimodal learning data, the preprocessed multimodal learning data including preprocessed video frames, preprocessed audio segments, preprocessed interactive text, and preprocessed operation time sequence segments. In this embodiment, face detection and alignment, light compensation, and frame extraction at a fixed frame rate (e.g., 1fps) are performed on the video behavior images to remove images without people or with blurry images. The voice response data is processed by removing silence, reducing environmental noise, normalizing volume, and segmenting into effective voice segments. Perform typo correction, stop word removal, and long text truncation on text interaction records (such as forum posts, code input, and answers); The operation sequence logs (such as mouse trajectory, click frequency, page dwell time, scroll speed) are timestamped, outlier filtered, and segmented into operation sequence segments. S2012: Input the preprocessed multimodal learning data into the data scheduling layer of the pre-trained AIGC multimodal large model, and schedule the corresponding single-modal feature encoder in the single-modal feature encoder layer according to the modality type of the preprocessed multimodal learning data. S2013: Using a single-modal feature encoder, feature encoding is performed on the corresponding preprocessed modal model in the preprocessed multimodal learning data to obtain several high-dimensional feature vectors. The high-dimensional feature vectors include visual feature vectors, audio feature vectors, text feature vectors, and behavioral feature vectors of different modalities. In this embodiment, visual feature encoding: the preprocessed video frame is input into the visual encoder to extract facial expression feature vectors and pose feature vectors, and then combined to obtain visual feature vectors; Audio feature encoding: The preprocessed speech segment is input into the audio encoder to extract acoustic features including intonation, speech rate and stress, and obtain the audio feature vector; Text feature encoding: Input the preprocessed interactive text into the text encoder to extract semantic features and obtain the text feature vector; Behavioral feature encoding: The preprocessed operation time sequence segment is input into the temporal encoder to extract behavioral pattern features and obtain behavioral feature vectors; S2013: Input several high-dimensional feature vectors into the cross-modal alignment and semantic mapping layer of the AIGC multimodal large model, use a modal adapter to perform projection space mapping on the high-dimensional feature vectors of different modalities, and obtain a projected token sequence containing all modalities. Linear / nonlinear projection: Dimensionality reduction and spatial transformation of high-dimensional feature vectors of each modality are performed through learnable multilayer perceptron (MLP) or Q-Former network. For example, visual feature vectors are projected into visual tokens, audio feature vectors are projected into audio tokens, and behavioral feature vectors are projected into behavioral tokens. Aligning the text semantic space: The text feature vector is usually the base space of LLM, without the need for additional projection. The projected non-text tokens (visual tokens, audio tokens, behavioral tokens) are aligned with the text tokens in vector distribution, enabling LLM to read images, sounds and behaviors as if reading text, and obtain the projected token sequence (text token alignment, visual tokens, audio tokens, behavioral tokens). S2014: Using the spatiotemporal alignment module, the projected token sequence is sequentially timestamped, windowed, and feature-fused to obtain a multimodal joint feature sequence containing temporal context. In cloud-based education scenarios, learners’ behavior is multimodal and evolves over time, with time delays between different modalities (e.g., a long pause followed by a sigh, then rubbing the eyes), requiring strict temporal alignment. Timestamp alignment and window segmentation: Based on the main timeline, the projected token sequence containing all modalities is segmented and mapped into a unified time window according to the timestamp information retained in the preprocessing stage; Temporal cross-attention: Within a time window, the cross-attention mechanism is used to calculate the temporal correlation between different modalities. For example, the attention weights are calculated by combining the audio token (sigh) at the current moment with the behavioral token (mouse pause) and visual token (frowning) a few seconds before and after. Strongly correlated cross-modal features are concatenated in the time dimension to form a multimodal joint feature sequence containing temporal context. S2015: Using the contrastive learning refinement module, by utilizing pre-built positive and negative sample pairs and minimizing the InfoNCE loss function, the multimodal joint feature sequence containing temporal context is refined through contrastive learning to obtain a multimodal feature sequence in a unified space. In this embodiment, cross-modal contrastive loss is calculated by using historical multimodal joint feature sequences containing temporal context (such as audio and video behavior segments containing the "confused" label) accumulated by the cloud-based education platform to construct positive and negative sample pairs. Positive sample pairs are historical multimodal joint feature sequences describing the same learning state (such as "sighing audio" and "frowning visual"), while negative sample pairs are historical multimodal joint feature sequences describing different states (such as "focused visual" and "sighing audio"). Feature space fine-tuning: By minimizing the InfoNCE loss function, the distance between positive sample pairs is shortened and the distance between negative sample pairs is widened in the unified semantic space. This process forces the model to eliminate modality-specific noise in each modality (such as differences in facial features between different people and different environmental noises), and extract pure "common semantic features" (such as the common representation of "frustration"), thereby eliminating the semantic gap between modalities and obtaining a multimodal feature sequence in the unified space. S2016: Input the multimodal feature sequence in the unified space into the multimodal deep fusion and inference layer of the AIGC multimodal large model, and concatenate the text instructions input by the learner with the multimodal feature sequence in the unified space to obtain an ultra-long input sequence; In this embodiment, attention weights are calculated: each token in the sequence (whether text, visual, or behavioral) is compared with all other tokens in the sequence to calculate a relevance score; Cross-modal evidence chain construction: This is the most critical generation logic. The engine automatically discovers implicit relationships between multimodal data; for example, even though it is not explicitly stated that "sighing" and "mouse pause" are related, in the self-attention calculation, the audio token representing "sighing" and the behavioral token representing "mouse pause" will be given extremely high attention weights to each other because they appear at the same time and point to the same semantics. S2017: Based on the self-attention mechanism, using the core Transformer engine, it performs inference on ultra-long input sequences to obtain a fused multimodal hidden state sequence; Eliminating Modal Isolations: Through iterations of multiple Transformer networks, the originally isolated modal tokens "look up" and "merge" each other's information, and eventually each token contains contextual information from other modalities, forming a fused latent state where each is intertwined with the others. The multimodal hidden state sequence is fused with the ultra-long input sequence in a one-to-one correspondence in dimensions. However, the vector at each position is no longer a simple modal feature, but a deep semantic representation that integrates instruction intent, temporal context and cross-modal interaction logic. For example, the hidden state vector representing "sighing" in the sequence now contains more than just acoustic features. It has integrated composite information such as "the current question is difficult (text / behavior)", "painful expression (visual)" and "needs emotional intervention (instruction intent)". S2018: The multi-dimensional evaluation feature vector decoupling extraction layer of the AIGC multi-modal large model is input into the multi-dimensional evaluation feature vector decoupling layer to decouple the multi-dimensional evaluation feature vector and obtain a multi-dimensional evaluation feature vector including cognitive dimension features, skill dimension features and emotional dimension features. In this embodiment, cognitive dimension features are decoupled: a "cognitive probe" task is added to guide the model to extract features related to knowledge mastery; cognitive features are obtained mainly by relying on text modality (answer accuracy, logical words) and behavioral modality (knowledge point dwell time) through dimensionality reduction using a fully connected layer; Skill dimension feature decoupling: Add a "skill probe" task to guide the model to extract features related to practical operation; mainly relying on behavioral modalities (code execution error rate, experimental step sequence, mouse click accuracy), and obtain skill features through dimensionality reduction of fully connected layers; Emotional dimension feature decoupling: Add an "emotional probe" task to guide the model to extract features related to emotional state; mainly relying on visual modalities (facial micro-expressions) and audio modalities (tone fluctuations, sighs), emotional features are obtained through dimensionality reduction using fully connected layers.
[0038] In one optional implementation, a teaching outcome evaluation model containing multi-dimensional weight vectors is constructed. A fitness function is built based on historical labeled data. This fitness function is used to calculate the error between the model's evaluation results and the actual evaluation labels, including: S2021: Using a multi-dimensional weight vector as the quantity to be optimized, a multi-dimensional evaluation feature vector as the input, and the final evaluation score as the output, a teaching outcome evaluation model is constructed, with the following formula: In the formula, X represents the final evaluation score corresponding to the quantity to be optimized; X is a multi-dimensional weight vector, corresponding to the position vector of an individual in the improved grey goose optimization algorithm; X is the quantity to be optimized, defined as an individual in the improved grey goose optimization algorithm. These are the multi-dimensional weights in the multi-dimensional weight vector. To evaluate the cognitive, skill, and affective dimensions of the feature vector in a multidimensional assessment; Scoring for the cognitive dimension; Scoring is based on the skills dimension; Score the emotional dimension; S2022: Based on historical labeled datasets, the root mean square error is defined as the fitness function, with the following formula: In the formula, Let X be the fitness value corresponding to the quantity to be optimized. For the historical labeled dataset, the first k The final evaluation score of historical labeled data under the quantity X to be optimized; For the historical labeled dataset, the first k The true labels of historical annotation data; k Indicator values for historical labeled data; NThe total number of historical labeled data; the fitness function is used to calculate the error between the model evaluation result and the actual evaluation label.
[0039] In one alternative implementation, an improved grey goose optimization algorithm is used to iteratively optimize the multi-dimensional weight vector of the teaching outcome evaluation model, resulting in an optimized teaching outcome evaluation model containing the optimal multi-dimensional weight vector, including: S2031: Encode the multi-dimensional weight vector of the teaching outcome evaluation model into the position vector of an individual in the improved grey goose optimization algorithm; S2032: Based on the fitness function, an improved grey goose optimization algorithm is used to iteratively optimize the multi-dimensional weight vector to obtain the optimal multi-dimensional weight vector. S2033: Based on the optimal multi-dimensional weight vector, optimize the teaching outcome evaluation model to obtain the optimized teaching outcome evaluation model containing the optimal multi-dimensional weight vector.
[0040] In one alternative implementation, based on the fitness function, an improved grey goose optimization algorithm is used to iteratively optimize the multi-dimensional weight vector to obtain the optimal multi-dimensional weight vector, including: S20321: Use Logistic mapping to generate chaotic sequences, and map the chaotic sequences to the solution space of individuals in the improved Grey Goose optimization algorithm to obtain an initial population including several initial individuals and the initial velocity of each initial individual. The formula is: In the formula, For the first n+ 1. n There are several chaotic variables whose values range from [0, 1]. The stability coefficient is typically 4. This sequence is ergodic and random, ensuring that the initial population is uniformly distributed in the solution space, avoiding getting trapped in local optima, which is superior to traditional random initialization. n Indicator of chaotic variables; In the formula, For the initial population, the first i An initial individual; For the first i One chaotic variable; These are the upper and lower bounds of the parameter space; i For individual indicators; t This is an indicator of the number of iterations. In the formula, For the initial population, the first iThe initial velocity of each initial individual; A random number in the interval (0,1); This represents the maximum speed. S20322: Based on the fitness function, calculate the fitness value of each individual in the initial population or the updated population of the previous iteration, and determine the optimal position of each individual and the global optimal individual of the population based on the fitness value. S20323: Based on the individual optimal position and the global optimal individual, a convergence factor and PSO mechanism are introduced, and an improved grey goose optimization algorithm is used to select the exploratory and development behaviors to perform with search control parameters, so as to obtain the updated population of the current iteration; In this embodiment, the search control parameters are calculated using the following formula: In the formula, For search control parameters; rand A random number between [0, 1]; This represents the maximum number of iterations. The convergence factor; These are the maximum and minimum values of the convergence factor; like If the gray goose moves away from the prey (optimal weight), the search area needs to be expanded, and the following two exploration mechanisms should be executed with a probability of 0.5 / 0.5: Mechanism 1: Execute the basic exploration fusion PSO speed mechanism, the formula is: In the formula, For the first t+ The first iteration i A newer individual; For the first t The iteration of the ... i Each updated individual is the same as the initial individual during the first iteration; A random number in the interval [0,1]. For the first t+ The first iteration i The speed at which each new pangolin is updated; For the first t The iteration of the ... i The speed of updating the pangolin, in the initial iteration, The initial velocity; Accelerate one's own cognition; The acceleration coefficient of social cognition; For the first t The globally optimal individual in the next iteration; The inertia coefficient; For the first t The iteration of the ... i The optimal position for an individual pangolin; Mechanism 2: Three individuals are randomly selected from the population to perform random migration exploration by greylag geese, using the following formula: In the formula, For the first t+ The first iteration i A newer individual; For individual movement weights; For the first t Three individuals are randomly selected from the updated population in the next iteration; To switch control parameters; For the first t The iteration of the ... i A newer individual; like If the gray goose approaches its prey, it needs to refine its development strategy near the optimal weight, and will execute one of the following two mechanisms with a probability of 0.5 / 0.5: Mechanism 1: Fine-grained development guided by the optimal individual sentinel: Based on the gray wolf cooperative concept, the best, second-best, and third-best individuals are selected from the initial population or the population updated in the previous iteration according to their fitness values. The position is updated based on the best, second-best, and third-best individuals, using the following formula: In the formula, For the first t The first, second, and third potential movement vectors of the next iteration; For the first t The next iteration Distance vectors between the best, second-best, and third-best individuals; This represents the vector of the first, second, and third control coefficients; A random number between [0, 1]; The oscillation coefficient; For search control parameters; The best, second-best, and third-best individuals are identified. The function is for finding the minimum value; Let X be the fitness value corresponding to the variable to be optimized, where X corresponds to the individual variable. Mechanism 2: Develop the optimal neighborhood, the formula is: In the formula, For individual movement weights; The convergence factor; S20324: Repeatedly update the position of the population. When the current iteration reaches the maximum number of iterations or the fitness value of the global best individual meets the requirements, terminate the iterative update of the population and output the final global best individual. S20325: Decode the position vector of the final globally optimal individual to obtain the optimal multi-dimensional weight vector.
[0041] In one alternative implementation, a multidimensional evaluation feature vector is input into the optimized teaching outcome evaluation model to generate a precise evaluation matrix that includes the final evaluation score and multidimensional independent scores. Based on the precise evaluation matrix, prompt words are generated, including: S2041: Input the multidimensional assessment feature vector into the optimized teaching outcome assessment model to generate the final assessment score, and extract the corresponding multidimensional independent scores, which include cognitive dimension scores, skill dimension scores and affective dimension scores. S2042: Combine the final assessment score, cognitive dimension score, skill dimension score, and affective dimension score to obtain a precise assessment matrix; S2043: Extract the dimension corresponding to the lowest score in the precise evaluation matrix as the key variable, and generate the corresponding prompt words based on the prompt word template.
[0042] In one alternative implementation, the prompt words are input into the AIGC multimodal large model, and the AIGC's generative capabilities are used to generate personalized attribution diagnostic reports and personalized learning paths, including: S2051: Input the prompt words and the corresponding precise evaluation matrix and multidimensional evaluation feature vector into the instruction fine-tuning and prompt word understanding layer of the AIGC multimodal large model, parse the task intent in the prompt words, transform the structured numerical features of the precise evaluation matrix and multidimensional evaluation feature vector into text-descriptive tokens that the AIGC multimodal large model can understand, and concatenate them with the fused multimodal hidden state sequence to obtain a guided input sequence containing complete instructions and multimodal evidence; S2052: Using the knowledge enhancement and memory layer of the AIGC multimodal large model, based on the key variables in the precise evaluation matrix, the corresponding external knowledge is retrieved from the external database, and the external knowledge is injected into the guided input sequence in the form of additional tokens to obtain the enhanced guided input sequence; In this embodiment, authoritative theories (such as Weiner's attribution theory and the zone of proximal development theory) are retrieved from an externally mounted educational knowledge base, and historical assessment files and learning trajectories are retrieved from the learner's long-term memory bank. S2053: The enhanced post-guided input sequence is input into the multimodal deep fusion and inference layer of the AIGC multimodal large model. Based on the self-attention mechanism, the core Transformer engine is used to infer the enhanced post-guided input sequence and obtain the attribution conclusion sequence. S2054: Input the attribution conclusion sequence into the multimodal generator layer of the AIGC multimodal large model, and use AIGC's generative capabilities to generate personalized attribution diagnosis reports and personalized learning paths; In this embodiment, the generation of personalized attribution diagnostic reports strictly follows the framework set by the prompt words: Comprehensive profile: An overview of the learner's current status; Interpretation of multidimensional evidence: Combining multimodal feature evidence to explain the scores of each dimension; Core attribution conclusions: Based on the retrieved educational theories, output a professional description of the root cause of the problem; After the attribution diagnosis is completed, the basic attribution conclusion sequence derivation intervention strategy continues, generating a personalized learning path containing the following elements: Prioritize emotional intervention: For abnormalities in the emotional dimension, generate steps to cool down the emotions (such as mindfulness guidance). Cognitive scaffolding construction: For those with weak cognitive dimensions, we recommend suitable micro-courses or advance organizers; Skill downgrading training: To address the lack of skill dimensions, step-by-step exercises or semi-structured tasks are generated; Dynamic goal setting: Based on the zone of proximal development theory, set short-term achievable goals.
[0043] In one optional implementation, the personalized learning path is parsed, the required knowledge point tags and recommendation difficulty coefficients are extracted, and the personalized attribution diagnostic report, knowledge point tags, and recommendation difficulty coefficients are input into the adaptive learning engine of the cloud-based education platform. This updates the learner's knowledge state graph and adjusts the push logic for the next stage of teaching content, including: S2061: Use named entity recognition algorithms or dependency parsing methods to parse personalized learning paths, extract core knowledge point entities from the cognitive intervention part of personalized learning paths, and map them to knowledge point tags in the platform knowledge base of cloud education platforms. S2062: Calculate the recommendation difficulty coefficient by combining the action modifiers in the personalized learning path and the conclusions of the personalized attribution diagnostic report; In this embodiment, the rule mapping is as follows: if words such as "downgraded training" or "semi-code mode" appear in the path, combined with the "frustration" in the emotional attribution, it is determined to be a downward difficulty level and assigned a low difficulty coefficient (e.g., D=0.3); if words such as "challenge advancement" or "expanded application" appear, combined with the attribution of "high skill - low cognition", a high difficulty coefficient is assigned (e.g., D=0.8). Knowledge graph localization: Locate the extracted knowledge point tags in the subject knowledge graph, obtain their preceding and related knowledge points, and provide anchor points for subsequent graph updates; S2063: The multidimensional evaluation feature vector, the core problem words in the personalized attribution diagnosis report (such as "learned helplessness" and "conceptual confusion"), knowledge point tags and recommendation difficulty coefficients are vectorized and fused to construct the input feature vector of the input state space of the adaptive learning engine. S2064: Based on the input feature vector, use an adaptive learning engine to update the learner's knowledge state graph and obtain the updated knowledge state graph; In this embodiment, the cognitive state probability is corrected by using knowledge point tags and cognitive features to update the mastery probability of the corresponding knowledge point node in the graph. If it is attributed to "concept confusion", it will not only reduce the probability of the current knowledge point, but also weaken the probability estimate of its predecessor knowledge points along the edge of the knowledge graph. Emotion and Metacognitive Labeling: Breaking through the limitations of traditional knowledge graphs that only record cognitive levels, this feature labels emotional characteristics and attribution conclusions as attributes on graph nodes. For example, it labels nodes with emotional tags such as resistance and attribution tags such as fear of difficulty. This provides a basis for subsequent emotionally sensitive push notifications. Edge weight adjustment: Based on the erroneous logical connections revealed in the attribution diagnosis (such as confusing two independent concepts), dynamically weaken or strengthen the connection weights between related nodes in the knowledge graph; S2065: Based on the updated knowledge state graph, an adaptive learning engine is used to recalculate the priority, path, and presentation format of teaching content, thereby adjusting the push logic for the next stage of teaching content. In this embodiment, the push priority is reordered: Prioritize strong intervention: If the emotion is marked as "severe frustration", prioritize pushing emotional comfort content or very low-difficulty achievement experience content (breaking the ice). Prioritize filling gaps in advance: If the cause is attributed to "weak foundation", prioritize pushing the prerequisite knowledge points of the current node in the graph, rather than the new content of the current progress. Dynamic difficulty adaptation: The filtering threshold of the question bank or resources is adjusted according to the recommended difficulty coefficient; if the recommended difficulty coefficient is low, the engine limits the probability of high-difficulty questions appearing and gradually increases the difficulty by adopting the "small step principle"; if the recommended difficulty coefficient is high, basic exercises are skipped and comprehensive application questions are pushed directly.
[0044] Multimodal morphology matching: Based on the modal preferences and emotional states revealed in the attribution report, adjust the presentation modality of the content; for example, for learners with "visual fatigue / auditory sensitivity", switch text and image materials to audio explanations; for learners with "behavioral dimension display operation deficiency", push interactive simulation experiments instead of passively watching videos; Scaffolding strategy injection: Based on the attribution conclusions, embed a scaffolding mechanism into the push logic; for example, for the attribution of "skill impairment", automatically enable code prompts or provide semi-finished code frameworks in the next stage of programming questions.
[0045] This invention also provides a cloud-based education platform teaching outcome evaluation system 300 based on AIGC, referring to... Figure 3 The system may include the following units: AIGC feature extraction unit 301 is used to collect learners' multimodal learning data in real time on the cloud education platform, input the collected multimodal learning data into the pre-trained AIGC multimodal large model, and extract multidimensional evaluation feature vectors through cross-modal alignment technology. The teaching outcome evaluation model construction unit 302 is used to construct a teaching outcome evaluation model containing multi-dimensional weight vectors and to construct a fitness function based on historical labeled data. The fitness function is used to calculate the error between the model evaluation result and the real evaluation label. Iterative optimization unit 303 is used to iteratively optimize the multi-dimensional weight vector of the teaching outcome evaluation model using the improved Grey Goose optimization algorithm, so as to obtain the optimized teaching outcome evaluation model containing the optimal multi-dimensional weight vector. The teaching outcome evaluation unit 304 is used to input the multi-dimensional evaluation feature vector into the optimized teaching outcome evaluation model, generate an accurate evaluation matrix including the final evaluation score and multi-dimensional independent scores, and generate prompt words based on the accurate evaluation matrix. The attribution and feedback generation unit 305 is used to input prompt words into the AIGC multimodal large model and use AIGC's generative capabilities to generate personalized attribution diagnostic reports and personalized learning paths. The closed-loop optimization unit 306 is used to analyze the personalized learning path, extract the knowledge point tags and recommendation difficulty coefficients to be pushed, input the personalized attribution diagnosis report, knowledge point tags and recommendation difficulty coefficients into the adaptive learning engine of the cloud education platform, update the learner's knowledge status graph, and adjust the push logic of the next stage of teaching content.
[0046] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus. Memory, used to store computer programs; When the processor executes the program stored in the memory, it implements the AIGC-based cloud education platform teaching outcome evaluation method of the present invention.
[0047] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EI) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. The communication interface is used for communication between the aforementioned terminal and other devices. The memory can include Random Access Memory (RAM), or non-volatile memory, such as at least one disk storage device. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor.
[0048] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0049] Furthermore, to achieve the above objectives, embodiments of the present invention also propose a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the AIGC-based cloud-based education platform teaching outcome evaluation method of the present invention.
[0050] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable hardware devices (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0051] The embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (apparatus), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0052] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0053] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0054] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. "And / or" indicates that either one or both can be chosen. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the element.
[0055] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An AIGC-based cloud education platform teaching achievement evaluation method, characterized in that, The method comprises: Real-time collection of multi-modal learning data of learners on a cloud education platform, input of the collected multi-modal learning data into a pre-trained AIGC multi-modal large model, and extraction of a multi-dimensional evaluation feature vector through cross-modal alignment technology; A teaching achievement evaluation model containing a multi-dimensional weight vector is constructed, and a fitness function is constructed based on historical labeled data, the fitness function being used to calculate the error between the model evaluation result and the true evaluation label; An improved grey goose optimization algorithm is used to iteratively optimize the multi-dimensional weight vector of the teaching achievement evaluation model, and an optimized teaching achievement evaluation model containing an optimal multi-dimensional weight vector is obtained; The multi-dimensional evaluation feature vector is input into the optimized teaching achievement evaluation model to generate an accurate evaluation matrix including a final evaluation score and multi-dimensional independent scores, and a prompt word is generated according to the accurate evaluation matrix; The prompt word is input into the AIGC multi-modal large model to generate a personalized attribution diagnosis report and a personalized learning path by using the generative ability of AIGC; The personalized learning path is analyzed to extract the required knowledge point label and recommended difficulty coefficient, and the personalized attribution diagnosis report, the knowledge point label and the recommended difficulty coefficient are input into the adaptive learning engine of the cloud education platform to update the knowledge state graph of the learner and adjust the push logic of the next stage of teaching content. 2.The AIGC-based cloud education platform teaching achievement evaluation method according to claim 1, characterized in that, The multi-modal learning data includes text interaction records, voice answer data, video behavior images and operation time sequence logs; The multi-dimensional evaluation feature vector includes cognitive dimension features, skill dimension features and emotional dimension features; The AIGC multi-modal large model includes a data scheduling layer, a single-modal feature encoder layer, a cross-modal alignment and semantic mapping layer, a multi-modal deep fusion and reasoning layer, a multi-dimensional evaluation feature vector decoupling extraction layer, a knowledge enhancement and memory layer, an instruction fine-tuning and prompt word understanding layer, and a multi-modal generator layer; The single-modal feature encoder layer includes a plurality of single-modal feature encoders, and the single-modal feature encoders include a visual encoder based on a ViT algorithm, an audio encoder based on a Whisper Encoder, a text encoder based on an LLaMA algorithm, and a time sequence encoder based on an LSTM algorithm; The cross-modal alignment and semantic mapping layer includes a modal adapter, a space-time alignment module, and a contrastive learning refinement module; The multi-modal deep fusion and reasoning layer includes a core Transformer engine. 3.The AIGC-based cloud education platform teaching achievement evaluation method according to claim 2, characterized in that, Real-time collection of multi-modal learning data of learners on a cloud education platform, input of the collected multi-modal learning data into a pre-trained AIGC multi-modal large model, and extraction of a multi-dimensional evaluation feature vector through cross-modal alignment technology, comprising: Real-time collection of multi-modal learning data of learners on a cloud education platform, pre-processing of the multi-modal learning data to obtain pre-processed multi-modal learning data, and the pre-processed multi-modal learning data including pre-processed video frames, pre-processed voice segments, pre-processed interactive texts, and pre-processed operation time sequence segments; The pre-processed multi-modal learning data is input into a data scheduling layer of the pre-trained AIGC multi-modal large model, and according to the modal type of the pre-processed multi-modal learning data, the corresponding single-modal feature encoder in the single-modal feature encoder layer is scheduled; The single-modal feature encoder is used to encode the corresponding pre-processed modal model in the pre-processed multi-modal learning data to obtain a plurality of high-dimensional feature vectors, wherein the high-dimensional feature vectors include visual feature vectors, audio feature vectors, text feature vectors and behavior feature vectors of different modalities; The plurality of high-dimensional feature vectors are input into a cross-modal alignment and semantic mapping layer of the AIGC multi-modal large model, and the modal adapter is used to project the high-dimensional feature vectors of different modalities to obtain a projected Token sequence containing all modalities; The time-space alignment module is used to sequentially align the timestamps, divide the windows and fuse the features of the projected Token sequence to obtain a multi-modal joint feature sequence containing the temporal context; The contrast learning refinement module is used to perform contrast learning refinement on the multi-modal joint feature sequence containing the temporal context by minimizing the InfoNCE loss function using the pre-constructed positive and negative sample pairs to obtain a multi-modal feature sequence in a unified space; The multi-modal feature sequence in the unified space is input into a multi-modal deep fusion and reasoning layer of the AIGC multi-modal large model, and the text instruction input by the learner and the multi-modal feature sequence in the unified space are spliced to obtain a super-long input sequence; Based on the self-attention mechanism, the core Transformer engine is used to reason the super-long input sequence to obtain a fusion multi-modal hidden state sequence; The fusion multi-modal hidden state sequence is input into a multi-dimensional evaluation feature vector decoupling extraction layer of the AIGC multi-modal large model to decouple the multi-dimensional evaluation feature vector to obtain a multi-dimensional evaluation feature vector including cognitive dimension features, skill dimension features and emotional dimension features.
4. The AIGC-based cloud education platform teaching achievement evaluation method according to claim 3, characterized in that, A teaching achievement evaluation model containing a multi-dimensional weight vector is constructed, and a fitness function is constructed based on historical labeled data, which is used to calculate the error between the model evaluation result and the true evaluation label, including: The multi-dimensional weight vector is used as the to-be-optimized quantity, the multi-dimensional evaluation feature vector is used as the input quantity, and the final evaluation score is used as the output quantity to construct the teaching achievement evaluation model, and the formula is: In the formula, is the final evaluation score corresponding to the to-be-optimized quantity X; X is a multi-dimensional weight vector corresponding to a position vector of an individual of the improved grey goose optimization algorithm; X is a to-be-optimized quantity defined as an individual of the improved grey goose optimization algorithm; is a multi-dimensional weight in the multi-dimensional weight vector; is a cognitive dimension feature, a skill dimension feature and an emotional dimension feature in the multi-dimensional evaluation feature vector; is a cognitive dimension score; is a skill dimension score; is an emotional dimension score; Based on the historical labeled data set, the root mean square error is defined as the fitness function, and the formula is: In the formula, is the fitness value corresponding to the to-be-optimized quantity X; is the historical labeled data in the historical labeled data set; k is the final evaluation score of the historical labeled data under the to-be-optimized quantity X; is the historical labeled data in the historical labeled data set; k is the true label of the historical labeled data; k is the to-be-optimized quantity X of the historical labeled data; N is the total number of the historical labeled data; and the fitness function is used to calculate the error between the model evaluation result and the true evaluation label. 5.The AIGC-based cloud education platform teaching achievement evaluation method of claim 4, wherein, An improved grey goose optimization algorithm is used to iteratively optimize the multi-dimensional weight vector of the teaching achievement evaluation model to obtain an optimized teaching achievement evaluation model containing an optimal multi-dimensional weight vector, including: The multi-dimensional weight vector of the teaching achievement evaluation model is encoded into the position vector of the individual of the improved grey goose optimization algorithm; Based on the fitness function, the multi-dimensional weight vector is iteratively optimized using the improved grey goose optimization algorithm to obtain an optimal multi-dimensional weight vector; According to the optimal multi-dimensional weight vector, the teaching achievement evaluation model is optimized to obtain an optimized teaching achievement evaluation model containing an optimal multi-dimensional weight vector. 6.The AIGC-based cloud education platform teaching achievement evaluation method according to claim 5, characterized in that, Based on the fitness function, an improved grey goose optimization algorithm is used to iteratively optimize the multi-dimensional weight vector to obtain the optimal multi-dimensional weight vector, including: The chaotic sequence is generated using the Logistic mapping and then mapped to the solution space of individuals in the improved Grey Goose optimization algorithm to obtain an initial population including several initial individuals and the initial velocity of each initial individual. Based on the fitness function, calculate the fitness value of each individual in the initial population or the updated population of the previous iteration, and determine the optimal position of each individual and the global optimal individual of the population based on the fitness value. Based on the individual optimal position and the global optimal individual, a convergence factor and PSO mechanism are introduced, and an improved grey goose optimization algorithm is used to select the exploratory and development behaviors to perform with search control parameters, so as to obtain the updated population of the current iteration. Repeatedly update the position of the population. When the current iteration reaches the maximum number of iterations or the fitness value of the global best individual meets the requirements, terminate the iterative update of the population and output the final global best individual. The position vector of the final globally optimal individual is decoded to obtain the optimal multi-dimensional weight vector. 7.The AIGC-based cloud education platform teaching achievement evaluation method according to claim 6, characterized in that, The multidimensional assessment feature vectors are input into the optimized teaching outcome assessment model to generate a precise assessment matrix that includes the final assessment score and multidimensional independent scores. Based on the precise assessment matrix, prompt words are generated, including: The multidimensional assessment feature vector is input into the optimized teaching outcome assessment model to generate the final assessment score, and the corresponding multidimensional independent scores are extracted. The multidimensional independent scores include cognitive dimension scores, skill dimension scores, and affective dimension scores. The final assessment scores, cognitive dimension scores, skill dimension scores, and affective dimension scores are combined to obtain a precise assessment matrix. Extract the dimension corresponding to the lowest score in the precise evaluation matrix as the key variable, and generate the corresponding prompt words based on the prompt word template. 8.The AIGC-based cloud education platform teaching achievement evaluation method according to claim 7, characterized in that, Input the prompt words into the AIGC multimodal large model, and leverage AIGC's generative capabilities to generate personalized attribution diagnosis reports and personalized learning paths, including: The prompt words and their corresponding precise evaluation matrix and multidimensional evaluation feature vector are input into the instruction fine-tuning and prompt word understanding layer of the AIGC multimodal big model. The task intent in the prompt words is parsed, and the structured numerical features of the precise evaluation matrix and multidimensional evaluation feature vector are transformed into text-descriptive tokens that the AIGC multimodal big model can understand. These tokens are then concatenated with the fused multimodal hidden state sequence to obtain a guided input sequence containing complete instructions and multimodal evidence. Using the knowledge enhancement and memory layer of the AIGC multimodal large model, based on the key variables in the accurate evaluation matrix, the corresponding external knowledge is retrieved from the external database, and the external knowledge is injected into the guided input sequence in the form of additional tokens to obtain the enhanced guided input sequence; The enhanced guided input sequence is fed into the multimodal deep fusion and inference layer of the AIGC multimodal large model. Based on the self-attention mechanism, the core Transformer engine is used to infer the enhanced guided input sequence and obtain the attribution conclusion sequence. The attribution conclusion sequence is input into the multimodal generator layer of the AIGC multimodal large model, and the generative capabilities of AIGC are used to generate personalized attribution diagnostic reports and personalized learning paths. 9.The AIGC-based cloud education platform teaching achievement evaluation method of claim 8, wherein, The system analyzes personalized learning paths, extracts the required knowledge point tags and recommendation difficulty coefficients, and inputs the personalized attribution diagnostic report, knowledge point tags, and recommendation difficulty coefficients into the adaptive learning engine of the cloud-based education platform. This updates the learner's knowledge status graph and adjusts the push logic for the next stage of teaching content, including: Using named entity recognition algorithms or dependency parsing methods, we parse personalized learning paths, extract core knowledge point entities from the cognitive intervention part of personalized learning paths, and map them to knowledge point tags in the platform knowledge base of cloud education platforms. The recommendation difficulty coefficient is calculated by combining the action modifiers in the personalized learning path and the conclusions of the personalized attribution diagnostic report. The multidimensional evaluation feature vector, the core problem words in the personalized attribution diagnosis report, the knowledge point tags, and the recommendation difficulty coefficient are vectorized and fused to construct the input feature vector of the input state space of the adaptive learning engine. Based on the input feature vector, the learner's knowledge state graph is updated using an adaptive learning engine to obtain the updated knowledge state graph; Based on the updated knowledge state graph, an adaptive learning engine is used to recalculate the priority, path, and presentation format of teaching content, thereby adjusting the logic for pushing teaching content in the next stage.
10. An AIGC-based cloud education platform teaching achievement evaluation system for implementing the cloud education platform teaching achievement evaluation method of any one of claims 1-9. The system includes: The AIGC feature extraction unit is used to collect learners' multimodal learning data in real time on the cloud education platform, input the collected multimodal learning data into the pre-trained AIGC multimodal large model, and extract multidimensional evaluation feature vectors through cross-modal alignment technology. The teaching outcome evaluation model construction unit is used to construct a teaching outcome evaluation model containing multi-dimensional weight vectors and to construct a fitness function based on historical labeled data. The fitness function is used to calculate the error between the model evaluation result and the actual evaluation label. The iterative optimization unit is used to iteratively optimize the multi-dimensional weight vector of the teaching outcome evaluation model using the improved Grey Goose optimization algorithm, so as to obtain the optimized teaching outcome evaluation model containing the optimal multi-dimensional weight vector. The teaching outcome assessment unit is used to input multi-dimensional assessment feature vectors into the optimized teaching outcome assessment model, generate a precise assessment matrix including the final assessment score and multi-dimensional independent scores, and generate prompt words based on the precise assessment matrix. The attribution and feedback generation unit is used to input prompt words into the AIGC multimodal large model and use AIGC's generative capabilities to generate personalized attribution diagnostic reports and personalized learning paths. The closed-loop optimization unit is used to analyze the personalized learning path, extract the knowledge point tags and recommendation difficulty coefficients to be pushed, input the personalized attribution diagnosis report, knowledge point tags and recommendation difficulty coefficients into the adaptive learning engine of the cloud education platform, update the learner's knowledge status graph, and adjust the push logic of the next stage of teaching content.