Intelligent evaluation method and device for 24-point game innovation ability based on multi-modal data

By using multimodal data acquisition and processing technology, combined with visual and language models, the problems of limited interactive scenarios and low accuracy in intelligent assessment of innovation ability have been solved, achieving a more accurate assessment of innovation ability that is suitable for the learning habits of teenagers and large-scale assessment.

CN121338338APending Publication Date: 2026-01-16BEIJING NORMAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511355416.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing intelligent assessment technologies for innovation ability suffer from problems such as limited interactive learning scenarios, low signal acquisition efficiency, and a small number of innovation ability assessment indicators with low accuracy, making it difficult to comprehensively evaluate the innovation ability of gifted children.

Method used

Employing multimodal data acquisition and processing techniques, the experiment recorded participants' operational processes and facial expressions through screen recording, video recording, and audio acquisition. Features were extracted using pre-trained visual and language models, and innovation capabilities were assessed using a Bi-LSTM temporal fusion architecture and a support vector machine classification model, covering problem-solving ability, learning agility, and resilience.

Benefits of technology

It enables more accurate and comprehensive quantitative assessment of innovation capabilities, especially problem-solving skills, learning agility, and resilience, with an accuracy rate improvement of 10%-15%. It aligns with the learning habits and real learning processes of teenagers and supports touchscreen interaction and large-scale online assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121338338A_ABST
    Figure CN121338338A_ABST
Patent Text Reader

Abstract

The invention discloses a 24-point game innovation ability intelligent evaluation method and device based on multi-modal data, and the method comprises the steps: 1, recording four segments of short videos for explaining 24-point intelligence development rapid calculation rules in advance; the system records the total learning time of the tested micro-course in real time; step 2, constructing a 24-point intelligence development rapid calculation test question bank containing five difficulty levels; step 3, recording the operation process of the subject by adopting a screen recording technology, synchronously acquiring expression action videos of the subject by adopting a video recording technology, and recording audio speech data of the subject by adopting an audio acquisition technology; step 4, processing the expression action video acquired in the step 3 by adopting a pre-trained visual network MANET, and extracting facial expression features; and step 5, constructing a time sequence fusion architecture based on a bidirectional long and short time memory network, and performing dynamic modeling on the features of the modes respectively. According to the invention, online micro-class learning is realized by using online short videos, learning agility is evaluated, and problem solving capability and anti-contusion capability are evaluated by using customs clearance tests.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent assessment technology for adolescents' innovation ability, specifically to a method and device for intelligent assessment of 24-point game-based innovation ability based on multimodal data. Background Technology

[0002] As humanity enters the digital age, the world has entered a new era of fierce competition surrounding the development and application of innovative technologies in Artificial Intelligence (AI), particularly Large Language Models (LLM). Talent, especially innovative leading talent, has become a crucial goal for high-quality educational development. Currently, numerous integrated training programs for gifted children have emerged in China, typically using paper-and-pencil exams to assess IQ and psychological states, serving as the basis for admission selection, process monitoring, and exit. In addition to solidly cultivating subject knowledge and enhancing learning and problem-solving abilities, the training process also consciously fosters practical skills, resilience, teamwork, and other innovative abilities.

[0003] With the rapid development of AI, especially LLM and MLLM (Multimodal LLM), a large amount of research has been conducted in recent years on intelligent assessment technology for innovation ability in order to select, cultivate and remove innovative talents more comprehensively and quickly. The assessment environment has transitioned from single-type environments such as exams, programming and discussions to multi-type environments. The signal acquisition method has transitioned from offline acquisition to online acquisition. The types of signals acquired have transitioned from single-modal and dual-modal to multimodal. The indicators of innovation ability have also gradually increased.

[0004] Innovation ability assessment is primarily a psychological assessment focusing on cognition, non-intellectual abilities, and creativity. It serves as the basis for identifying gifted children and nurturing their mathematical and intellectual development. Various assessment indicator systems for student learning abilities have been developed both domestically and internationally. Currently, the four indicators of creativity are mainly used for innovation ability assessment. Multiple assessment scales can be used for creativity evaluation, such as the Guilford Divergent Thinking Test, which measures three characteristics of divergent thinking—fluency, flexibility, and originality—from semantic, symbolic, and graphic perspectives. The Torrance Creative Thinking Test, through verbal and graphic tests, measures four characteristics of divergent thinking: fluency, flexibility, originality, and refinement. The advantages of using creativity assessment tools for innovation ability assessment are that the methods are mature and easy to use, effectively evaluating the divergent and creative thinking problem-solving abilities of test takers. The disadvantage is that it is difficult to assess indicators such as practical skills, teamwork, resilience, and agile learning.

[0005] To address the shortcomings of existing creativity assessments, various improvements have emerged. For example, Zheng Qinhua et al. proposed a perseverance assessment method for evaluating resilience, evaluating perseverance from six perspectives: behavioral, emotional, and cognitive. This method enables semi-automatic assessment through human-computer collaboration, particularly in online learning scenarios. To comprehensively assess innovation capabilities intelligently, Fang Fang et al. proposed a 5-dimensional, 14-angle innovation capability assessment index system, facilitating the assessment of innovation capabilities from different dimensions and perspectives for various learning scenarios.

[0006] Currently, many paper-and-pencil test-based tools are used to identify gifted children, including intelligence tests, academic achievement tests, creativity tests, and personality trait tests. Their advantages are that they facilitate the organization of standardized tests and quickly provide objective and fair assessment results, ensuring the fairness and authority of the selection of gifted children. However, their disadvantages are that they can only assess cognitive level and intelligence, and it is difficult to effectively assess the innovative abilities of test takers from more perspectives, such as learning agility, practical skills, teamwork ability, and resilience.

[0007] With the rapid development of AI, LLM, and MLLM technologies, intelligent assessment technologies for classroom teaching are maturing and becoming more practical. These technologies collect multimodal data, including audio, video, and spoken text, to conduct multimodal intelligent assessments of teacher and student behaviors and emotions, teacher-student interactions, and teaching activities. Compared to the diverse and large-scale teaching activities of a classroom, existing intelligent assessment methods related to innovation ability are geared towards specific, small-scale learning activities involving individuals or groups. These methods are characterized by specific learning activities, fewer participants, and ease of obtaining high-quality single-modal or multimodal data. They can intelligently assess specific innovation ability indicators using relatively simple feature extraction and classification algorithms.

[0008] Current research on intelligent assessment of innovation ability mainly focuses on assessing problem-solving abilities that combine creativity and divergent thinking. In 2021, Kovalkovd et al., in their literature, collected images and audio generated by Scratch programming, extracted image and audio features using shallow networks such as pyAudioAnalysis, and then input them into XGBoost for feature classification, quantifying three indicators of creativity: fluency, flexibility, and originality. In 2025, Hadas et al., in their literature, collected verbal text data for group discussions, extracted text features using ChatGPT4.0, and input them into a shallow classification network, quantifying three indicators of creativity: fluency, flexibility, and originality. Also in 2025, Zhang et al. and Acar et al., in their literature, collected images and their text annotations from children's drawings, extracted multimodal features of text and images using non-LLM networks like ResNet and LLM networks like CLIP, respectively, and input them into shallow classification networks like Random Forest, quantifying creativity indicators such as fluency, flexibility, originality / innovation, and refinement, thus achieving intelligent assessment of divergent thinking.

[0009] In 2025, Zheng Qinhua et al. conducted research on intelligent assessment of resilience for innovation capabilities. They proposed the definition and evaluation indicators of grit, and collected multimodal data such as system logs, operation videos, self-reflection reports, and test questions and answers for online experiments. They used non-LLM tools such as YOLO and BERT to extract multimodal features such as behavior, expression, and text. They adopted a decision fusion strategy to quantify six indicators, including focus and persistence, positive and negative emotions, goal awareness and self-monitoring. Then, through simple weighted calculation, they calculated the quantitative evaluation values ​​of three indicators of grit: behavioral grit, affective grit, and cognitive grit. Finally, they calculated the quantitative evaluation result of grit.

[0010] For example, Chinese invention patent application number CN201811484130.1 discloses a 24-point arithmetic teaching system and method. The 24-point arithmetic teaching system includes at least one AR display device. The AR display device is used to identify, generate, or receive four playing cards, calculate the result of 24 based on the numbers on the four playing cards, and provide calculation prompts for the calculation steps. The beneficial effect of this invention is that it uses an AR display device to enable 24-point arithmetic and 24-point arithmetic teaching anytime and anywhere, increasing the fun and practicality of 24-point arithmetic.

[0011] For example, Chinese invention patent application number CN202310736578.2 discloses a 24-point game device with a buzzer, including a device shell. The surface of the device shell is provided with a main control module and several slave display screen modules, and the buzzer switch module is located in the center of the device shell. The main control module includes a power module, a main control board, a main display screen module, and a touch button module. The main control board has embedded a main control chip, a Bluetooth module, a memory, a data debugging interface, and a slave display screen interface. This device can support multiplayer games and comes with a buzzer, as well as functions for displaying the number of answers and a countdown. The overall structure is simple and clear, the circuit layout is scientific and reasonable, and the operation and display are vivid and interesting. It has high practical value and great potential for promotion and popularization.

[0012] For example, Chinese invention patent application number CN202121022844.8 discloses a math enlightenment game machine, relating to the field of educational game tools. It includes a base with a supporting column fixedly mounted in the middle. A rotating shaft is movably sleeved on the surface of the supporting column, and a rotating disk is fixedly mounted on the outer top of the rotating shaft. Number cards are fixedly mounted on the outer ring of the rotating disk. A housing is fixedly mounted on the top of the base, with the rotating disk inside the housing. A display window is fixedly mounted on the surface of the housing, and the number cards are matched with the display window. This invention uses a drive gear on a drive motor to rotate a driven gear. A limiting block and a limiting slot are matched and engaged to stop the rotating disk, causing the number cards to be randomly displayed in the display window. Based on the displayed numbers, users can play a math puzzle game of calculating 24. The game machine is easy to use and effectively enhances the enjoyment of learning.

[0013] Existing technologies for intelligent assessment of innovation ability suffer from problems such as limited interactive learning scenarios, low signal acquisition efficiency, and a lack of innovation ability assessment indicators with low accuracy. In order to provide a feasible innovation ability assessment scheme for large-scale, normalized selection and digital cultivation of outstanding talents for gifted children, this invention provides a 24-point game-based intelligent assessment method and device for innovation ability based on multimodal data. Summary of the Invention

[0014] To address the aforementioned technical problems in existing technologies, this invention provides a method and apparatus for intelligent evaluation of 24-point game innovation capabilities based on multimodal data.

[0015] The present invention adopts the following technical solution:

[0016] This invention provides a 24-point intelligent evaluation method for game innovation ability based on multimodal data, including:

[0017] Step 1: Pre-record 4 short videos explaining the rules of the 24-point mental arithmetic puzzle. The short videos should at least cover the game instructions, rule introduction, calculation rules, and interface operation demonstration. The participants watch the short videos in order, choosing to watch them in their entirety or skip to the next video. The system records the total learning time t0 of the participants in real time. After watching, the participants will automatically enter the pass assessment stage.

[0018] Step 2: Construct a 24-point mental arithmetic test question bank containing 5 difficulty levels. Randomly select 2 questions from each difficulty level to form a test set of 10 questions. Set the maximum completion time to 20 minutes. When answering question i, the subject enters the equation and submits the answer. If the answer is correct, the subject moves on to the next question. If the answer is incorrect, the subject can resubmit the equation or choose to give up and move on to the next question. The system records the completion time t for each question. i Number of submissions k i and score s i Simultaneously, the total score S, total number of submissions K, and total time T of the 10 questions are calculated to form active measurement data; if the subject does not complete all the questions within 20 minutes, the system automatically terminates the assessment and enters the feature extraction stage;

[0019] Step 3: During the micro-lesson learning and assessment process, screen recording technology was used to record the participants' operation process, video recording technology was used to simultaneously collect the participants' facial expressions and actions, and audio acquisition technology was used to record the participants' audio and speech data, forming audio and video data; through the test design with rapidly increasing difficulty, the participants' focus, engagement, and resilience in learning were stimulated, and data related to learning emotions were collected.

[0020] Step 4: The pre-trained visual network MANET is used to process the facial expression and action videos collected in Step 3 to extract facial expression features; the Chinese-HuBERT-base model is used to parse the audio speech data to extract intonation, prosody, and semantic pause features; the Chinese-RoBERTa-wwm-ext-large-4 deep language model is used to process the arithmetic text input by the subjects to extract text semantic vectors; at the same time, features of the active measurement data in Step 2 are extracted, including the total score S, the total time T, and the time for each question t. i k i s i and the total duration of micro-lesson learning, t0;

[0021] Step 5: Construct a Bi-LSTM temporal fusion architecture based on a bidirectional long short-term memory network. Perform dynamic modeling and multimodal feature fusion on the facial expression features, intonation and prosody features, semantic pause features, text semantic vector features, and active measurement data features extracted in Step 4. Train a support vector machine classification model and use the fused features to evaluate the subjects' problem-solving ability, learning agility, and resilience. Output the quantitative results of each ability index, with the accuracy rate of problem-solving ability evaluation not less than 73% and the accuracy rate of learning agility evaluation not less than 76%.

[0022] Furthermore, step 1 involves pre-recording four short videos explaining the rules of the 24-point mental arithmetic puzzle, including:

[0023] The duration of each short video segment is set to 3-5 minutes, and each short video segment contains at least 2 interactive prompt nodes to prompt the participants to pay attention to the key points in the calculation rules.

[0024] Furthermore, step 2, which involves constructing a 24-point mental arithmetic question bank with 5 difficulty levels, includes:

[0025] Difficulty level 1, containing only mixed addition and multiplication operations, where the sum of two numbers is multiplied by two other numbers to get 24;

[0026] Difficulty level 2, involving three mixed arithmetic operations including one subtraction operation, and obtaining 24 through the combination of "addition, subtraction, and multiplication";

[0027] Difficulty level 3, involving mixed arithmetic operations with one integer division, where 24 is obtained through a combination of addition, subtraction, multiplication, and division;

[0028] Difficulty level 4, involving four arithmetic operations with two different operations, requiring the combination of the two operations to obtain 24;

[0029] Difficulty level 5, a complex arithmetic operation involving parentheses that change the order of operations, requiring adjustment of priority using parentheses to obtain 24.

[0030] Furthermore, the technical parameters for data acquisition in step 3 include:

[0031] The screen recording frame rate is greater than or equal to 30fps, the recording resolution is greater than or equal to 1080P, the frame rate is greater than or equal to 25fps, the audio sampling rate is greater than or equal to 44.1kHz, the bit depth is greater than or equal to 16bit, and the clarity of the acquired audio and video data meets the requirements for subsequent feature extraction.

[0032] Furthermore, step 4 includes:

[0033] The MANET network is used to extract one key frame every 10 frames from the facial expression video. After detecting 68 facial key points on the key frame, a 128-dimensional facial expression feature vector is generated.

[0034] The audio data was divided into frames with a frame length of 20ms and a frame shift of 10ms using the Chinese-HuBERT-base model, and a 256-dimensional intonation and prosodic feature vector was extracted.

[0035] After segmenting the arithmetic text using the Chinese-RoBERTa-wwm-ext-large-4 model, a 768-dimensional semantic vector of the text is generated.

[0036] Furthermore, the timing fusion architecture in step 5 includes:

[0037] The input layer receives a 512-dimensional vector after concatenating multimodal features. The hidden layer has two hidden units, each containing 128 neurons, and uses the tanh activation function. The output layer outputs a 256-dimensional fused feature vector. The support vector machine classification model uses the radial basis function, with a penalty parameter C = 1.0 and a gamma parameter = 0.1. The model training effect is optimized through 5-fold cross-validation.

[0038] This invention also provides a 24-point intelligent assessment device for game innovation ability based on multimodal data, comprising:

[0039] The micro-lesson learning unit stores four short videos of 24-point mental arithmetic rules, has video playback control function, and records the total micro-lesson learning time t0 of the subjects in real time.

[0040] The assessment unit contains a 24-question bank across 5 difficulty levels, with 2 questions drawn from each level to form a 10-question test set. A maximum completion time of 20 minutes is set, and a calculation input interface and a give-up button are provided. The time taken for each question is recorded. i k i s i The total score (S), total number of attempts (K), and total time (T) are used to generate active measurement data. If the assessment is not completed within the time limit, the assessment will automatically terminate.

[0041] The data acquisition unit consists of a screen recording module, a video recording module, and an audio acquisition module, and outputs multimodal audio and video data;

[0042] The feature extraction unit is equipped with a pre-trained MANET visual network, Chinese-HuBERT-base model, and Chinese-RoBERTa-wwm-ext-large-4 model to extract facial expression features, intonation and prosody features, semantic pause features, and text semantic vectors, while also extracting features from actively measured data.

[0043] The intelligent evaluation unit includes a Bi-LSTM temporal fusion module and a support vector machine classification module, which outputs quantitative evaluation results for each capability indicator.

[0044] Furthermore, the intelligent assessment device is developed based on the Android tablet system and has touch screen interaction function; the micro-lesson learning unit, the pass assessment unit, the data acquisition unit, the feature extraction unit, and the intelligent assessment unit are integrated through an Android application. The collected active measurement data, audio and video data, and system logs are stored in the device's local storage module, or uploaded to the GPU server via a wireless network for subsequent data processing.

[0045] Furthermore, the test question bank of the assessment unit supports background updates and maintenance: by adding, deleting or modifying questions of different difficulty levels through the background management interface, the updated test question bank is automatically synchronized to the assessment unit, ensuring the diversity and timeliness of the test questions.

[0046] Furthermore, the intelligent assessment unit also has a result visualization function, displaying the test taker's total score S and the score s for each question in real time. i t i k i The total learning time of the micro-lessons, t0, is displayed, and the quantitative assessment results of problem-solving ability, learning agility, and resilience are presented in the form of bar charts / line graphs for easy viewing by assessors.

[0047] Compared with the prior art, the superior effects of the present invention are as follows:

[0048] 1. The 24-point game innovation ability intelligent assessment method based on multimodal data described in this invention quantifies the subject's rapid learning ability for new rules through the total learning time t0 of micro-lessons and active measurement data. Actual verification shows that the accuracy rate of this indicator reaches 76.9%, filling a gap in existing technology. The method uses a five-level difficulty-increasing level design to stimulate frustration, combined with audio and video recording and the number of submissions k per question. i This study captures participants' perseverance and emotional regulation abilities when facing difficulties, offering a more objective assessment than existing self-report grit tests with an accuracy rate of 65.4%. It utilizes scores (S) and problem-solving time (t) across 10 graded questions. i Text semantic vectors quantify the participants' computational logic and problem-solving abilities with an accuracy rate of 73.1%, making them more suitable for basic ability scenarios for teenagers than existing single-modal programming assessments.

[0049] 2. The 24-point game innovation ability intelligent assessment method based on multimodal data described in this invention uses short video micro-lessons that conform to the fragmented learning habits of teenagers. The progressively increasing difficulty level design can naturally stimulate real learning emotions such as focus, engagement, frustration, and epiphany, which is closer to the real learning process than existing laboratory simulation scenarios. Active behavior data and passive emotion data are recorded simultaneously in the scenario, avoiding the one-sidedness of existing technologies that only collect result data or only collect single behavior data, and providing a data foundation for multi-dimensional ability assessment.

[0050] 3. The intelligent evaluation method for 24-point game innovation ability based on multimodal data described in this invention covers four modalities: active measurement data, visual data, audio data, and text data, which is 1-2 more key modalities than existing technologies. It matches dedicated pre-trained models for different modalities, avoiding the problem of poor adaptability of general models. It adopts Bi-LSTM temporal fusion, which can capture the dynamic correlation between emotions and behaviors during the answering process, improving the accuracy of evaluation by 10%-15% compared with existing static feature fusion.

[0051] 4. The 24-point game innovation ability intelligent assessment method based on multimodal data described in this invention achieves full functionality on an Android tablet, supports touchscreen interaction, requires no professional equipment, and is suitable for scenarios such as schools and training institutions; data can be stored locally or uploaded to a GPU server for processing, accommodating both offline sampling and large-scale online assessment; the question bank supports background updates and can be replaced with other puzzle scenarios such as 16-point mental arithmetic and Sudoku, requiring only adjustments to the micro-lesson content and difficulty grading rules, without reconstructing the core algorithm, making it more reusable than existing scenario-bound assessments; the intelligent assessment unit can display process data and ability indicator charts in real time, facilitating quick interpretation by assessors, and is more practical than existing technologies that only output abstract scores;

[0052] 5. The 24-point intelligent assessment method for game innovation ability based on multimodal data described in this invention achieves process-oriented assessment through full-process data recording and process feature extraction. It records the skipping behavior during micro-lesson learning, the number of submissions / time for each question, and the emotional change curve, rather than only focusing on the final score of 10 questions. The ability assessment results based on process data output can directly provide direction for training. Compared with the single function of existing technologies that only screen and do not provide guidance, it is more in line with the full-link needs of innovative talents in "selection-training-monitoring". Attached Figure Description

[0053] Figure 1 This is a schematic diagram illustrating the working principle of the 24-point intelligent evaluation method for game innovation ability based on multimodal data described in this embodiment of the invention.

[0054] Figure 2 This is a schematic diagram of a multimodal representation framework and a multimodal intelligent evaluation model based on the fusion of visual-speech-text ternary features in an embodiment of the present invention;

[0055] Figure 3 This is a schematic diagram of the startup of the 24-point intelligent evaluation device system for game innovation ability based on multimodal data described in this embodiment of the invention;

[0056] Figure 4 This is a schematic diagram of the micro-lesson learning interface of the 24-point game innovation ability intelligent assessment device based on multimodal data described in this embodiment of the invention;

[0057] Figure 5 This is a schematic diagram of the level-passing assessment interface of the 24-point game innovation ability intelligent assessment device based on multimodal data described in this embodiment of the invention;

[0058] Figure 6 This is a schematic diagram of the display interface of the active evaluation data and system log data of the 24-point intelligent evaluation device for game innovation ability based on multimodal data described in this embodiment of the invention;

[0059] Figure 7 This is a schematic diagram of the test subject's audio and video interface for the online micro-lesson of the 24-point game innovation ability intelligent assessment device based on multimodal data described in this embodiment of the invention;

[0060] Figure 8 This is a schematic diagram of the learning expressions stimulated by the online micro-lesson of the 24-point game innovation ability intelligent assessment device based on multimodal data described in this embodiment of the invention;

[0061] Figure 9 This is a schematic diagram illustrating the problem-solving ability prediction in the 24-point intelligent evaluation method for game innovation ability based on multimodal data described in this embodiment of the invention.

[0062] Figure 10 This is a schematic diagram of learning agility prediction in the 24-point intelligent evaluation method for game innovation ability based on multimodal data described in this embodiment of the invention;

[0063] Figure 11 This is a schematic diagram illustrating the prediction of resilience in the 24-point intelligent assessment method for game innovation ability based on multimodal data described in this embodiment of the invention. Detailed Implementation

[0064] To better understand the above-mentioned objectives, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.

[0065] This invention utilizes a 24-point mental arithmetic puzzle to construct an interactive intelligent assessment method for innovation ability. The method includes five modules: micro-lesson learning, pass assessment, learning status perception, feature extraction, feature fusion and intelligent assessment.

[0066] like Figure 1 As shown, the intelligent evaluation method for 24-point game innovation ability based on multimodal data includes:

[0067] Step 1: Micro-lesson learning. Four short videos explaining the rules of the 24-point mental arithmetic puzzle are pre-recorded. The content of the short videos should at least cover the game instructions, rule introduction, calculation rules, and interface operation demonstration. The participants watch the short videos in order, choosing to watch them completely or skip to the next video. The system records the total micro-lesson learning time t0 of the participants in real time. After watching, the participants will automatically enter the pass assessment stage.

[0068] Step 2: Assessment and Test Completion. A 24-point mental arithmetic test question bank with 5 difficulty levels is constructed. Two questions are randomly selected from each difficulty level, forming a test set of 10 questions. The maximum completion time is set to 20 minutes. When answering question i, the subject enters the equation and submits the answer. If the answer is correct, the subject proceeds to the next question; if the answer is incorrect, the subject can resubmit the equation or choose to skip and proceed to the next question. The system records the completion time t for each question. i Number of submissions k i and score s i Simultaneously, the total score S, total number of submissions K, and total time T of the 10 questions are calculated to form active measurement data; if the subject does not complete all the questions within 20 minutes, the system automatically terminates the assessment and enters the feature extraction stage;

[0069] Step 3, Learning State Perception Module: During the micro-lesson learning and assessment process, screen recording technology is used to record the participants' operation process, video recording technology is used to simultaneously collect the participants' facial expressions and actions, and audio acquisition technology is used to record the participants' audio and speech data, forming audio and video data; through the test design with rapidly increasing difficulty, the participants' focus, engagement, and resilience in learning are stimulated, and learning emotion-related data are collected.

[0070] Step 4, Feature extraction module, such as Figure 2As shown, to construct a dynamic evaluation model for innovation ability in multi-person interactive scenarios, a multimodal representation framework based on visual-speech-text ternary feature fusion is used. The pre-trained visual network MANET is employed to process the facial expression and action videos collected in step 3, extracting facial expression features. The Chinese-HuBERT-base model is used to parse the audio speech data, extracting intonation, prosody, and semantic pause features. The Chinese-RoBERTa-wwm-ext-large-4 deep language model is used to process the arithmetic text input by the participants, extracting text semantic vectors. Simultaneously, features of the actively measured data in step 2 are extracted, including the total score S, total time T, and the time for each question t. i k i s i And the total duration of micro-lesson learning, t0.

[0071] Step 5, Feature Fusion and Intelligent Assessment Module: Construct a Bi-LSTM temporal fusion architecture based on a bidirectional long short-term memory network. Perform dynamic modeling and multimodal feature fusion on the facial expression features, intonation and prosody features, semantic pause features, text semantic vector features, and active measurement data features extracted in Step 4. Train a support vector machine classification model and use the fused features to assess the subjects' problem-solving ability, learning agility, and resilience. Output the quantitative results of each ability indicator, with the accuracy rate of problem-solving ability assessment not less than 73% and the accuracy rate of learning agility assessment not less than 76%.

[0072] Furthermore, step 1 involves pre-recording four short videos explaining the rules of the 24-point mental arithmetic puzzle, including:

[0073] The duration of each short video segment is set to 3-5 minutes, and each short video segment contains at least 2 interactive prompt nodes to prompt the participants to pay attention to the key points in the calculation rules.

[0074] Furthermore, step 2, which involves constructing a 24-point mental arithmetic question bank with 5 difficulty levels, includes:

[0075] Difficulty level 1, containing only mixed addition and multiplication operations, where the sum of two numbers is multiplied by two other numbers to get 24;

[0076] Difficulty level 2, involving three mixed arithmetic operations including one subtraction operation, and obtaining 24 through the combination of "addition, subtraction, and multiplication";

[0077] Difficulty level 3, involving mixed arithmetic operations with one integer division, where 24 is obtained through a combination of addition, subtraction, multiplication, and division;

[0078] Difficulty level 4, involving four arithmetic operations with two different operations, requiring the combination of the two operations to obtain 24;

[0079] Difficulty level 5, a complex arithmetic operation involving parentheses that change the order of operations, requiring adjustment of priority using parentheses to obtain 24.

[0080] Furthermore, the technical parameters for data acquisition in step 3 include:

[0081] The screen recording frame rate is greater than or equal to 30fps, the recording resolution is greater than or equal to 1080P, the frame rate is greater than or equal to 25fps, the audio sampling rate is greater than or equal to 44.1kHz, the bit depth is greater than or equal to 16bit, and the clarity of the acquired audio and video data meets the requirements for subsequent feature extraction.

[0082] Furthermore, step 4 includes:

[0083] The MANET network is used to extract one key frame every 10 frames from the facial expression video. After detecting 68 facial key points on the key frame, a 128-dimensional facial expression feature vector is generated.

[0084] The audio data was divided into frames with a frame length of 20ms and a frame shift of 10ms using the Chinese-HuBERT-base model, and a 256-dimensional intonation and prosodic feature vector was extracted.

[0085] After segmenting the arithmetic text using the Chinese-RoBERTa-wwm-ext-large-4 model, a 768-dimensional semantic vector of the text is generated.

[0086] Furthermore, the timing fusion architecture in step 5 includes:

[0087] The input layer receives a 512-dimensional vector after concatenating multimodal features. The hidden layer has two hidden units, each containing 128 neurons, and uses the tanh activation function. The output layer outputs a 256-dimensional fused feature vector. The support vector machine classification model uses the radial basis function, with a penalty parameter C = 1.0 and a gamma parameter = 0.1. The model training effect is optimized through 5-fold cross-validation.

[0088] This invention also provides a 24-point intelligent assessment device for game innovation ability based on multimodal data, comprising:

[0089] The micro-lesson learning unit stores four short videos of 24-point mental arithmetic rules, has video playback control function, and records the total micro-lesson learning time t0 of the subjects in real time.

[0090] The assessment unit contains a 24-question bank across 5 difficulty levels, with 2 questions drawn from each level to form a 10-question test set. A maximum completion time of 20 minutes is set, and a calculation input interface and a give-up button are provided. The time taken for each question is recorded. i k i s iThe total score (S), total number of attempts (K), and total time (T) are used to generate active measurement data. If the assessment is not completed within the time limit, the assessment will automatically terminate.

[0091] The data acquisition unit consists of a screen recording module, a video recording module, and an audio acquisition module, and outputs multimodal audio and video data;

[0092] The feature extraction unit is equipped with a pre-trained MANET visual network, Chinese-HuBERT-base model, and Chinese-RoBERTa-wwm-ext-large-4 model to extract facial expression features, intonation and prosody features, semantic pause features, and text semantic vectors, while also extracting features from actively measured data.

[0093] The intelligent evaluation unit includes a Bi-LSTM temporal fusion module and a support vector machine classification module, which outputs quantitative evaluation results for each capability indicator.

[0094] Furthermore, such as Figures 3 to 7 As shown, the intelligent assessment device is developed based on the Android tablet system and has touch screen interaction function. The micro-lesson learning unit, the pass assessment unit, the data collection unit, the feature extraction unit, and the intelligent assessment unit are integrated through an Android application. The collected active measurement data, audio and video data, and system logs are stored in the device's local storage module or uploaded to the GPU server via a wireless network for subsequent data processing.

[0095] Furthermore, the test question bank of the assessment unit supports background updates and maintenance: by adding, deleting or modifying questions of different difficulty levels through the background management interface, the updated test question bank is automatically synchronized to the assessment unit, ensuring the diversity and timeliness of the test questions.

[0096] Furthermore, the intelligent assessment unit also has a result visualization function, displaying the test taker's total score S and the score s for each question in real time. i t i k i The total learning time of the micro-lessons, t0, is displayed, and the quantitative assessment results of problem-solving ability, learning agility, and resilience are presented in the form of bar charts / line graphs for easy viewing by assessors.

[0097] Based on the content of this invention, and according to the successfully developed "Intelligent Assessment Tool for Innovation Ability in Online Micro-lesson Learning Scenarios V1.0", on-site sampling was carried out. Among them, 21 samples were collected on the campus of Beijing Normal University, including 10 second-year female students, 10 fourth-year male students, and 1 fifth-grade male student. 3 first-year female students were collected from Ciqu Middle School in Tongzhou District, and a total of 10 junior high school students were collected in Chongqing.

[0098] This invention utilizes online micro-lessons to assess the innovation abilities of junior high school students, primary school students, and university students, achieving results such as... Figure 6 The test results and system logs shown, and even more, achieved the following: Figure 8 The learning expressions shown; for Figure 8 The expressions shown in the nine-square grid are as follows: the first row of three expressions comes from three first-year junior high school girls, all showing focused and engaged expressions of rapid thinking; the second row of three expressions comes from a fifth-grade elementary school boy, showing engaged and frustrated expressions of rapid thinking; the third row of three expressions comes from a fourth-year undergraduate boy who won first prize in a provincial mathematics competition. His first expression is focused and engaged, his second expression is engaged and frustrated, and his third expression is insightful and joyful. The 24-point mental arithmetic assessment tool V1.0 created by this invention provides a sufficiently high level of difficulty to stimulate learning emotions in students of all ages, including focus, engagement, frustration, insight, and joy, thereby assessing the subjects' resilience.

[0099] To quickly conduct intelligent assessments of innovation capabilities, this invention uses... Figure 6 The test scores and system logs shown extract multimodal features from the micro-lesson learning scenario, including total score, learning time, test time, scores for questions 1 to 10, number of submissions, solution time, and longest time taken. Finally, a support vector machine classification model is used to train and identify students' problem-solving ability, learning agility, and resilience. Under a verification protocol with only one student remaining, the overall average recognition accuracy achieved is as follows: Figures 9 to 11 The figures show the prediction accuracy rates for problem-solving ability, learning agility, and resilience, respectively. The prediction accuracy rate for problem-solving ability reached 73.1%, the prediction accuracy rate for learning agility reached 76.9%, and the prediction accuracy rate for resilience reached 65.4%.

[0100] This invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the invention as claimed. The scope of protection of this invention is defined by the appended claims.

Claims

1. A 24-point game innovation ability intelligent evaluation method based on multi-modal data, characterized by, The method comprises the following steps: Step 1, pre-recording 4 short videos explaining the rules of 24-point mental arithmetic, the content of the short videos covering at least game guide, rule introduction, operation rule and interface operation demonstration; the subjects watch the short videos in order, choose to watch completely or skip, and the system records the total learning time t0 of the subjects in real time; after watching, the subjects automatically enter the pass test evaluation link; Step 2, construct a 24-point puzzle speed calculation test question bank containing 5 difficulty levels, randomly select 2 questions from each difficulty level to form a test set of a total of 10 questions; set the longest pass time as 20 minutes, when the subject answers the ith question, input the calculation formula and submit the answer, if the answer is correct, proceed to the next question, if the answer is wrong, re-submit the calculation formula or choose to give up and proceed to the next question; the system records the completion time t of each question i , the number of submissions k i and the score si, while counting the total score S of the 10 questions, the total number of submissions K and the total time T, forming the active measurement data; if the subject does not complete all the questions within 20 minutes, the system automatically terminates the evaluation and enters the feature extraction link; Step 3, during the process of micro-lesson learning and pass test evaluation of the subjects, the operation process of the subjects is recorded by screen recording technology, the expression and action video of the subjects is synchronously collected by video recording technology, and the audio speech data of the subjects is recorded by audio collection technology, forming audio and video data; through the test design of difficulty rapid promotion, the learning emotions of the subjects such as concentration, involvement and anti-frustration are stimulated, and the collection of learning emotion related data is realized; Step 4, the expression and action video collected in step 3 is processed by using the pre-trained visual network MANET to extract facial expression features; The audio speech data is analyzed by the Chinese-HuBERT-base model to extract the prosodic rhythm and semantic pause features; the Chinese-RoBERTa-wwm-ext-large-4 deep language model is used to process the formula text input by the subject to extract the text semantic vector; meanwhile, the features of the active measurement data in step 2 are extracted, including the total score S, the total time T, the t of each question i , k i , si and the total micro-lesson learning time t0; Step 5, a Bi-LSTM time sequence fusion architecture based on a bidirectional long short-term memory network is constructed to dynamically model and multi-modal feature fuse the facial expression features, tone prosody and semantic pause features, text semantic vector features and active measurement data features extracted in step 4; The support vector machine classification model is trained, the fused features are used to evaluate the problem solving ability, learning agility and anti-frustration ability of the subjects, and the quantitative results of each ability index are output, wherein the problem solving ability evaluation accuracy is not less than 73%, and the learning agility evaluation accuracy is not less than 76%.

2. The 24-point game innovation ability intelligent assessment method based on multi-modal data according to claim 1, characterized in that, The step 1 of pre-recording 4 short videos explaining the rules of 24-point mental arithmetic comprises: The time length of a single short video is set to 3-5 minutes, and at least 2 interactive prompt nodes are embedded in each short video to prompt the subjects to pay attention to the key points in the operation rules.

3. The 24-point game innovation ability intelligent assessment method based on multi-modal data according to claim 1, characterized in that, The step 2 of constructing a 24-point mental arithmetic test question bank containing 5 difficulty levels comprises: Difficulty level 1, only mixed operation of addition and multiplication, 24 is obtained by adding two numbers and multiplying the other two numbers; Difficulty level 2, three mixed operations containing one subtraction operation, 24 is obtained by combining "addition, subtraction and multiplication"; Difficulty level 3, four mixed operations containing one integer division operation, 24 is obtained by combining "addition, subtraction, multiplication and division"; Difficulty level 4, four mixed operations containing two different operations, 24 is obtained by combining two operations; Difficulty level 5, complex four mixed operations containing operation order change by parentheses, 24 is obtained by adjusting the priority after the parentheses.

4. The 24-point game innovation ability intelligent assessment method based on multi-modal data according to claim 1, characterized in that, The technical parameters of data collection in step 3 comprise: The frame rate of screen recording is greater than or equal to 30fps, the resolution of video recording is greater than or equal to 1080P, the frame rate is greater than or equal to 25fps, the audio sampling rate is greater than or equal to 44.1kHz, the bit depth is greater than or equal to 16bit, and the clarity of collected audio and video data meets the subsequent feature extraction requirements.

5. The 24-point game innovation ability intelligent assessment method based on multi-modal data according to claim 1, characterized in that, The step 4 comprises: Through the MANET network, 1 key frame is extracted every 10 frames of the expression and action video, and after 68 face key point detection is performed on the key frame, a 128-dimensional face expression feature vector is generated; The audio data is divided into frames by the Chinese-HuBERT-base model with a frame length of 20 ms and a frame shift of 10 ms, and 256-dimensional prosodic feature vectors are extracted; After the Chinese-RoBERTa-wwm-ext-large-4 model is used to perform word segmentation on the formula text, a 768-dimensional text semantic vector is generated.

6. The 24-point game innovation ability intelligent assessment method based on multi-modal data according to claim 1, characterized in that, The time sequence fusion architecture in step 5 includes: The input layer receives a 512-dimensional vector after the multi-modal features are spliced, the hidden layer has 2 hidden units, each unit contains 128 neurons, uses a tanh activation function, and the output layer outputs a 256-dimensional fusion feature vector; the support vector machine classification model uses a radial basis kernel function, the penalty parameter C = 1.0, the gamma parameter = 0.1, and the model training effect is optimized through 5-fold cross-validation.

7. A 24-point game innovation ability intelligent evaluation device based on multi-modal data, characterized by It includes: The micro-lesson learning unit stores 4 short videos of 24-point intelligence speed calculation rules, has video playback control function, and records the total learning time t0 of the subject in real time; The pass-through evaluation unit contains a 24-point test question bank with 5 difficulty levels, and 10 test questions are selected from each level to form a test set; Set 20 minutes maximum time, provide formula input interface and abandon button, record each question t i , k i , si and total score S, total number K, total time T, form active measurement data, automatically terminate if not completed within the time limit; The data acquisition unit is composed of a screen recording module, a video recording module, and an audio acquisition module, and outputs multi-modal audio and video data; The feature extraction unit is equipped with pre-trained MANET visual network, Chinese-HuBERT-base model, Chinese-RoBERTa-wwm-ext-large-4 model, respectively extracts facial expression features, prosodic and semantic pause features, text semantic vectors, and active measurement data features; The intelligent evaluation unit includes a Bi-LSTM time sequence fusion module and a support vector machine classification module, and outputs the quantitative evaluation results of each ability index.

8. The 24-point game innovation ability intelligent assessment device based on multi-modal data according to claim 7, characterized in that, The intelligent evaluation device is developed based on the Android tablet system and has touch screen interaction function; the micro-lesson learning unit, the pass-through evaluation unit, the data acquisition unit, the feature extraction unit, and the intelligent evaluation unit are integrated through the Android application program, and the collected active measurement data, audio and video data, and system logs are stored in the local storage module of the device or uploaded to the GPU server for subsequent data processing.

9. The 24-point game innovation ability intelligent assessment device based on multi-modal data according to claim 7, characterized in that, The test question bank of the pass-through evaluation unit supports background update and maintenance: new, delete or modify test questions of different difficulty levels through the background management interface, and the updated test question bank is automatically synchronized to the pass-through evaluation unit, ensuring the diversity and timeliness of the test questions.

10. The 24-point game innovation ability intelligent assessment device based on multi-modal data according to claim 7, characterized in that, The intelligent assessment unit also has a result visualization function, displaying the test taker's total score S, and the scores si, ti, and k for each question in real time. i The total learning time of the micro-lessons, t0, is displayed, and the quantitative assessment results of problem-solving ability, learning agility, and resilience are presented in the form of bar charts / line graphs for easy viewing by assessors.

Citation Information

Patent Citations

  • 24-point operation teaching system and teaching method

    CN109300365A

  • 24-point game device with responder

    CN117357884A

  • Mathematics enlightenment game machine

    CN214541216U