A multimodal scientific experiment evaluation auxiliary method and system based on model coupling
By combining video acquisition, attitude evaluation model and object detection model, multi-dimensional features are extracted and experimental evaluation reports are generated using multi-modal large models, the problem of one-sided experimental evaluation results in the existing technology is solved, and a comprehensive automated evaluation of scientific experiments and efficient evaluation report generation is realized.
Patent Information
- Application Number
- CN202411037765.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-07-31
AI Technical Summary
The existing experimental evaluation methods mainly rely on manual supervision, making it difficult to comprehensively evaluate dynamic changes in the experiment process, resulting in one-sided and lack of flexibility and versatility in the evaluation results.
Multimodal scientific experimental evaluation auxiliary method based on model coupling is used to extract multidimensional features through the combination of video acquisition, pose evaluation model and object detection model, and a detailed experimental evaluation report is generated using multimodal large model.
It realizes a comprehensive and automated evaluation of scientific experiments, can objectively analyze experimental operations, generate detailed evaluation reports, improve the efficiency and quality of experimental teaching, reduce artificial deviations, and improve the accuracy of evaluation through continuous learning.
Smart Images

Figure CN118968382B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital experiment evaluation, and in particular to a multimodal scientific experiment evaluation auxiliary method and system based on model coupling. Background Art
[0002] At present, schools are increasingly paying attention to experimental teaching and assessment in middle schools. At the same time, they emphasize the comprehensive evaluation and monitoring of experimental teaching, require that experimental teaching conditions be included in the education quality evaluation and monitoring system, and require schools to strengthen the process management of experimental teaching to ensure the standardization of experimental teaching content and class hours, as well as the strict implementation of experimental teaching procedures and norms.
[0003] However, the existing experimental evaluation and guidance work mainly relies on manual supervision. Even if some experimental scoring and evaluation schemes based on target detection are used, they mainly capture the static characteristics of the final results or key nodes of the experiment, ignoring the dynamic changes of the experimental process. They fail to fully grasp the correctness of the experimental steps, the standardization of the operations, and the importance of the intermediate results, resulting in one-sided evaluation results. They are also difficult to adapt to the evaluation needs of experiments of different types and difficulties, and lack flexibility and versatility.
[0004] Therefore, finding an experimental evaluation method that can not only comprehensively evaluate the operational standardization of students during the experiment, but also provide multi-dimensional evaluation results to improve the evaluation accuracy is a technical problem that needs to be urgently solved by technical personnel in this field. Summary of the invention
[0005] In view of this, the present invention proposes a multimodal scientific experiment evaluation auxiliary method and system based on model coupling, which can realize comprehensive automated evaluation of scientific experiments, not only can objectively analyze students' experimental operations, but also can generate detailed evaluation reports, greatly improving the efficiency and quality of experimental teaching.
[0006] The technical solution of the present invention is achieved in this way:
[0007] In a first aspect, the present invention provides a multimodal scientific experiment evaluation auxiliary method based on model coupling, comprising the following steps:
[0008] S1. Use a collection device to collect videos of students' experimental operation processes, obtain scientific experiment video sequences, and upload the scientific experiment video sequences to a streaming media server;
[0009] S2, the streaming media server uses the posture assessment model and the target detection model to extract features from the scientific experiment video sequence and fuses them to obtain multi-dimensional features;
[0010] S3, classifying and evaluating the scientific experiment video sequence according to the multidimensional features to obtain classification results and evaluation results;
[0011] S4, obtaining the scoring standard document of the experimental operation, and constructing a vector database of the experimental scoring standard according to the scoring standard document of the experimental operation;
[0012] S5. Generate an experimental evaluation report using a multimodal large model based on a vector database of multidimensional features, classification results, evaluation results, and experimental scoring criteria;
[0013] S6. Use the experimental evaluation report to iteratively optimize the training of the multimodal large model.
[0014] On the basis of the above technical solution, preferably, step S5 specifically includes:
[0015] Generate corresponding prompt information for the multidimensional feature using the experimental scoring criteria in the vector database of the experimental scoring criteria;
[0016] The prompt information is input into the multimodal large model to generate preliminary experimental evaluation content, and the experimental evaluation content is structured to generate a complete experimental evaluation report; the experimental evaluation content includes experimental operation scoring, experimental step scoring, experimental phenomenon scoring, experimental conclusion scoring and operation error prompts.
[0017] On the basis of the above technical solution, preferably, step S2 specifically includes:
[0018] S21, extracting the student's body key point information from the scientific experiment video sequence according to a predetermined frequency using a posture assessment model;
[0019] S22, using the target detection model to identify and locate the position information of the experimental objects and equipment from the scientific experiment video sequence at set intervals;
[0020] S23. Information fusion is performed based on the students’ human body key point information, the experimental objects and the location information of the equipment to form multi-dimensional features.
[0021] On the basis of the above technical solution, preferably, the feature extraction network includes downsampling, a visual attention neural network based on LKA and a Norm layer, wherein the visual attention neural network based on LKA includes deep convolution, deep dilated convolution and channel convolution;
[0022] The prediction key point network adopts the first multi-stage architecture, and each stage gradually refines the prediction of the key point position; each stage consists of a 3×3 convolutional layer, and each stage is optimized using the L2 loss function; the loss function of the prediction key point network is:
[0023]
[0024] Among them, F1(Loss) represents the loss function of the prediction key point network, S represents the total number of stages in the first multi-stage architecture, p represents the key point of the human body, w(p) represents the weight of the key point p, gt(p) represents the true position of the key point p, and pv(p) represents the predicted position of the key point p;
[0025] The limb network adopts a second multi-stage architecture, where each stage consists of a 1×1 convolutional layer.
[0026] Based on the above technical solution, preferably, the network structure of the target detection model includes a lightweight classification head, a downsampling layer and a rank-guided block, wherein the lightweight classification head includes a lightweight architecture of 2 depth-separable convolutions and 1 channel convolution; the downsampling layer includes point convolution and depth-wise convolution.
[0027] Based on the above technical solutions, preferably, in the downsampling layer, point-wise convolution is used to adjust the channel dimension, and depth-wise convolution is used for spatial downsampling. The computational cost and parameter cost are as follows:
[0028]
[0029] N(param)=2c 2 +18c
[0030] Among them, h represents the height of the input feature map, w represents the width of the input feature map, c represents the number of input feature map channels, O represents the computational cost, and N represents the parameter cost.
[0031] More preferably, step S3 specifically includes:
[0032] S31, determining the operation type of each frame of the scientific experiment video sequence according to the multi-dimensional features; the operation types include experimental step arrangement, correct operation step capture, abnormal operation capture, drug dosage detection, experimental environment detection and experimental result capture;
[0033] S32, screening high-confidence operation classification results through a set probability threshold;
[0034] S33, associating the screened operation classification results with the time series to form an experimental step sequence;
[0035] S34, based on the location information of the experimental objects and equipment, calculating the experimental usage in the scientific experiment video sequence and generating an environmental description;
[0036] S35. Extract and classify experimental phenomena from the location information of experimental objects and equipment.
[0037] In a second aspect, the present invention provides a multimodal scientific experiment evaluation assistance system based on model coupling, which adopts the multimodal scientific experiment evaluation assistance method as described above, including:
[0038] The video acquisition module is configured to use an acquisition device to collect video of the student's experimental operation process, obtain a scientific experiment video sequence, and upload the scientific experiment video sequence to a streaming media server;
[0039] A feature extraction module, which is configured as a streaming media server to extract features from the scientific experiment video sequence using a posture assessment model and a target detection model respectively and fuse them to obtain multi-dimensional features;
[0040] An operation classification module is configured to classify and evaluate the scientific experiment video sequence according to the multi-dimensional features to obtain classification results and evaluation results;
[0041] A knowledge base generation module is configured to obtain a scoring standard document for experimental operations and construct a vector database of experimental scoring standards based on the scoring standard document for experimental operations;
[0042] A report generation module configured to generate an experimental evaluation report using a multimodal large model based on a vector database of multidimensional features, classification results, evaluation results, and experimental scoring criteria;
[0043] The model training module is configured to use the experimental evaluation report to iteratively optimize the multimodal large model.
[0044] In a third aspect, the present invention provides a multimodal scientific experiment evaluation auxiliary device, comprising: at least one processor, at least one memory, a communication interface and a bus; wherein:
[0045] The processor, memory, and communication interface communicate with each other via the bus;
[0046] The memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to implement the multimodal scientific experiment evaluation auxiliary method as described above.
[0047] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to implement the multimodal scientific experiment evaluation assistance method as described above.
[0048] The multimodal scientific experiment evaluation auxiliary method of the present invention has the following beneficial effects compared with the prior art:
[0049] (1) By combining video acquisition, feature extraction and classification of scientific experiment video sequences, a comprehensive and automated evaluation of scientific experiments can be achieved. This not only objectively analyzes students' experimental operations, but also generates detailed evaluation reports, greatly improving the efficiency and quality of experimental teaching, significantly reducing human bias, and providing more comprehensive and accurate experimental evaluation results. At the same time, the accuracy of the evaluation is continuously improved through continuous learning.
[0050] (2) By constructing a vector database of experimental scoring standards and combining it with a multimodal large model, we can achieve the automatic generation of high-quality, personalized experimental evaluation reports, which not only provide standardized scores but also give targeted feedback and suggestions based on specific circumstances, greatly improving the depth and educational value of the evaluation.
[0051] (3) By integrating the human body key point information extracted by the posture assessment model and the experimental object and equipment position information identified by the target detection model, multi-dimensional features are formed, which significantly improves the ability to understand the experimental process and makes the classification and evaluation of scientific experiment video sequences more accurate and comprehensive. It can not only evaluate students' operating skills, but also detect the dosage of drugs, experimental environment and experimental results, providing a multi-angle and multi-level experimental evaluation, which greatly improves the comprehensiveness and accuracy of the evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0053] Figure 1 A flowchart of the multimodal scientific experiment evaluation assistance method of the present invention;
[0054] Figure 2 A block diagram of the multimodal scientific experiment evaluation assistance method of the present invention;
[0055] Figure 3 A network structure diagram of a posture assessment model of the multimodal scientific experiment assessment auxiliary method of the present invention;
[0056] Figure 4 A flow chart of an experimental evaluation report for the multimodal scientific experiment evaluation assistance method of the present invention. DETAILED DESCRIPTION
[0057] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0058] like Figure 1 and Figure 2 As shown, the present invention provides a multimodal scientific experiment evaluation auxiliary method based on model coupling, comprising the following steps:
[0059] S1. Use acquisition equipment to collect videos of students' experimental operation processes, obtain scientific experiment video sequences, and upload the scientific experiment video sequences to a streaming media server.
[0060] In one embodiment of the present application, each experimental workbench is equipped with a camera, which is horizontally facing the students and has a height set to the average height of the students. When students conduct experiments, it is required that the experimental objects and experimental phenomena cannot be blocked by their bodies during operation.
[0061] S2. The streaming media server uses a posture assessment model and a target detection model to extract features from the scientific experiment video sequence and fuses them to obtain multi-dimensional features.
[0062] Specifically, step S2 specifically includes:
[0063] S21, extracting the student's body key point information from the scientific experiment video sequence according to a predetermined frequency using a posture assessment model;
[0064] S22, using the target detection model to identify and locate the position information of the experimental objects and equipment from the scientific experiment video sequence at set intervals; wherein the experimental objects include flames, liquids, etc., and the experimental equipment include test tubes, stirring rods, droppers, etc.;
[0065] S23. Information fusion is performed based on the students’ human body key point information, the experimental objects and the location information of the equipment to form multi-dimensional features.
[0066] The predetermined frequency and interval can be set according to actual needs, and this application does not impose any specific restrictions on this.
[0067] By integrating the human body key point information extracted by the posture assessment model and the experimental object and equipment position information identified by the target detection model, rich multi-dimensional features are formed, which significantly improves the ability to understand the experimental process, making the classification and evaluation of scientific experiment video sequences more accurate and comprehensive. It can accurately determine various operation types and screen high-confidence results by setting probability thresholds to form a reliable experimental step sequence. It can not only evaluate students' operating skills, but also detect drug dosage, experimental environment and experimental results, providing multi-angle and multi-level experimental evaluation, which greatly improves the comprehensiveness and accuracy of the evaluation.
[0068] In an embodiment of the present application, the predetermined frequency is 1 frame / s, and the interval is 1s / frame.
[0069] In one embodiment of the present application, the posture assessment model refers to the PAF technology of openpose, and replaces the feature extraction network VGG16 of the PAF technology with VAN to make the model more adaptable. At the same time, the model only needs to extract information of 12 key points of the human body: eyes, nose, mouth, neck, shoulders, elbows, and wrists.
[0070] like Figure 3 As shown, specifically, the feature extraction network includes downsampling, a visual attention neural network based on LKA, and a Norm layer, wherein the visual attention neural network based on LKA includes deep convolution, deep dilated convolution, and channel convolution;
[0071] The prediction key point network adopts the first multi-stage architecture, and each stage gradually refines the prediction of the key point position; each stage consists of a 3×3 convolutional layer, and each stage is optimized using the L2 loss function; the loss function F1 of the prediction key point network is:
[0072]
[0073] Among them, F1(Loss) represents the loss function of the prediction key point network, S represents the total number of stages in the first multi-stage architecture, p represents the key point of the human body, w(p) represents the weight of the key point p, gt(p) represents the true position of the key point p, and pv(p) represents the predicted position of the key point p;
[0074] The limb network adopts the second multi-stage architecture, where each stage consists of a 1×1 convolutional layer. The loss function F2 of the limb network is the same as the loss function F1 of the prediction keypoint network.
[0075] By replacing the VGG16 feature extraction network in OpenPose's PAF technology with VAN and introducing a visual attention neural network based on LKA, the adaptability and accuracy of the posture evaluation model are greatly improved. It can more accurately capture the students' subtle movements during the experiment, thereby better evaluating the accuracy and standardization of the experimental operations.
[0076] It can be understood that the posture evaluation model can include multiple prediction key point networks and limb networks, each prediction key point network and limb network is a stage, that is, the posture evaluation model includes multiple stages, and the input of each stage is the output of the previous stage, that is, the characteristics after the prediction key point network and the limb network are fused.
[0077] In one embodiment of the present application, the posture assessment model includes two stages, namely, a feature extraction network, a first predicted key point network, a first limb network, a second predicted key point network and a second limb network. The features after fusion of the feature extraction network, the first predicted key point network and the first limb network are used as input, and are processed respectively using the second predicted key point network and the second limb network, and feature fusion is performed with the results of the feature extraction network to obtain human body key point information.
[0078] In one embodiment of the present application, the visual attention neural network based on LKA introduces the self-attention mechanism in the Transformer framework, which includes deep convolution, deep dilated convolution and channel convolution. The weight formula of the self-attention mechanism is as follows:
[0079] attention=conv 1×1 (DWDConv(DWConv(F)))
[0080] Among them, F represents the input feature, conv 1×1 It represents channel convolution with a convolution kernel of 1×1, DWDConv represents depth-dilated convolution, DWConv represents depth-wise convolution, and attention represents the importance of each feature in the input feature F.
[0081] In one embodiment of the present application, the network structure of the target detection model includes a lightweight classification head, a downsampling layer and a rank-guided block, wherein the lightweight classification head includes a lightweight architecture of 2 depth-separable convolutions and 1 channel convolution; the downsampling layer includes point convolution and depth-wise convolution.
[0082] In one embodiment of the present application, classification and regression in the lightweight classification head simultaneously use the same head architecture.
[0083] In one embodiment of the present application, in the downsampling layer, point-wise convolution is used to adjust the channel dimension, and depth-wise convolution is used for spatial downsampling. The computational cost and parameter cost are as follows:
[0084]
[0085] N(param)=2c 2 +18c
[0086] Among them, h represents the height of the input feature map, w represents the width of the input feature map, c represents the number of input feature map channels, O represents the computational cost, and N represents the parameter cost.
[0087] By adopting a network structure including a lightweight classification head, a downsampling layer, and a rank-guided block, the computational complexity and number of parameters are significantly reduced while maintaining efficient target detection performance. It is able to quickly and accurately identify and locate experimental objects and equipment with limited computing resources, and achieve real-time evaluation.
[0088] S3. Classify and evaluate the scientific experiment video sequence according to the multi-dimensional features to obtain classification results and evaluation results.
[0089] Specifically, step S3 specifically includes:
[0090] S31, determining the operation type of each frame of the scientific experiment video sequence according to the multi-dimensional features; the operation type includes experimental steps, correct operation steps, abnormal operations, drug dosage detection, experimental environment detection and experimental results;
[0091] S32, screening high-confidence operation classification results through a set probability threshold;
[0092] S33, associating the screened operation classification results with the time series to form an experimental step sequence;
[0093] S34, based on the location information of the experimental objects and equipment, calculating the experimental usage in the scientific experiment video sequence and generating an environmental description;
[0094] S35. Extract and classify experimental phenomena from the location information of experimental objects and equipment.
[0095] In one embodiment of the present application, the student's human body key point information, the location information of the experimental objects and equipment are used to classify the operations of each frame in the scientific experiment video sequence. For each type of operation, only the data with the highest confidence level and greater than the probability threshold is retained, and various features and the corresponding time series are recorded to form the correct operation feature values and abnormal operation feature values of the scientific experiment.
[0096] In one embodiment of the present application, the probability threshold is 0.8.
[0097] S4. Obtain the scoring standard document for the experimental operation, and construct a vector database of the experimental scoring standard based on the scoring standard document for the experimental operation.
[0098] Specifically, step S4 specifically includes:
[0099] S41, collecting experimental operation scoring standard documents, extracting text content, preprocessing the text content, and segmenting the preprocessed text to obtain knowledge units;
[0100] S42, using the embedding model to vectorize the knowledge unit and store it in the vector database;
[0101] S43. Use the vector database to build a vector index structure and add metadata tags to each vectorized knowledge unit; the metadata tags include experimental operations, experimental steps, experimental phenomena, experimental conclusions, etc.
[0102] In one embodiment of the present application, the NLP tool is used to segment the content of the preprocessed text.
[0103] S5. Generate an experimental evaluation report using a multimodal large model based on a vector database of multidimensional features, classification results, evaluation results, and experimental scoring criteria.
[0104] like Figure 4 As shown, specifically, step S5 specifically includes:
[0105] Generate corresponding prompt information for the multidimensional feature using the experimental scoring criteria in the vector database of the experimental scoring criteria;
[0106] The prompt information is input into the multimodal large model to generate preliminary experimental evaluation content, and the experimental evaluation content is structured to generate a complete experimental evaluation report; the experimental evaluation content includes experimental operation scoring, experimental step scoring, experimental phenomenon scoring, experimental conclusion scoring and operation error prompts.
[0107] By constructing a vector database of experimental scoring standards and combining it with a multimodal large model, we can achieve automatic generation of high-quality, personalized experimental evaluation reports. This not only provides standardized scoring, but also gives targeted feedback and suggestions based on specific circumstances, greatly improving the depth and educational value of the evaluation.
[0108] Specifically, the step S5 is described by taking the acid-base neutralization titration experiment as an example:
[0109] By analyzing each frame of the science experiment video sequence, the student's operation actions are identified, such as "pick up the burette", "adjust the acid height", "add acid", etc. If at the 10th second, the confidence level of the operation "add acid" is 0.85, which exceeds the threshold of 0.8, then this operation is recorded, along with the relevant key point positions (such as wrist position) and target positions (such as burette and beaker positions);
[0110] Combine the identified operation sequence (such as "pick up the burette" → "adjust the acid solution height" → "add acid solution") with the time information and generate a text description by comparing it with the experimental operation manual: "10 seconds after the experiment started, the student correctly performed the operation of adding acid solution, which complies with the standard experimental steps."
[0111] Through target detection, the change in the liquid level in the burette is identified. If the initial height is 20 mL and it is 5 mL after the titration, the amount of acid used is calculated to be 15 mL and recorded as the experimental dosage characteristic value;
[0112] The laboratory table is detected to have beakers, burettes, pH test papers and other items, and their placement is identified, based on which the environment description is generated: "The laboratory table is clean and tidy, and the necessary experimental equipment is placed reasonably and in accordance with safety operating specifications. The student wears protective glasses but not a lab coat."
[0113] Continuously monitor changes in phenomena during the experiment, such as identifying changes in the color of the solution and generating a description: "During the titration process, the solution gradually changed from colorless to pink. Finally, the solution showed a stable light pink color, indicating that the titration endpoint was reached."
[0114] S6. Use the experimental evaluation report to iteratively optimize the training of the multimodal large model.
[0115] In one embodiment of the present application, based on the large language model (LLM) and the Langchain framework, an experimental report generation system based on a vector database is implemented to generate a large number of scientific experiment evaluation data sets, and a high-quality scientific experiment evaluation training set is obtained after manual error correction.
[0116] Wherein, step S6 specifically includes:
[0117] Use the multi-dimensional features as the input of the multi-modal large model, match the corresponding experimental items in the vector database, and generate corresponding prompts;
[0118] Use a multimodal large model to conduct experimental evaluation from multiple dimensions, and use manual annotation adjustment to adjust the generated experimental evaluation to generate corresponding scientific experimental reports and evaluation reports;
[0119] Based on the pre-trained multimodal large model, the P-Tuning large model fine-tuning method is adopted to set hyperparameters such as learning rate, epchos, batch size, and iteratively optimize and train the multimodal large model.
[0120] In the embodiment of the present application, P-Tuning is used to fine-tune the model, so that the system can continuously learn and adapt to new experimental types and evaluation criteria, ensuring the long-term effectiveness and adaptability of the evaluation system.
[0121] This application realizes comprehensive automated evaluation of scientific experiments by combining video acquisition, feature extraction and classification of scientific experiment video sequences. It can not only objectively analyze students' experimental operations, but also generate detailed evaluation reports, greatly improving the efficiency and quality of experimental teaching. This multimodal scientific experiment evaluation auxiliary method significantly reduces human bias, provides more comprehensive and accurate experimental evaluation results, and continuously improves the accuracy of evaluation through continuous learning.
[0122] In one embodiment of the present application, taking the experiment of "Removal of Insoluble Impurities in Crude Salt" as an example, a multimodal scientific experiment evaluation auxiliary method based on model coupling is specifically implemented:
[0123] S1: A camera is installed on the operating table in the laboratory, and the height is set to 170cm (assuming that 170cm is the average height of students). When student a is conducting the experiment, the camera is facing student a horizontally to capture the entire experimental process;
[0124] S2: After the experimental video is uploaded to the streaming server, the streaming server starts feature extraction:
[0125] S21: The posture evaluation model extracts 12 key points (eyes, nose, mouth, neck, shoulders, elbows, wrists) of student a from the experimental video every second.
[0126] S22: The target detection model identifies the location of experimental equipment (such as beakers, filter paper, glass rods) and experimental objects (such as crude salt solution) every second.
[0127] S23: The streaming media server combines the posture information of student a with the positions of the experimental equipment and the experimental objects to form a multi-dimensional feature.
[0128] S3: Classification and evaluation of experimental videos based on multi-dimensional features:
[0129] S31: Perform time series analysis on multidimensional features to identify the start and end time points of the experimental operation;
[0130] S32: Classify experimental operations based on time series characteristics; such as measuring distilled water, dissolving crude salt, filtering, etc.;
[0131] S33: Correct and complete the classification results based on the standard sequence of experimental operations;
[0132] S34: The amount of distilled water detected is 12 ml (2 ml more than the standard);
[0133] S35: The filtered solution was observed to be clear and transparent.
[0134] S4: Obtain the scoring standard document for the experiment "Removal of insoluble impurities in crude salt" and build a vector database:
[0135] S41: Collect the scoring standard document of the experiment "Removal of Insoluble Impurities in Crude Salt", extract the plain text content from the document, remove the format information, preprocess the text content (such as removing extra spaces and line breaks, unifying punctuation, correcting possible spelling errors, etc.), and use nlp tools (such as NLTK or spaCy) to segment the preprocessed text content into knowledge units; for example:
[0136] "Measure 10 ml of distilled water (5 points)"
[0137] "Add appropriate amount of coarse salt and stir to dissolve (5 minutes)"
[0138] "Fold the filter paper into four pieces and then unfold it into a funnel shape (5 points)"
[0139] "Use a rubber-tipped dropper to absorb an appropriate amount of water to moisten the filter paper and stick it close to the inner wall of the funnel (5 points)"
[0140] "Assembling the filtration device (20 points)"
[0141] "Put the lower end of the glass rod close to the three layers of filter paper and slowly pour the turbid salt water along the glass rod (25 points)"
[0142] "Experimental supplies arrangement (10 points)"
[0143] "There are filtered impurities on the filter paper, and the clear, transparent filtrate flows into the beaker (20 minutes)".
[0144] S42: Use the embedding model to vectorize the knowledge unit and store it in the vector database FAISS; for example:
[0145] "Measure 10 ml of distilled water (5 minutes)"->[0.1, -0.3, 0.5, ..., 0.2]
[0146] "Add appropriate amount of coarse salt and stir to dissolve (5 minutes)"->[-0.2, 0.4, 0.1, ..., -0.3]
[0147] S43: Use the vector database FAISS to build a vector index structure and add metadata tags; for example:
[0148] "Measure 10 ml of distilled water (5 points)"->
[0149] {
[0150] "Operation Type": "Experimental Step",
[0151] "Step number": 1,
[0152] "score": 5,
[0153] "Keywords": ["measurement", "distilled water"]
[0154] }
[0155] "There are impurities on the filter paper after filtering, and the clear, transparent filtrate flows into the beaker (20 points)"->
[0156] {
[0157] "Operation Type": "Experimental Phenomenon",
[0158] "score": 20,
[0159] "Keywords": ["filter paper", "impurities", "clarification", "transparency", "filtrate"]
[0160] }.
[0161] S5: Generate experimental evaluation report using multimodal large model:
[0162] S51: The system retrieves relevant scoring criteria from the vector database based on the multi-dimensional features.
[0163] S52: The system generates preliminary assessment content, such as:
[0164] "Experimental step score: 75 / 80 points. Reason for deduction: The glass rod was not close to the filter paper during filtration (-5 points)." "Experimental phenomenon score: 15 / 20 points. Reason for deduction: Failure to confirm whether there are impurities on the filter paper (-5 points).
[0165] "
[0166] "Operation error tips: 1. Make sure the bottom of the glass rod is close to the filter paper when filtering. 2. Wash your hands after the experiment.
[0167] "
[0168] S6: Use this evaluation report to optimize the multimodal large model and improve the accuracy of future evaluations.
[0169] The present invention also provides a multimodal scientific experiment evaluation auxiliary system based on model coupling, which adopts the multimodal scientific experiment evaluation auxiliary method as described above, including:
[0170] The video acquisition module is configured to use an acquisition device to collect video of the student's experimental operation process, obtain a scientific experiment video sequence, and upload the scientific experiment video sequence to a streaming media server;
[0171] The feature extraction module is configured as a streaming media server to extract features from the scientific experiment video sequence using a posture assessment model and a target detection model and perform fusion to obtain a multi-dimensional feature
[0172] An operation classification module is configured to classify and evaluate the scientific experiment video sequence according to the multi-dimensional features to obtain classification results and evaluation results;
[0173] A knowledge base generation module is configured to obtain a scoring standard document for experimental operations and construct a vector database of experimental scoring standards based on the scoring standard document for experimental operations;
[0174] A report generation module configured to generate an experimental evaluation report using a multimodal large model based on a vector database of multidimensional features, classification results, evaluation results, and experimental scoring criteria;
[0175] The model training module is configured to use the experimental evaluation report to iteratively optimize the multimodal large model.
[0176] The present invention also provides a multimodal scientific experiment evaluation auxiliary device, comprising: at least one processor, at least one memory, a communication interface and a bus; wherein the processor, memory, and communication interface communicate with each other through the bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to implement the multimodal scientific experiment evaluation auxiliary method as described above.
[0177] The present invention also provides a computer-readable storage medium, which stores computer instructions, and the computer instructions enable the computer to implement the multimodal scientific experiment evaluation auxiliary method as described above.
[0178] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A multimodal scientific experiment evaluation auxiliary method based on model coupling, characterized in that: The following steps are involved: S1. Use a collection device to collect videos of students' experimental operation processes, obtain scientific experiment video sequences, and upload the scientific experiment video sequences to a streaming media server; S2, the streaming media server uses the posture assessment model and the target detection model to extract features from the scientific experiment video sequence and fuses them to obtain multi-dimensional features; S3, classifying and evaluating the scientific experiment video sequence according to the multidimensional features to obtain classification results and evaluation results; S4, obtaining the scoring standard document of the experimental operation, and constructing a vector database of the experimental scoring standard according to the scoring standard document of the experimental operation; S5. Generate an experimental evaluation report using a multimodal large model based on a vector database of multidimensional features, classification results, evaluation results, and experimental scoring criteria; S6. Use the experimental evaluation report to iteratively optimize the multimodal large model; Step S2 specifically includes: S21, extracting the student's body key point information from the scientific experiment video sequence according to a predetermined frequency using a posture assessment model; S22, using the target detection model to identify and locate the position information of the experimental objects and equipment from the scientific experiment video sequence at set intervals; S23, integrating information based on the students' human body key point information, experimental objects and equipment location information to form multi-dimensional features; The feature extraction network includes downsampling, LKA-based visual attention neural network and Norm layer, where the LKA-based visual attention neural network includes deep convolution, deep dilated convolution and channel convolution; The prediction key point network adopts the first multi-stage architecture, and each stage gradually refines the prediction of the key point position; each stage consists of a 3×3 convolutional layer, and each stage is optimized using the L2 loss function; the loss function of the prediction key point network is: Among them, F1(Loss) represents the loss function of the prediction key point network, S represents the total number of stages in the first multi-stage architecture, p represents the key point of the human body, w(p) represents the weight of the key point p, gt(p) represents the true position of the key point p, and pv(p) represents the predicted position of the key point p; The limb network adopts the second multi-stage architecture, where each stage consists of a 1×1 convolutional layer; The network structure of the target detection model includes a lightweight classification head, a downsampling layer and a rank guidance block, wherein the lightweight classification head includes a lightweight architecture of 2 depth-separable convolutions and 1 channel convolution; the downsampling layer includes point convolution and depth convolution; In the downsampling layer, point-wise convolution is used to adjust the channel dimension, and depth-wise convolution is used for spatial downsampling. The computational cost and parameter cost are as follows: N(param)=2c 2 +18c Among them, h represents the height of the input feature map, w represents the width of the input feature map, c represents the number of input feature map channels, O represents the computational cost, and N represents the parameter cost.
2. A multimodal scientific experiment evaluation auxiliary method based on model coupling as claimed in claim 1, characterized in that: Step S5 specifically includes: Generate corresponding prompt information for the multidimensional feature using the experimental scoring criteria in the vector database of the experimental scoring criteria; The prompt information is input into the multimodal large model to generate preliminary experimental evaluation content, and the experimental evaluation content is structured to generate a complete experimental evaluation report; the experimental evaluation content includes experimental operation scoring, experimental step scoring, experimental phenomenon scoring, experimental conclusion scoring and operation error prompts.
3. A multimodal scientific experiment evaluation auxiliary method based on model coupling as claimed in claim 1, characterized in that: Step S3 specifically includes: S31, determining the operation type of each frame of the scientific experiment video sequence according to the multi-dimensional features; the operation types include experimental step arrangement, correct operation step capture, abnormal operation capture, drug dosage detection, experimental environment detection and experimental result capture; S32, screening high-confidence operation classification results through a set probability threshold; S33, associating the screened operation classification results with the time series to form an experimental step sequence; S34, based on the location information of the experimental objects and equipment, calculating the experimental usage in the scientific experiment video sequence and generating an environmental description; S35. Extract and classify experimental phenomena from the location information of experimental objects and equipment.
4. A multimodal scientific experiment evaluation auxiliary system based on model coupling, characterized in that: The multimodal scientific experiment evaluation auxiliary method according to any one of claims 1 to 3 is adopted, comprising: The video acquisition module is configured to use an acquisition device to collect video of the student's experimental operation process, obtain a scientific experiment video sequence, and upload the scientific experiment video sequence to a streaming media server; A feature extraction module, which is configured as a streaming media server to extract features from the scientific experiment video sequence using a posture assessment model and a target detection model respectively and fuse them to obtain multi-dimensional features; An operation classification module is configured to classify and evaluate the scientific experiment video sequence according to the multi-dimensional features to obtain classification results and evaluation results; A knowledge base generation module is configured to obtain a scoring standard document for experimental operations and construct a vector database of experimental scoring standards based on the scoring standard document for experimental operations; A report generation module configured to generate an experimental evaluation report using a multimodal large model based on a vector database of multidimensional features, classification results, evaluation results, and experimental scoring criteria; The model training module is configured to use the experimental evaluation report to iteratively optimize the multimodal large model.
5. A multimodal scientific experiment evaluation auxiliary device, characterized in that: include: At least one processor, at least one memory, a communication interface and a bus; wherein, The processor, memory, and communication interface communicate with each other via the bus; The memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to implement the multimodal scientific experiment evaluation auxiliary method as described in any one of claims 1-3.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to implement the multimodal scientific experiment evaluation assistance method as described in any one of claims 1-3.
Citation Information
Patent Citations
Middle school experiment scoring method and device, computing equipment and storage medium
CN117037016A
Ultrasonic image breast tumor classification method based on feature fusion and attention mechanism
CN117746119A