Scratch virtual teaching demonstration method based on large language model
Through the virtual teaching demonstration method based on the large language model, the problems of unbalanced educational resources and inefficiency in the Scratch teaching model are solved, personalized teaching and efficient learning are achieved, teaching costs are reduced, and programming education is promoted.
Patent Information
- Application Number
- CN202510182278.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The existing Scratch teaching model has problems such as unbalanced educational resources, inefficiency, and difficulty in dynamic response to user needs, making it difficult to provide personalized teaching content and guidance.
Using a virtual teaching demonstration method based on a large language model, a large language model is trained to format user learning intentions, predict interactive actions, perform results state evaluation and generate logical code. Combined with OpenCV and mouse and keyboard simulation software packages, it dynamically responds to user needs, builds component state trees and interacts.
It has achieved personalized teaching content and guidance based on each student's learning intention and progress, which has improved learning efficiency and interest, reduced teaching costs, and promoted the popularization of programming education.
Smart Images

Figure CN120124664A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of virtual teaching, and particularly relates to a Scratch virtual teaching demonstration method based on a large language model. Background Art
[0002] With the continuous development of the education field, the traditional teacher-led teaching method has gradually revealed deficiencies in terms of efficiency and personalization. Teachers have limited time and energy and it is difficult to provide personalized tutoring for each student's learning progress and comprehension ability. Especially in some programming learning platforms, students often need personalized guidance when understanding complex concepts and operations. To address this challenge, educational technology is seeking automated teaching methods to replace the role of traditional teachers.
[0003] As a graphical programming platform for children and beginners, Scratch has become an important tool for promoting programming education worldwide with its intuitive interface and interactivity. However, the traditional Scratch teaching mode relies on on-site teacher explanations and students' independent operations, resulting in unbalanced educational resources, low efficiency, and difficulty in dynamically responding to user needs in Scratch teaching. Therefore, how to combine virtual teaching methods with the Scratch platform to replace the role of teachers has become an important research direction for improving educational efficiency and quality. This automated teaching method can facilitate knowledge transfer. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a Scratch virtual teaching demonstration method based on a large language model that integrates open source resources, reduces costs, and dynamically responds to user needs.
[0005] The technical solution adopted to solve the above technical problem is: A Scratch virtual teaching demonstration method based on a large language model, comprising the following steps:
[0006] Step 1. Train the large language model
[0007] Select an open source large language model and train it through a predefined data set to enable it to have the functions of formatting the user's learning intention, predicting interactive actions and targets, evaluating the execution result status, and generating logical codes for the component status tree;
[0008] The formatting of the user's learning intention is to process the learning intention text input by the user into a specific format with clear key data fields;
[0009] The prediction of interactive actions and targets is to predict the next interactive action and target component based on the current component status tree, user intention, and interaction history, and generate a text explanation of the prediction result;
[0010] The execution result status evaluation determines whether the current interaction has achieved the final learning intention based on user feedback;
[0011] The component state tree logic code generation converts the state subtree of the code editing area in the component state tree into a logical code text sequence;
[0012] Step 2. Determine the user's learning intention
[0013] Format the learning intention text input by the user through a large language model to ensure that necessary key data fields are included;
[0014] Step 3. Build a component state tree
[0015] The system obtains the window handle of the Scratch desktop application by calling the Win32 API interface, and then obtains the position and size of the visible content area of the window, ensures that the window is in the front and in the focused state, uses the OpenCV software package to intercept the image of the visible content area of the window and processes the intercepted image, divides the functional area, and extracts the size, position, and text information of the components in each area to build a component state tree;
[0016] Step 4. Prediction by the large language model
[0017] The trained large language model predicts the next interaction action, target component, and generates a corresponding text explanation based on the formatted user learning intention, component state tree, and interaction history;
[0018] Step 5. Execute the interaction action
[0019] The system calls the mouse and keyboard simulation software package according to the result predicted in Step 4 to execute the predicted interaction action;
[0020] Step 6. Determine the differences between component state trees
[0021] After the interaction action is completed, re-execute Step 3 to obtain a new component state tree, calculate the differences between the new and old component state trees to evaluate the effect of the interaction action;
[0022] Step 7. Result evaluation, loop interaction, and state storage
[0023] Judge whether the final learning intention has been achieved according to the evaluation result. If not, continue to loop and execute Step 4, Step 5, and Step 6, and store the result of each interaction as an interaction history in the form of a quadruple for subsequent analysis and adjustment of teaching strategies.
[0024] Preferably, in step 3, the functional areas include a building block type switching area, a building block selection area, a code editing area, a stage display area, a character editing area, and a stage switching area.
[0025] Preferably, the method for extracting the size, position, and text information of components in each area and constructing a component status tree is as follows:
[0026] Convert the color image of each area component into a grayscale image, and then convert the grayscale image into a binary image to facilitate identifying the component shape. Remove the noise in the binary image to improve the clarity of the component contour. Use an edge detection algorithm to detect the edges in the image, use a contour detection algorithm to detect the contours in the image, extract the bounding rectangle of the component to obtain the position and size of the component, filter according to the size of the component contour area, remove the noise that does not meet the requirements, use a shape matching algorithm to identify the shape of the component, and confirm the type of the component;
[0027] Apply an optical character recognition method to extract text information from the component image to obtain the type, function, and status value of the component;
[0028] Organize the extracted component information in a tree structure to form a component status tree, where each node represents a component and contains the size, position, and text information of the component.
[0029] Preferably, the method for calculating the difference between the old and new component status trees is one of a recursive comparison method based on depth-first or breadth-first search, a hash value comparison method, and a tree edit distance method.
[0030] Preferably, the method for step 7 to determine whether the final learning intention is achieved according to the evaluation result is as follows:
[0031] The system prompts the user to manually execute the code for the current building block construction. After the user executes the code, the system obtains the difference between the code execution result and the expectation through a feedback mechanism;
[0032] Set the prompting task of the large language model to the "evaluation mode". The large language model determines whether the current interaction has achieved the final learning intention according to the user's feedback result. If the user's feedback indicates that the code execution result is consistent with the expectation, the large language model outputs "yes", indicating that the learning intention has been achieved; otherwise, it outputs "no", indicating that the learning intention has not been achieved.
[0033] Preferably, the method for step 7 to determine whether the final learning intention is achieved according to the evaluation result is as follows:
[0034] Set the prompting task of the large language model to the "generation mode". The large language model converts the user's learning intention into corresponding building block code, extracts the status subtree of the code editing area from the component status tree, and further obtains the building block code;
[0035] Use the code comparison method to analyze whether the current interaction meets the preset learning objectives and the final learning intention;
[0036] The code comparison method is the comparison method of Token sets, or the embedding representation comparison method based on a deep learning model, or the similarity comparison method based on a syntax tree.
[0037] The beneficial effects of the present invention are as follows:
[0038] The present invention can provide personalized teaching content and guidance according to the learning intentions and progress of each student, meeting the learning needs of different students. By simulating real interaction actions, the system can interact with students, enhancing students' learning interest and participation. The present invention is not restricted by time and space, and students can learn at any time and any place, improving the flexibility of learning. Through the automated teaching process, the system can quickly respond to students' learning needs, provide timely feedback and guidance, and improve teaching efficiency. At the same time, schools can reduce the demand for professional programming teachers, reducing teaching costs.
[0039] Through scientific evaluation methods and rigorous teaching steps, the present invention enables the system to accurately evaluate students' learning effects, helping students better master programming knowledge and skills, lowering the threshold of programming learning, and promoting the popularization of programming education.
[0040] The present invention promotes high-quality teaching resources to more regions through open-source large language models and the Scratch platform, helping schools and students in education resource-poor areas obtain high-quality programming education, narrowing the education gap between different regions, bringing new teaching models and methods to the education field, and promoting educational innovation. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a schematic diagram of the Scratch virtual teaching demonstration method of the present invention based on a large language model.
[0042] Figure 2 is a schematic diagram of the process of extracting components based on the OpenCV software package.
[0043] Figure 3 is a schematic diagram of the completed component status tree.
[0044] Figure 4 is a schematic diagram of the process of executing interaction actions on components. DETAILED DESCRIPTION OF THE INVENTION
[0045] The present invention will be further described in detail below with reference to the drawings and embodiments, but the present invention is not limited to the following embodiments.
[0046] Embodiment 1
[0047] This embodiment will be described by "designing a chasing game".
[0048] In Figure 1 a Scratch virtual teaching demonstration method based on a large language model according to this embodiment includes the following steps:
[0049] Step 1. Train the large language model
[0050] Select an open-source large language model for training to enable it to have the functions of formatting the user's learning intention, predicting interactive actions and targets, evaluating the execution result status, and generating the logical code of the component status tree;
[0051] Format the user's learning intention: Process the learning intention text input by the user into a specific format with clear key data fields. The key data fields are: project, role, logic;
[0052] Predict interactive actions and targets: Predict the next interactive action and target component based on the current component status tree, user intention, and interaction history, and generate a text explanation of the prediction result;
[0053] Evaluate the execution result status: Judge whether the current interaction reaches the final learning intention according to the user's feedback;
[0054] Generate the logical code of the component status tree: Convert the status subtree of the code editing area in the component status tree into a logical code text sequence;
[0055] This embodiment selects the GLM-4-9B-Chat large language model open-sourced by Zhipu AI and trains it in a multi-round dialogue manner through the LoRA dataset. Among them: Format the user's learning intention, and the Prompt task is set to Format; Predict interactive actions and targets, and the Prompt task is set to Predict; Evaluate the execution result status, and the Prompt task is set to Evaluate; Generate the logical code of the component status tree, and the Prompt task is set to Generate;
[0056] Step 2. Determine the user's learning intention
[0057] Format the learning intention text input by the user through the trained large language model to ensure that it contains the necessary key data fields;
[0058] For example, when the input text is "Design a program that makes a kitten chase a basketball. When the ball hits the wall or the kitten, it will bounce off. When the kitten hits the ball, it will make a sound. Use the arrow keys to move the position of the kitten", the specific formatted text output by the large language model is "Project: Chase Game; Characters: Kitten, Basketball; Logic: When the ball hits the wall or the kitten, it will bounce off. Use the arrow keys to move the kitten. When the kitten hits the basketball, it will make a sound".
[0059] Step 3. Build the component state tree
[0060] The system obtains the window handle of the Scratch desktop application by calling the Win32 API interface, and then obtains the position and size (x, y, w, h) of the visible content area of the window. (x, y) represents the coordinates of the upper left corner of this area, x = 0, y = 29, and w and h represent the pixel width and height of this area respectively, w = 2560, h = 1440, ensuring that the window is in the front and in the focused state;
[0061] Use the OpenCV package to capture the image of the visible content area of the window and process the captured image. Generate an image mask based on the background color of the Scratch application window. Apply the contour detection method based on this mask to obtain a list of all contours in the image for distinguishing different functional areas;
[0062] Exclude contours with an area less than 2×10 4 Sort them from left to right and from top to bottom to obtain 6 functional areas, namely the block type switching area, the block selection area, the code editing area, the stage display area, the character editing area, and the stage switching area;
[0063] Extract the size, position, and text information of the components in each functional area to build the component state tree, as shown in Figure 2 、 3 , specifically as follows:
[0064] Convert the color image of each area component into a grayscale image, and then convert the grayscale image into a binary image to facilitate identifying the component shape. Remove the noise in the binary image to improve the clarity of the component contour. Use the edge detection algorithm to detect the edges in the image, use the contour detection algorithm to detect the contours in the image, extract the bounding rectangle of the component to obtain the position and size of the component. Filter according to the size of the component contour area to remove the noise that does not meet the requirements. Use the shape matching algorithm to identify the shape of the component and confirm the type of the component;
[0065] Apply the optical character recognition method to extract text information from the component image to obtain the type, function, and status value of the component;
[0066] Organize the extracted component information in a tree structure to form the component state tree T 1, each node represents a component, including the size, position, and text information of the component.
[0067] Step 4. Prediction by the large language model
[0068] Based on the formatted user learning intention, component status tree, and interaction history, the trained large language model predicts the next interaction action, target component, and generates the corresponding text explanation. Among them, the interaction action is a finite set composed of mouse and keyboard events. In this embodiment, the elements in this set include: {click, drag, move, scroll wheel, input};
[0069] In this embodiment, the interaction history is empty during the first prediction. Taking the code logic "a cat will meow when it touches a basketball" as an example, the output result of the large language model is "Target component: Building block selection area - Sound - Play sound <meow>Wait until it finishes playing, Code Editor - C3 - If <encounter <basketball>>Then the interaction action: Drag; Text explanation: Now we have completed the conditional judgment code for the kitten touching the basketball. Next, let's write the specific content of the conditional code. Put the "Play sound" in the "Sound" block <meow>Wait for the broadcast to finish and drag the building block into the conditional judgment code block. In this way, when the kitten touches the basketball, it will make a sound.
[0070] Step 5. Execute the interaction action
[0071] The system calls the mouse and keyboard simulation software package according to the result predicted in Step 4 and executes the predicted interaction action;
[0072] In this embodiment, the target component types are: building block 1 and building block 2; the position information is: (85, 191, 212, 51) and (1020, 729, 264, 109) respectively; the center positions of building blocks 1 and 2 are (191, 217) and (1152, 783) respectively; the specific effect of the interaction action "drag" is: move the mouse to the position (191, 217), press the left mouse button to trigger the Press event, move the mouse to the position (1152, 783), and release the left mouse button to trigger the Release event. As Figure 4 shown.
[0073] Step 6. Determine the differences between the component state trees
[0074] After the interaction action is completed, re-execute Step 3 to obtain a new component state tree T 2 , calculate the differences between the new and old component state trees to evaluate the effect of the interaction action;
[0075] The method for calculating the differences between the new and old component state trees is one of the recursive comparison methods based on depth-first or breadth-first search, the hash value comparison method, and the tree edit distance method. In this embodiment, the tree edit distance method is adopted. The component state tree T 2 has one more building block code than the component state tree T 1 , that is, "play sound <meow>"Waiting for playback to finish", so the tree edit distance is 1, and the component state tree is T 1 Only need to add a node to the tree to reach the state of the component state tree T 2 of the state;
[0076] Step 7. Result evaluation, loop interaction, state storage
[0077] Judge whether the final learning intention is achieved according to the evaluation result. If not, continue to loop and execute Step 4, Step 5, and Step 6, and store the results of each interaction in the form of a quadruple as the interaction history for subsequent analysis and adjustment of teaching strategies;
[0078] In this embodiment, the current state is "The logic that the kitten makes a sound when it touches the ball and moves the position of the kitten through the arrow keys has been completed, but the logic that the ball bounces off the wall and the kitten has not been completed yet".
[0079] 7.1 Result evaluation
[0080] Judge whether the current interaction has achieved the user's final learning intention. The system will prompt the user to manually execute the code of the current block building. The system obtains the difference between the code execution result and the expectation through the feedback mechanism, sets the Prompt task of the large language model to Evaluate, inputs the user's feedback text, and the large language model judges whether the current interaction has achieved the final learning intention according to the feedback result.
[0081] Among them, the user feedback is "The ball does not bounce off the wall and the kitten", and the large language model outputs the result "No", indicating that the user's final learning intention has not been achieved yet.
[0082] 7.2 Loop interaction
[0083] The result evaluation shows that the final learning intention has not been achieved. Go back to Step 3, capture the content of the application window again, build a new component state tree, and based on the new component state tree, re-execute Steps 4, 5, 6, and 7. This is a loop process until the user's final learning intention is met.
[0084] In the example, since the logic that the ball does not bounce off the wall and the kitten has not been completed, the system goes back to Step 3 to rebuild the component state tree and continues to execute the prediction task.
[0085] 7.3 State storage
[0086] Add the output results of each step of this round to the global interaction history list in the form of a quadruple to help the system analyze the user's operation sequence, learning progress, and difficulties, provide reference for subsequent teaching, and ensure that students can gradually achieve the final learning goal. The quadruple includes the component state tree, the target component, the interaction action, and the action result.
[0087] Example 2
[0088] In this embodiment, the method for judging whether the final learning intention is achieved according to the evaluation result in step 7 is as follows:
[0089] Set the Prompt task of the large language model to "Generate". The large language model converts the user's learning intention into corresponding block codes, extracts the state subtree of the code editing area from the component state tree, and further obtains the block codes;
[0090] In the example, the block codes generated by the large language model for the cat character are as follows:
[0091] # When the green flag is clicked
[0092] # Repeat execution
[0093] # If it touches (basketball)
[0094] # Play sound <meow>Waiting for playback to finish
[0095] #When the → key is pressed# Increase the x coordinate by 10
[0096] #When the ↑ key is pressed# Increase the y coordinate by 10
[0097] #When the ← key is pressed# Decrease the x coordinate by 10
[0098] #When the ↓ key is pressed# Decrease the y coordinate by 10
[0099] #End
[0100] Use the code comparison method to analyze whether the current interaction meets the preset learning goals and final learning intentions. Among them, the code comparison method is the comparison method of Token sets or the embedding representation comparison method based on deep learning models or the similarity comparison method based on syntax trees. In this embodiment, the embedding representation comparison method based on deep learning models is used, and the output result is: {No}, add the missing "Play sound" in the block code <meow>Wait for the "play to the end" part to complete the function.
[0101] Other steps are the same as those in Embodiment 1.< / meow> < / meow> < / meow> < / meow> < / basketball> < / meow>
Claims
1. A Scratch virtual teaching demonstration method based on a large language model, characterized in that: The following steps are involved: Step 1. Train a large language model Select an open source large language model and train it with a predefined data set to enable it to have the functions of formatting user learning intentions, interactive actions and goal predictions, execution result status evaluation, and component state tree logic code generation; The formatting of user learning intention is to process the learning intention text input by the user into a specific format with clear key data fields; The interaction action and target prediction is to predict the next interaction action and target component based on the current component state tree, user intention and interaction history, and generate a textual explanation of the prediction result; The execution result status evaluation is to judge whether the current interaction achieves the final learning intention based on user feedback; The component state tree logic code generation is to convert the state subtree of the code editing area in the component state tree into a logic code text sequence; Step 2. Determine user learning intent The learning intention text entered by the user is formatted through a large language model to ensure that the necessary key data fields are included; Step 3. Build the component state tree The system obtains the window handle of the Scratch desktop application by calling the Win32 API interface, and then obtains the position and size of the visible content area of the window, ensuring that the window is in the front and in focus. The system uses the OpenCV software package to capture the image of the visible content area of the window and processes the captured image to divide the functional area, and extracts the size, position and text information of the components in each area to build a component status tree. Step 4. Large language model prediction The trained large language model predicts the next interaction action and target component based on the formatted user learning intention, component state tree, and interaction history, and generates corresponding text explanations; Step 5. Perform interactive actions According to the predicted result in step 4, the system calls the mouse and keyboard simulation software package to execute the predicted interactive action; Step 6. Identify differences between component state trees After the interaction is completed, re-execute step 3 to obtain a new component state tree, and calculate the difference between the new and old component state trees to evaluate the effect of the interaction; Step 7. Result evaluation, loop interaction, and state storage Determine whether the final learning intention is achieved based on the evaluation results. If not, continue to loop through steps 4, 5, and 6, and store the results of each interaction in the form of a four-tuple as an interaction history for subsequent analysis and adjustment of teaching strategies.
2. The Scratch virtual teaching demonstration method based on a large language model according to claim 1 is characterized in that: The functional areas in step 3 include a block type switching area, a block selection area, a code editing area, a stage display area, a character editing area, and a stage switching area.
3. The Scratch virtual teaching demonstration method based on a large language model according to claim 1 is characterized in that: The method of extracting the size, position and text information of the components in each area and constructing the component status tree is as follows: Convert the color image of each component in each area into a grayscale image, and then convert the grayscale image into a binary image to facilitate the identification of the component shape, remove the noise in the binary image, and improve the clarity of the component outline. Use the edge detection algorithm to detect the edge in the image, use the contour detection algorithm to detect the contour in the image, extract the circumscribed rectangle of the component, and obtain the position and size of the component. Filter according to the size of the component outline area to remove the noise that does not meet the requirements. Use the shape matching algorithm to identify the shape of the component and confirm the type of the component. Apply optical character recognition methods to extract text information from component images and obtain component type, function and status value; The extracted component information is organized in a tree structure to form a component status tree. Each node represents a component and contains the size, position, and text information of the component.
4. The Scratch virtual teaching demonstration method based on a large language model according to claim 1 is characterized in that: The method for calculating the difference between the new and old component state trees is: a recursive comparison method based on depth-first or breadth-first search, a hash value comparison method, or a tree edit distance method.
5. The Scratch virtual teaching demonstration method based on a large language model according to claim 1 is characterized in that: The method of step 7 to determine whether the final learning intention is achieved according to the evaluation results is: The system prompts the user to manually execute the code built by the current building blocks. After the user executes the code, the system obtains the difference between the code execution result and the expected result through the feedback mechanism; Set the prompt task of the large language model to "evaluation mode". The large language model determines whether the current interaction achieves the final learning intention based on the user's feedback. If the user's feedback indicates that the code execution result is consistent with expectations, the large language model outputs "yes", indicating that the learning intention is achieved; Otherwise, the output is "No", indicating that the learning intention has not been achieved.
6. The Scratch virtual teaching demonstration method based on a large language model according to claim 1 is characterized in that: The method of step 7 to determine whether the final learning intention is achieved according to the evaluation results is: Set the prompt task of the large language model to "generate mode". The large language model converts the user's learning intention into the corresponding building block code, extracts the state subtree of the code editing area from the component state tree, and further obtains the building block code; Use code comparison method to analyze whether the current interaction meets the preset learning objectives and ultimate learning intentions; The code comparison method is a Token set comparison method, an embedded representation comparison method based on a deep learning model, or a syntax tree-based similarity comparison method.
Citation Information
Patent Citations
Teenager algorithm code auxiliary learning system and method based on large language model
CN117235347A
Consistency detection method and device for decentralized application, medium and equipment
CN118689487A
Large model thinking chain evaluation method and system for formal theorem proof
CN118709791A
Web application content learning and interaction system based on large language model
CN118861210A
Gneral modeling method for design evaluation and performance analisysof ATM protocol
KR1019980023365A