Image identification tracking method based on interactive intelligent experiment teaching system
By integrating multimodal sensor data and dynamically embedding teaching knowledge graphs, the problem that a single sensor in the existing technology is susceptible to the environment is solved, precise tracking of students' operating behaviors and adaptive teaching intervention are realized, and teaching efficiency and effect are improved.
Patent Information
- Application Number
- CN202510369781.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
AI Technical Summary
The existing experimental teaching system only relies on a single sensor and is susceptible to environmental factors, resulting in loss of targets, making it difficult to accurately monitor students' operating steps in real time, making it difficult to correct errors in a timely manner.
By integrating multimodal sensor data, dynamically embedding the teaching knowledge graph, and building a personalized feedback mechanism, it realizes accurate tracking, compliance judgment and adaptive teaching intervention of students' operating behavior.
It realizes accurate tracking and compliance determination of students' operating behaviors, improves teaching efficiency, corrects wrong operations in a timely manner, and improves teaching effectiveness.
Smart Images

Figure CN120220071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent experimental teaching systems, and specifically, to an image recognition and tracking method based on an interactive intelligent experimental teaching system. Background Art
[0002] The Intelligent Tutoring System (ITS) is an important research field in educational technology. It plays an important role in helping learners acquire knowledge and skills without the guidance of a human tutor by means of artificial intelligence technology.
[0003] However, in the process of using the existing experimental teaching systems, only a single sensor is relied on to identify the operation behaviors of students, which is vulnerable to environmental factors and easily leads to target loss, and it is difficult to accurately monitor the operation steps of each student in real time, resulting in difficult timely correction of incorrect operations.
[0004] Therefore, providing an image recognition and tracking method based on an interactive intelligent experimental teaching system that realizes accurate tracking, compliance determination, and adaptive teaching intervention of students' operation behaviors, and improves teaching efficiency by fusing multi-modal sensor data, dynamically embedding teaching knowledge graphs, and constructing a personalized feedback mechanism during use is an urgent problem to be solved by the present invention. Summary of the Invention
[0005] Aiming at the above technical problems, the purpose of the present invention is to overcome the problem in the prior art that only a single sensor is relied on to identify the operation behaviors of students, which is vulnerable to environmental factors and easily leads to target loss, and it is difficult to accurately monitor the operation steps of each student in real time, resulting in difficult timely correction of incorrect operations. Thus, an image recognition and tracking method based on an interactive intelligent experimental teaching system that realizes accurate tracking, compliance determination, and adaptive teaching intervention of students' operation behaviors, and improves teaching efficiency by fusing multi-modal sensor data, dynamically embedding teaching knowledge graphs, and constructing a personalized feedback mechanism during use is provided.
[0006] To achieve the above purpose, the present invention provides an image recognition and tracking method based on an interactive intelligent experimental teaching system, and the method includes the following steps:
[0007] S101. Collect experimental scenario data through a spatio-temporal alignment multi-modal sensor array;
[0008] S102. Use a cross-modal spatio-temporal alignment network to achieve feature-level fusion of optical, thermodynamic, and electromagnetic wave data;
[0009] S103. Inject teaching semantic constraints, generate semantic constraint items based on a predefined teaching knowledge graph, and correct the target tracking trajectory;
[0010] S104. Dynamically select feedback strategies according to the student portrait model, including security blocking instructions and teaching prompt instructions.
[0011] Preferably, the method for realizing feature-level fusion of optical, thermodynamic, and electromagnetic wave data by using a cross-modal spatio-temporal alignment network includes the following steps:
[0012] S201. Cross-modal feature extraction: Input the calibrated multi-modal data stream, where the multi-modal data stream includes: RGB image I rgb , radar point cloud image p radar and heat map M thermal ;
[0013] S202. Generate a fused feature map F fusion through CTA-Net, and the calculation formula is:
[0014]
[0015] where W k is a learnable weight matrix, and ⊕ represents the channel concatenation operation.
[0016] Preferably, the method for injecting teaching semantic constraints includes the following steps:
[0017] S301. Load the pre-compiled teaching knowledge graph, where the teaching knowledge graph is in JSO format;
[0018] S302. Extract the SOP rule set R = {r1,..., r n} of the current experiment;
[0019] S303. Encode the rule set R into a state transition matrix P ∈ R n×n , and the matrix element P ij represents the allowable probability from the operation step S i to S j ;
[0020] S304. Add a semantic constraint term to the loss function of the tracking algorithm as:
[0021]
[0022] Preferably, the matrix element P ij satisfies the following conditions:
[0023]
[0024] Preferably, the student portrait model is:
[0025] M student = {E skill ,Erisk , H history};
[0026] Among them, E skill : Skill level; E risk : Risk tolerance; H history : Historical operations.
[0027] Preferably, based on Q-Learning to generate the feedback action a t , the dynamic selection feedback strategy function can be obtained as:
[0028]
[0029] Among them, R immediate is the immediate reward, such as +1 for correct operation, -5 for dangerous actions, and γ = 0.9 is the discount factor.
[0030] Preferably, the triggering conditions of the dynamic selection feedback strategy include:
[0031] When a high-risk operation is detected, immediately activate the safety block;
[0032] When the operation deviates from the SOP but there is no direct danger, push an error operation warning and correction suggestions through voice prompts, AR projections or the teaching management system.
[0033] Preferably, the multi-modal sensor includes: an infrared sensor, a radar, and an RGB camera.
[0034] According to the above technical solution, the beneficial effects of the cleaning device for machining cylinder head castings provided by the present invention during use are:
[0035] By integrating multi-modal sensor data, dynamically embedding teaching knowledge graphs, and constructing a personalized feedback mechanism, the present invention realizes precise tracking, compliance determination, and adaptive teaching intervention for students' operation behaviors, improving teaching efficiency.
[0036] Other features and advantages of the present invention will be described in detail in the subsequent specific implementation part; and the parts not involved in the present invention are the same as or can adopt the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification, and are used to explain the present invention together with the following specific implementation manners, but do not constitute a limitation to the present invention. In the drawings:
[0038] Figure 1 is a workflow block diagram of an image recognition and tracking method based on an interactive intelligent experimental teaching system provided in a preferred embodiment of the present invention. Detailed implementation manners
[0039] The following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0040] As Figure 1 shown, the image recognition and tracking method based on the interactive intelligent experimental teaching system provided by the present invention includes the following steps:
[0041] S101. Collect experimental scenario data through a spatio-temporal alignment multi-modal sensor array;
[0042] S102. Use a cross-modal spatio-temporal alignment network to achieve feature-level fusion of optical, thermodynamic, and electromagnetic wave data;
[0043] S103. Inject teaching semantic constraints, generate semantic constraint items based on a predefined teaching knowledge graph, and correct the target tracking trajectory;
[0044] S104. Dynamically select feedback strategies according to the student portrait model, including safety blocking instructions and teaching prompt instructions.
[0045] In the above solution, the synchronization accuracy of the multi-modal sensor array is controlled within 0.5 ms by the IEEE 1588v2 protocol to ensure that the time stamp deviation of the multi-modal data < 0.5 ms.
[0046] This method is applied to the tracking of biological experiment micromanipulation. More specifically, it is used to track the operation of pipetting cell culture solution under a microscope:
[0047] Monitor the temperature change of the pipette through a multi-modal sensor, identify the operation start signal, and fuse the data of the microscopic camera and the radar through a cross-modal spatio-temporal alignment network (CTA-Net) to track the sub-millimeter displacement of the tip of the pipette. When it is detected that "the tip of the pipette touches different culture solutions without changing the pipette tip", a safety block is triggered.
[0048] In summary, this method realizes precise tracking, compliance determination, and adaptive teaching intervention of students' operation behaviors by fusing multi-modal sensor data, dynamically embedding a teaching knowledge graph, and constructing a personalized feedback mechanism, thereby improving the teaching efficiency.
[0049] In a preferred implementation manner of the present invention, the multi-modal sensor includes: an infrared sensor, a radar, and an RGB camera.
[0050] In the above solution, the infrared sensor can select a non-cooled infrared thermal imager (FLIR A700, temperature measurement accuracy ±2°C), the radar can select an FMCW millimeter-wave radar (TI AWR1843, 4D point cloud output), and the RGB camera can select a polarized RGB-D camera (Hikvision MV-CH250-90GM, supporting 2560×1440@90fps).
[0051] In a preferred embodiment of the present invention, the method for realizing feature-level fusion of optical, thermodynamic, and electromagnetic wave data by using a cross-modal spatio-temporal alignment network includes the following steps:
[0052] S201. Cross-modal feature extraction: Input the calibrated multi-modal data stream, and the multi-modal data stream includes: RGB image I rgb , radar point cloud image p radar and thermal map M thermal ;
[0053] S202. Generate a fused feature map F fusion through CTA-Net, and the calculation formula is:
[0054]
[0055] where W k is a learnable weight matrix, and ⊕ represents the channel splicing operation.
[0056] In the above solution, the cross-modal spatio-temporal alignment network (CAT-Net) is used to solve the resolution mismatch problem between the millimeter-wave radar point cloud (sparse) and the RGB image (dense), so as to accurately collect experimental scenario data.
[0057] In a preferred embodiment of the present invention, the method for injecting teaching semantic constraints includes the following steps:
[0058] S301. Load the pre-compiled teaching knowledge graph, and the teaching knowledge graph is in JSO format;
[0059] S302. Extract the SOP rule set R = {r1,..., r n} of the current experiment;
[0060] S303. Encode the rule set R into a state transition matrix P ∈ R n×n , and the matrix element P ij represents the allowable probability from the operation step S i to S j ;
[0061] S304. Add a semantic constraint term to the tracking algorithm loss function as:
[0062]
[0063] In the above solution, converting the experimental SOP into the probability transfer constraint of the tracking algorithm can reduce the error rate of experimental operations.
[0064] In a preferred embodiment of the present invention, the matrix element P ij satisfies the following conditions:
[0065]
[0066] In a preferred embodiment of the present invention, the student portrait model is:
[0067] M student = {E skill ,E risk ,H history};
[0068] Among them, E skill : skill level; E risk : risk tolerance; H history : historical operations.
[0069] In the above solution,.
[0070] In a preferred embodiment of the present invention, based on Q-Learning to generate the feedback action a t , the dynamic selection feedback strategy function can be obtained as:
[0071]
[0072] Among them, R immediate is the immediate reward, such as +1 for correct operation, -5 for dangerous actions, and γ = 0.9 is the discount factor.
[0073] In a preferred embodiment of the present invention, the triggering conditions of the dynamic selection feedback strategy include:
[0074] When a high-risk operation is detected, immediately initiate safety blocking;
[0075] When the operation deviates from the SOP but there is no direct danger, push error operation warnings and correction suggestions through voice prompts, AR projections, or the teaching management system.
[0076] In the above solution, when this method is applied to monitor the real-time distance between the student's hand and the robotic arm, specifically, when the student's arm enters the working area of the robotic arm, calculate the optimal feedback action based on the Q(s,a) decision function:
[0077] When the student is a novice, (E skill<0.3), immediately initiate safety blocking and pause the robotic arm;
[0078] When the student is a proficient operator, (E skill > 0.7), only issue a voice warning.
[0079] Therefore, by dynamically selecting the conditions for the feedback strategy to achieve high-risk operations, such as when contacting corrosive substances without wearing gloves, immediately initiate safety blocking so that students can well correct their experimental operation behaviors; and without the supervision of a teacher, it can automatically track whether the students' operation behaviors are appropriate, avoiding dangers caused by students' incorrect operations. Intelligent teaching is realized and the teaching efficiency is improved.
[0080] In summary, the cleaning device for machining cylinder head castings provided by the present invention overcomes the problems in the prior art that only a single sensor is relied on to identify students' operation behaviors, which is easily affected by environmental factors and prone to target loss, and it is difficult to accurately monitor the operation steps of each student in real time, resulting in the difficulty of timely correcting incorrect operations.
[0081] The preferred embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solutions of the present invention, and these simple modifications all fall within the protection scope of the present invention.
[0082] In addition, it should be noted that, among the various specific technical features described in the above specific embodiments, they can be combined in any appropriate manner without conflict. To avoid unnecessary repetition, the present invention will not separately describe various possible combination methods.
[0083] Furthermore, any combination can be made between various different embodiments of the present invention as long as it does not violate the idea of the present invention, and it should also be regarded as the content disclosed by the present invention.
Claims
1. An image recognition and tracking method based on an interactive intelligent experimental teaching system, characterized in that: The method comprises the following steps: S101, collecting experimental scene data through a spatiotemporally aligned multimodal sensor array; S102. Feature-level fusion of optical, thermodynamic and electromagnetic wave data using a cross-modal spatiotemporal alignment network; S103, injecting teaching semantic constraints, generating semantic constraint items based on a predefined teaching knowledge graph, and correcting the target tracking trajectory; S104. Dynamically select feedback strategies based on the student portrait model, including safety blocking instructions and teaching prompt instructions.
2. The image recognition and tracking method based on the interactive intelligent experimental teaching system according to claim 1 is characterized in that: The method for realizing feature-level fusion of optical, thermodynamic and electromagnetic wave data using a cross-modal spatiotemporal alignment network comprises the following steps: S201, cross-modal feature extraction: input the calibrated multi-modal data stream, the multi-modal data stream includes: RGB image I rgb , radar point cloud image p radar and heat map M thermal ; S202, generate fusion feature map F through CTA-Net fusion , the calculation formula is: Among them, W k is a learnable weight matrix, and ⊕ represents a channel concatenation operation.
3. The image recognition and tracking method based on the interactive intelligent experimental teaching system according to claim 1 is characterized in that: The method of injecting teaching semantic constraints comprises the following steps: S301, loading a precompiled teaching knowledge graph, wherein the teaching knowledge graph is in a JSO format; S302, extract the SOP rule set R of the current experiment = {r1, ..., r n }; S303, encode the rule set R into a state transfer matrix P∈R n×n , the matrix element P ij Indicates that from operation step S i To S j The allowed probability of S304, adding a semantic constraint term to the tracking algorithm loss function:
4. The image recognition and tracking method based on the interactive intelligent experimental teaching system according to claim 3 is characterized in that: The matrix element P ij The following conditions are met:
5. The image recognition and tracking method based on the interactive intelligent experimental teaching system according to claim 1 is characterized in that: The student portrait model is: M student ={E skill ,E risk ,H history }; Among them, E skill : Skill level; E risk : Risk tolerance; H history : Historical operations.
6. The image recognition and tracking method based on the interactive intelligent experimental teaching system according to claim 1 is characterized in that: Generate feedback action a based on Q-Learning t , it can be concluded that the dynamic selection feedback strategy function is: Among them, R immediate For timely rewards, such as +1 for correct operation and -5 for dangerous action, γ = 0.9 is the discount factor.
7. The image recognition and tracking method based on the interactive intelligent experimental teaching system according to claim 6, characterized in that: The triggering conditions for dynamically selecting the feedback strategy include: When high-risk operations are detected, security blocking is initiated immediately; When the operation deviates from the SOP but there is no direct danger, incorrect operation warnings and correction suggestions are pushed through voice prompts, AR projection or teaching management system.
8. The image recognition and tracking method based on the interactive intelligent experimental teaching system according to claim 1 is characterized in that: The multimodal sensor includes an infrared sensor, a radar and an RGB camera.