Handmade ceramic teaching system applying visual sensor and use method of handmade ceramic teaching system

Through visual sensors to collect multimodal data and combine data processing and interactive guidance, the problems of low efficiency and strong feedback subjectivity in traditional manual ceramic teaching are solved, real-time standardized evaluation and systematic storage of process knowledge are realized, and teaching efficiency and beginner experience are improved.

CN120340323AInactive Publication Date: 2025-07-18ZHEJIANG DONGDU CULTURAL CREATIVITY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510778918.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional handmade ceramics have low efficiency in teaching, strong subjectivity in feedback, difficult to systematically process knowledge, and lack objective quantitative evaluation of students' operational norms, resulting in high difficulty in getting started by novices and waste of teaching resources.

Method used

Vision sensors are used to collect multimodal data, combine data processing modules for feature extraction and process compliance analysis, real-time guidance is provided through AR display, voice prompts and tactile feedback, and optimize the process knowledge base in combination with machine learning to form a closed-loop teaching system.

Benefits of technology

It realizes objective quantitative evaluation of manual ceramic teaching, improves teaching efficiency, reduces learning difficulty, enhances the real-time nature of systematic storage and feedback of process knowledge, and reduces material waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340323A_ABST
    Figure CN120340323A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of ceramic teaching, in particular to a handmade ceramic teaching system applying a visual sensor and a use method thereof, and the system comprises a visual perception module, a data processing module, an interaction guidance module, a process knowledge base and terminal equipment integrating the above modules. The method is suitable for scenes such as ceramic technology teaching, ceramic art training and non-inheritance, and aims to improve understanding and operation normalization of learners on the handmade ceramic manufacturing process through intelligent perception and interaction technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ceramic teaching, and specifically to a manual ceramic teaching system applying a vision sensor and its usage method. Background Art

[0002] Manual ceramic production is an important field that combines traditional craftsmanship and artistic creation. Its teaching process usually relies on the experience transfer mode of "master leading apprentice". There are many technical bottlenecks in this traditional teaching mode: Firstly, the teaching efficiency is low. Since it requires one-on-one guidance from the master, limited by labor costs and time investment, it is difficult to achieve large-scale teaching. Secondly, the teaching feedback is highly subjective. The judgment of the standardization of trainees in key operations such as throwing, trimming, and modeling completely depends on the personal experience of the master, lacking objective quantitative standards, resulting in novice trainees often affecting the quality of their works due to slight movement deviations. Thirdly, it is difficult to effectively retain process knowledge. Key parameters such as the control of the thickness of the blank body and the matching of the rotation speed in ceramic production mainly rely on oral instruction, lacking systematic visual recording means. Finally, it is difficult for beginners to get started. Beginners are difficult to intuitively understand the dynamic relationship between hand strength, clay deformation, and the shape of the blank body, often needing to repeatedly try and error, resulting in material waste and an extended learning cycle. These problems seriously restrict the inheritance and development of manual ceramic skills. In view of the above problems, the existing technology urgently needs to be improved. Summary of the Invention

[0003] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a manual ceramic teaching system applying a vision sensor and its usage method.

[0004] To achieve the above purpose, the present invention provides the following technical solution: A manual ceramic teaching system applying a vision sensor, including a visual perception module, a data processing module, an interactive guidance module, a process knowledge base, and a terminal device integrating the above modules; The visual perception module is used to collect multi-modal visual data of the trainee's operation scene in real time, The multi-modal visual data includes RGB images, depth information, and infrared thermal distribution data; The data processing module is used to preprocess, extract features, recognize actions, and analyze process compliance of the multi-modal visual data, generating operation deviation data; The interactive guidance module is used to output guidance information through multi-modal feedback according to the operation deviation data; The process knowledge base is used to store standard ceramic production processes, key action specifications, and typical error cases, and support learning optimization based on operation data; The terminal device is used to integrate each module and provide a human-computer interaction interface, supporting the interaction between the trainee's operation and the system.

[0005] In some of these embodiments, the visual perception module includes: An RGB camera for collecting color images of the operation scene, with a resolution of not less than 1080P and a frame rate of not less than 30fps; A depth camera for collecting depth information of the operation scene, with an accuracy of not less than ±2cm; An infrared thermal imager for collecting the surface temperature distribution of the mud material, with a wavelength band range of 8 - 14μm and a temperature detection range of -20°C to 1200°C; A sensor calibration unit for synchronizing timestamps and calibrating spatial coordinates of the RGB camera, depth camera, and infrared thermal imager to ensure the spatio-temporal consistency of multi-modal data.

[0006] In some of these embodiments, the data processing module includes: A preprocessing unit for denoising, background segmentation, and region of interest extraction of multi-modal visual data. The background segmentation uses a U-Net network model; A feature extraction unit for extracting hand key point coordinates through an HRNet network and predicting the 3D hand pose through a 3D convolutional neural network; An action recognition unit for modeling the time series features of continuous actions through an LSTM network to identify the current operation step; A parameter matching unit for comparing action features with standard parameter thresholds in the process knowledge base and calculating the deviation value.

[0007] In some of these embodiments, the interaction guidance module includes: An AR display unit for overlaying standard action trajectory lines or key point reference diagrams in the real operation scene through AR glasses or a mobile phone AR engine; A voice interaction unit for generating natural language guidance information through TTS technology, including operation correction prompts or step guidance; A tactile feedback unit for outputting vibration reminders through wearable devices when the operation deviation exceeds the threshold, and the vibration mode is positively correlated with the deviation degree.

[0008] In some of these embodiments, the process knowledge base includes: A structured database storing standard process flows, key action specifications, and typical error cases; A machine learning optimization unit for updating standard parameter thresholds based on trainee operation data.

[0009] In some of these embodiments, the terminal device includes: A main control unit using a high-performance embedded computer for coordinating the operations of each module; A human-computer interaction interface for displaying real-time operation screens, guidance information, learning progress statistics, and historical operation playback; A power supply module using a rechargeable lithium battery with a battery life of no less than 8 hours.

[0010] In some embodiments, the system further includes an environmental perception unit for collecting temperature and humidity data of the operation environment and inputting the environmental parameters into the data processing module to correct the temperature detection error of the infrared thermal imager.

[0011] In some embodiments, when performing action recognition, the data processing module uses a multi-task learning model to simultaneously output the classification result of the operation steps and the action quality score, and the action quality score is used to quantify the operation standardization.

[0012] In some embodiments, when generating a learning report, the interactive guidance module counts the operation deviation frequency, average deviation value, and typical error types of each step of the trainee, and provides improvement suggestions based on the process knowledge base.

[0013] To achieve the above object, the present invention provides the following technical solution: A method for using a manual ceramic teaching system applying a vision sensor, the steps of which are as follows: (1) System initialization, collecting multi-modal data of the operation scenario through the vision perception module, and the sensor calibration unit completes spatio-temporal calibration; (2) The trainee starts to operate, and the data processing module collects and preprocesses multi-modal data in real time, and extracts the key features of the hand, clay, and tools; (3) Action recognition and compliance analysis, identifying the current operation step through 3DCNN, analyzing the time series features of continuous actions through LSTM, comparing the action features with the standard parameters of the process knowledge base, and calculating the deviation value; (4) Multi-modal feedback output, if the deviation value ≤ the tolerance threshold, the interactive guidance module outputs a positive prompt; if the deviation value > the tolerance threshold, corrective guidance is output through AR display, voice prompt, and tactile vibration; (5) Learning data recording and optimization, storing the operation data of this time into the process knowledge base, and the machine learning optimization unit updates the standard parameter threshold; (6) After completing all operation steps, a learning report is generated to display the scores of each step, common error types, and improvement suggestions.

[0014] Compared with the prior art, the beneficial effects of the present invention are: By collecting multi-modal data in real time through the vision perception module, combining the compliance analysis of the data processing module and the multi-modal feedback of the interactive guidance module, the problems of low efficiency and strong subjectivity of feedback in the traditional teaching mode are solved, and it has the advantages of improving teaching efficiency, providing objective and quantitative feedback, realizing systematic retention of process knowledge, and reducing the learning difficulty.

[0015] Details of one or more embodiments of the present application are set forth in the following drawings and description, so that other features, objects, and advantages of the present application will become more concise and understandable. The present application will be described in detail and understood through the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is the overall system architecture diagram of the present invention; Figure 2 It is the data processing flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] In the traditional existing manual ceramic teaching process, relying on the experience transmission mode leads to multi-dimensional technical defects. There is a lack of multi-modal data acquisition and fusion capabilities in the teaching scenario, and it is impossible to realize the real-time quantitative analysis of hand movements and mud deformation, resulting in the lack of an objective benchmark for evaluating operation norms. The key process parameters do not establish a structured knowledge system, resulting in the difficulty of forming a closed-loop optimization mechanism for error pattern recognition and guidance feedback. There is a lack of a cross-modal data collaborative processing architecture at the system level, and it is impossible to support the spatio-temporal consistency modeling in dynamic operation scenarios, restricting the standardization and replicability of the teaching process. For example, in the ceramic throwing process, there is a non-linear coupling relationship between the three-dimensional movement trajectory of the trainee's hand and the plastic deformation of the mud. The existing teaching system only relies on monocular vision to capture two-dimensional plane movements and cannot obtain synchronous data on the change rate of hand joint angles and the mud thickness distribution. When the trainee operates the turntable, the traditional method cannot quantitatively analyze the influence degree of the wrist pitch angle deviation on the ellipticity of the blank. In the data processing flow, no time-series correlation model is established between action feature extraction and process parameter matching, resulting in the inability to accurately model the dynamic relationship between the tool contact pressure and the rotation speed during the trimming stage. If the above problems are not solved, it will lead to the accumulation of systematic errors in the transmission of key process parameters, and the wrong operation mode cannot be iteratively optimized through data-driven methods. The dynamic feedback delay of hand-eye coordination in the teaching process will directly affect the quality stability of the blank forming, and the lack of multi-dimensional feature association in the process knowledge base will hinder the generation of personalized guidance strategies. The data processing bottleneck at the system level will limit the real-time concurrent processing ability in large-scale teaching scenarios and reduce the utilization efficiency of teaching resources.

[0019] When facing the above problems, this application first considers how to construct a multi-modal data acquisition and analysis system to break through the perceptual limitations of traditional teaching. Traditional methods rely on a single visual dimension, resulting in the lack of analysis of the correlation between actions and deformations. This application attempts to fuse RGB images, depth information, and infrared thermal distribution data, and establish a multi-dimensional mapping relationship between hand movements and the state of the clay through a spatio-temporal calibration mechanism. For the systematic error in the transmission of process parameters, this application explores the establishment of a structured knowledge base, matches the features of the standard process and dynamic operation data, and introduces a machine learning optimization mechanism to achieve adaptive adjustment of parameter thresholds. To solve the problem of feedback delay, this application designs a multi-modal interaction channel, combines augmented reality, voice prompts, and tactile feedback to form a closed-loop guidance system, ensuring the real-time and intuitive nature of operation correction. Through a modular architecture, the perception, computing, and interaction functions are integrated to finally form a systematic solution covering the entire process of data acquisition, analysis, and guidance.

[0020] In response, as Figures 1 to 2 shown, this application proposes a manual ceramic teaching system using a vision sensor, including: a visual perception module, a data processing module, an interaction guidance module, a process knowledge base, and a terminal device integrating the above modules; the visual perception module is used to collect multi-modal visual data of the student operation scene in real time, and the multi-modal visual data includes RGB images, depth information, and infrared thermal distribution data; the data processing module is used to preprocess, extract features, recognize actions, and analyze process compliance of the multi-modal visual data, and generate operation deviation data; the interaction guidance module is used to output guidance information through multi-modal feedback according to the operation deviation data; the process knowledge base is used to store the standard ceramic production process, key action specifications, and typical error cases, and support learning optimization based on operation data; the terminal device is used to integrate each module and provide a human-computer interaction interface, supporting student operations and system interactions.

[0021] Among them, the visual perception module's real-time collection of multi-modal visual data of the trainee's operation scenario refers to synchronously obtaining visual information of the operation scenario through multiple sensors. Specifically, it can be achieved by combining an RGB camera, a depth camera, and an infrared thermal imager to collect color images, spatial depth, and temperature distribution data respectively, for comprehensively capturing hand movements, the shape of the clay, and temperature changes, and solving the problem that dynamic relationships are difficult to intuitively understand in traditional teaching. Among them, the data processing module's preprocessing, feature extraction, action recognition, and process compliance analysis of multi-modal visual data refer to processing the original data through algorithms and extracting key features. Specifically, the U-Net network can be used for background segmentation, HRNet for extracting hand key points, 3DCNN for estimating three-dimensional postures, and LSTM for modeling action sequences, and combining with the standard parameters in the process knowledge base for deviation calculation, solving the problems of strong subjectivity in feedback and lack of quantitative standards. Among them, the interactive guidance module's output of guidance information through multi-modal feedback refers to transmitting corrective information by combining multiple sensory channels. Specifically, it can be achieved by using AR to overlay the standard action trajectory, TTS to generate voice prompts, and the tactile vibration of a smart bracelet, for guiding the trainee to adjust the operation in real time and reducing material waste caused by action deviations. Among them, the process knowledge base storing the standard ceramic production process, key action specifications, and typical error cases refers to converting process experience into structured data. Specifically, the step sequence, angle threshold, and error feature data can be stored in a database, and the parameters can be dynamically updated by combining with the machine learning optimization unit, solving the problems of difficult knowledge retention and low teaching efficiency. Among them, the terminal device integrating each module and providing a human-computer interaction interface refers to realizing system functions through hardware integration. Specifically, an embedded computer can be used to coordinate operations, a touch screen to display the operation screen and statistical information, supporting real-time interaction between the trainee and the system and improving the operability of teaching. The core innovation of this application lies in forming a closed-loop teaching system through the multi-modal data collection of the visual perception module, the intelligent analysis of the data processing module, the multi-modal feedback of the interactive guidance module, and the dynamic optimization of the process knowledge base, realizing the objective quantitative evaluation and real-time guidance of ceramic production actions, and solving the problems of low efficiency, subjective feedback, and difficult knowledge retention in traditional master-apprentice teaching.

[0022] The working process and principle of this application are as follows: The visual perception module collects multi-modal visual data of the trainee's operation scenario through an RGB camera, a depth camera, and an infrared thermal imager. This data includes RGB images, depth information, and infrared thermal distribution data, providing comprehensive scene information. The data processing module preprocesses, extracts features, recognizes actions, and analyzes process compliance for the collected multi-modal visual data. In the preprocessing stage, noise is removed and regions of interest are extracted. In the feature extraction stage, hand key points and postures are recognized. In the action recognition stage, the current operation step is recognized through temporal modeling. In the process compliance analysis stage, the extracted features are compared with the standard parameters in the process knowledge base to generate operation deviation data. The interactive guidance module outputs guidance information through multi-modal methods such as augmented reality display, voice prompts, and tactile feedback based on the operation deviation data. The process knowledge base stores standard production processes, action specifications, and error cases, and supports learning and optimization based on operation data to continuously improve the knowledge system. The terminal device integrates each functional module, provides a human-computer interaction interface, and supports the interaction between trainees' operations and the system. Each module works together to achieve real-time perception, analysis, and guidance of trainees' operations, forming a closed-loop teaching process.

[0023] As a preferred embodiment, the solution of this application is specifically implemented as follows: The visual perception module uses an RGB camera, a depth camera, and an infrared thermal imager. The RGB camera collects color images of the operation scenario. The depth camera collects three-dimensional depth information of the scenario. The infrared thermal imager collects the temperature distribution on the surface of the clay. The three sensors are time-synchronized and spatially calibrated through a calibration unit to ensure the consistency of multi-modal data. The data processing module first denoises and segments the background of the raw data. Then it extracts the coordinates and three-dimensional postures of hand key points. Next, it recognizes the current operation steps, such as throwing and trimming. Finally, it compares the action features with the standard parameters in the process knowledge base to calculate the deviation value. The interactive guidance module includes AR glasses, a voice broadcast device, and a smart bracelet. The AR glasses superimpose and display the standard action trajectory. The voice broadcast device outputs voice guidance. The smart bracelet outputs a vibration reminder when the operation deviation is large. The process knowledge base uses a relational database to store process flows, action specifications, and error cases. The machine learning unit updates the standard parameters based on historical data. The terminal device uses an embedded computer equipped with a touch display screen to display real-time operation images and guidance information.

[0024] Through the above solution, this application realizes the objective quantitative analysis of the manual ceramic teaching process. The multi-modal visual perception breaks through the limitations of the traditional single visual dimension and realizes the real-time correlation analysis of hand movements and clay deformation. The structured process knowledge base and machine learning optimization mechanism effectively reduce the systematic error of parameter transmission. The multi-modal interactive guidance improves the real-time and intuitiveness of operation deviation correction. The overall solution significantly improves the teaching efficiency and standardization, reduces the entry threshold for novices, and provides technical support for the inheritance and innovation of manual ceramic skills.

[0025] In some of the above solutions of the present application, the visual perception module needs to collect multi-modal visual data in real time. However, in actual applications, the data collected by different sensors may have problems such as unsynchronized timestamps or misaligned spatial coordinates, resulting in the subsequent data processing module being unable to accurately associate multi-modal data and affecting the accuracy of action recognition and process compliance analysis. The present application further proposes that the visual perception module includes an RGB camera, a depth camera, an infrared thermal imager, and a sensor calibration unit. Among them, the RGB camera is configured to collect color images of the operation scene, with a resolution of not less than 1080P and a frame rate of not less than 30fps; the depth camera is configured to collect depth information of the operation scene, with an accuracy of not less than ±2cm; the infrared thermal imager is configured to collect the surface temperature distribution of the mud material, with a wavelength range of 8-14μm and a temperature detection range of -20°C to 1200°C; the sensor calibration unit is configured to perform timestamp synchronization and spatial coordinate calibration on the RGB camera, the depth camera, and the infrared thermal imager. Specifically, the RGB camera, through its high-resolution and high-frame-rate configuration, ensures that the color images of the operation scene can clearly capture the details of hand movements and the morphological changes of the mud material; the depth camera, through high-precision depth information collection, provides spatial coordinate data for 3D hand pose estimation; the infrared thermal imager covers the typical temperature change range of the mud material during the ceramic production process through its wide temperature detection range. For example, the temperature of the mud material during the process of pulling the blank is usually in the range of 20°C to 60°C. The sensor calibration unit eliminates the time difference in multi-sensor data collection through timestamp synchronization, and at the same time unifies the coordinate systems of different sensors to the same reference system through spatial coordinate calibration. For example, it registers the point cloud data of the depth camera with the image pixel coordinates of the RGB camera. Thus, the spatio-temporal consistency of multi-modal visual data is guaranteed, providing accurate input for feature extraction and action recognition by the subsequent data processing module. For example, during the blank-pulling stage, the calibrated multi-modal data can accurately reflect the correlation between the hand movement trajectory and the mud material temperature distribution, avoiding misjudgment of process compliance due to data misalignment.

[0026] As a preferred embodiment, the solution of the present application is specifically implemented as follows: The visual perception module includes an RGB camera, a depth camera, an infrared thermal imager, and a sensor calibration unit. The RGB camera collects color images of the operation scene with a resolution of 1920x1080 pixels and a frame rate of 60fps. The depth camera collects depth information of the operation scene with an accuracy of ±1.5cm. The infrared thermal imager collects the surface temperature distribution of the mud material, with a wavelength range of 8-14μm and a temperature detection range of -20°C to 1200°C. The sensor calibration unit performs timestamp synchronization and spatial coordinate calibration on the RGB camera, the depth camera, and the infrared thermal imager to ensure the spatio-temporal consistency of multi-modal data.

[0027] Specifically, the RGB camera uses a high-resolution CMOS sensor and is equipped with a large-aperture lens, which can clearly capture the details of hand movements in low-light environments. The depth camera uses structured light technology to calculate depth information by projecting an infrared structured light pattern and analyzing its deformation. The infrared thermal imager uses a non-cooled microbolometer to accurately measure the temperature change on the surface of the clay. The sensor calibration unit uses the Zhang Zhengyou calibration method to jointly calibrate multiple cameras and achieves hardware-level trigger synchronization through a time synchronization board.

[0028] Through the above technical solutions, the present application realizes all-round and high-precision visual perception of the manual ceramic production process. The RGB camera provides clear images of the operation scene, the depth camera obtains three-dimensional information of hand movements, and the infrared thermal imager monitors the temperature change of the clay. The collaborative work and precise calibration of multiple sensors ensure the integrity and consistency of the collected data, providing a reliable data basis for subsequent motion analysis and guidance. Thus, the system can accurately capture the operation details of the trainees, realize precise analysis of hand postures, force control, and clay deformation, and effectively overcome the problem of strong subjectivity in feedback in traditional teaching.

[0029] In some of the above solutions of the present application, the data processing module needs to process multi-modal visual data to achieve action recognition and compliance analysis. However, in practical applications, traditional data processing methods are difficult to effectively extract the spatio-temporal features of hand movements, resulting in insufficient action recognition accuracy and inability to accurately quantify operation deviations. The present application further proposes that the data processing module includes a preprocessing unit, a feature extraction unit, an action recognition unit, and a parameter matching unit. Among them, the preprocessing unit uses the U-Net network model to denoise, segment the background, and extract the region of interest from multi-modal visual data, and improves the background segmentation accuracy through a deep learning model; the feature extraction unit extracts the coordinates of hand key points through the HRNet network and estimates the 3D pose of the hand, including the pitch angle, yaw angle, and roll angle, in combination with a 3D convolutional neural network; the action recognition unit models the time series features of continuous actions through the LSTM network to identify the current operation step; the parameter matching unit compares the action features with the standard parameter thresholds in the process knowledge base and calculates the deviation value. Specifically, the preprocessing unit first denoises the original data to eliminate environmental interference, and then segments the background through a U-Net network, retaining the interaction area between the hand and the clay, and reducing the computational amount of redundant data. In the feature extraction unit, the HRNet network maintains the positioning accuracy of hand key points through high-resolution feature maps, and the 3D convolutional neural network extracts three-dimensional pose parameters of the hand from multi-frame depth information to capture the stereo features of hand movements. The action recognition unit uses an LSTM network to model the hand pose sequence of consecutive frames and recognizes the temporal correlation of operations such as throwing and trimming. The parameter matching unit compares features such as displacement trajectory, rotational angular velocity, and contact distance with the standard thresholds in the process knowledge base, and realizes a quantitative evaluation of operation compliance through dynamic deviation calculation. Thus, through multi-stage data processing and feature fusion, the action recognition accuracy and the objectivity of deviation analysis are improved.

[0030] As a preferred embodiment, the solution of the present application is specifically implemented as follows: The data processing module includes a preprocessing unit, a feature extraction unit, an action recognition unit, and a parameter matching unit.

[0031] The preprocessing unit performs denoising, background segmentation, and extraction of the region of interest on multi-modal visual data. The background segmentation uses a U-Net network model. Specifically in implementation, Gaussian filtering denoising is first performed on RGB images, depth images, and infrared thermal images. Then, the pre-trained U-Net model is used to perform semantic segmentation on the images, separating target regions such as the operation tabletop, clay, and tools from the background. Finally, the region of interest containing the hand and the clay is extracted.

[0032] The feature extraction unit extracts the hand key point coordinates through the HRNet network and estimates the 3D pose of the hand through a 3D convolutional neural network (3DCNN). Specifically in implementation, the pre-trained HRNet model is used to extract the 2D coordinates of 21 hand key points from the RGB image. Then, the 2D coordinates are fused with the corresponding depth information to obtain 3D coordinates. These are then input into the 3DCNN network to estimate the pitch angle, yaw angle, and roll angle of the hand.

[0033] The action recognition unit models the temporal sequence features of consecutive actions through an LSTM network and recognizes the current operation step. Specifically in implementation, the 3D pose data of 30 consecutive frames of the hand is used as the input sequence, and temporal modeling is performed through a bidirectional LSTM network. The output layer uses a softmax classifier to recognize the current operation step category, such as throwing, trimming, or kneading.

[0034] The parameter matching unit compares the action features with the standard parameter thresholds in the process knowledge base and calculates the deviation value. Specifically in implementation, features such as the hand displacement trajectory, rotational angular velocity, and the contact distance between the hand and the clay are extracted, and Euclidean distance calculations are performed with the preset standard parameters to obtain the deviation values of each feature.

[0035] Through the above technical solutions, this application realizes the intelligent processing and analysis of multi-modal visual data in the process of manual ceramic production. The preprocessing unit effectively removes background interference and improves the accuracy of subsequent feature extraction. The feature extraction unit accurately captures the key information of hand movements through a deep learning network. The action recognition unit realizes the real-time recognition of continuous actions. The parameter matching unit objectively evaluates the standardization of operations through quantitative comparison. This series of processes provides reliable data support for subsequent interactive guidance, helping to improve the accuracy and effectiveness of teaching.

[0036] In some of the above solutions of this application, the interactive guidance module outputs operation deviation data only through a single feedback method, resulting in difficulty for trainees to quickly understand the correction direction in complex operation scenarios and lacking the guiding effect of multi-sensory collaboration, which affects learning efficiency.

[0037] This application further proposes that the interactive guidance module includes an AR display unit, a voice interaction unit, and a tactile feedback unit.

[0038] Among them, the AR display unit superimposes a standard action trajectory line or a key point reference diagram in the real operation scenario through AR glasses or a mobile phone AR engine. For example, during the pottery throwing stage, it displays the rotation trajectory that the hand should follow; the voice interaction unit generates natural language guidance information through TTS technology, including operation correction prompts or step guidance. For example, when a wrist angle deviation is detected, it generates a voice instruction of "the wrist needs to rotate 15° inward"; the tactile feedback unit outputs a vibration reminder through a wearable device when the operation deviation exceeds a threshold, and the vibration mode is positively correlated with the deviation degree. For example, when the hand trajectory deviation exceeds 5 mm, a high-frequency vibration is triggered.

[0039] Specifically, the AR display unit superimposes the standard action trajectory line on the visual scene of the trainee's actual operation, enabling the trainee to intuitively compare the differences between their own actions and the standard actions; the voice interaction unit provides real-time step guidance through natural language instructions. For example, during the trimming stage, it prompts "currently entering the thickness control stage"; the tactile feedback unit dynamically adjusts the vibration intensity according to the deviation degree. For example, when the contact distance deviation exceeds 10 mm, it outputs continuous vibration. Through the synergistic effect of multi-modal feedback of vision, hearing, and touch, trainees can receive multi-dimensional correction information synchronously during the operation process, shortening the correction time of incorrect actions and improving learning efficiency.

[0040] As a preferred embodiment, the solution of this application is specifically implemented as follows: The interactive guidance module includes an AR display unit, a voice interaction unit, and a tactile feedback unit.

[0041] The AR display unit overlays the standard action trajectory line or key point reference diagram in the real operation scenario through AR glasses or the mobile phone AR engine. Specifically, the AR glasses adopt waveguide display technology, with a resolution of 1920x1080 and a field of view angle of 52°. The mobile phone AR engine is developed based on the ARCore framework and supports functions such as plane detection and image tracking. The standard action trajectory line is presented in the form of a 3D curve, and the key point reference diagram is displayed with semi-transparent overlay.

[0042] The voice interaction unit generates natural language guidance information through TTS technology. Among them, the TTS engine adopts a deep learning model, supports bilingual output in Chinese and English, and the voice tone can be selected as male or female. The guidance information includes two categories: operation correction prompts and step guidance. The operation correction prompts give correction suggestions for specific action deviations, such as "the wrist needs to rotate 15° inward". The step guidance is used to prompt the current operation stage, such as "currently entering the stage of centering the clay blank".

[0043] The tactile feedback unit outputs a vibration reminder through a wearable device when the operation deviation exceeds the threshold. Further, the wearable device is in the form of a smart bracelet, with a linear vibration motor built-in, and the vibration frequency range is 160 - 260Hz. The vibration mode is positively correlated with the degree of deviation. A slight deviation corresponds to a single short vibration, and a serious deviation corresponds to multiple long vibrations.

[0044] Through the above technical solutions, this application realizes multi-modal, real-time, and accurate operation guidance. The AR display intuitively shows the standard actions, the voice interaction provides timely feedback, and the tactile feedback strengthens the error correction reminder. Thus, the trainees can obtain an immersive and personalized learning experience, improving the standardization of actions and learning efficiency. Further, the multi-modal feedback adapts to different learning styles, enhancing knowledge understanding and skill mastery.

[0045] In some of the above solutions of this application, the process knowledge base only stores static standard parameter thresholds and cannot be dynamically optimized according to the actual operation data of the trainees, resulting in a lack of personalized adaptation ability in teaching guidance.

[0046] This application further proposes that the process knowledge base includes a structured database and a machine learning optimization unit. The structured database stores the standard ceramic production process, key action specifications, and typical error cases. The machine learning optimization unit updates the standard parameter thresholds based on the trainees' operation data.

[0047] Among them, the structured database stores the standard step sequences of drawing, trimming, and kneading in a relational database. The key action specifications include the numerical range of the contact angle between the hand and the clay material. The typical error cases contain the eigenvectors of uneven body thickness. The machine learning optimization unit processes the historical correct operation data through a supervised learning algorithm to establish the mapping relationship between action features and standard parameters. At the same time, the reinforcement learning algorithm is used to adjust the action tolerance range according to the incorrect operation data. The structured database and the machine learning optimization unit achieve two-way communication through a data interface, and the standard parameter thresholds are automatically synchronized to the database after being updated.

[0048] Specifically, the structured database stores the standard hand contact angle in the drawing stage as 75° - 85° in tabular form, and the tool movement speed threshold in the trimming stage is 0.5 - 1.2 cm / s. The machine learning optimization unit collects the hand angle data of the trainee's 10 consecutive drawing operations. When 80% of the data falls within the range of 78° - 82°, the standard parameter threshold is adjusted to 78° - 82° through the reinforcement learning algorithm. The adjusted parameter threshold is written into the structured database through the data interface, and the subsequent action compliance analysis is compared using the updated threshold. The eigenvectors of the typical error cases contain data samples with a body thickness standard deviation greater than 0.3 cm. The machine learning optimization unit triggers the parameter update logic when detecting similar features.

[0049] As a preferred embodiment, the solution of the present application is specifically implemented as follows: The process knowledge base includes a structured database and a machine learning optimization unit. The structured database stores the standard process flow, key action specifications, and typical error cases. The machine learning optimization unit updates the standard parameter thresholds based on the trainee's operation data.

[0050] The structured database is implemented using the relational database MySQL. The database includes a process flow table, an action specification table, and an error case table. The process flow table records the step sequences such as drawing, trimming, and kneading. The action specification table stores the key parameters of each step, such as the hand contact angle with the clay material during drawing is 75° - 85°. The error case table records the characteristic data of common errors, such as the numerical range of uneven body thickness.

[0051] The machine learning optimization unit is implemented using the reinforcement learning algorithm. Taking the drawing action as an example, parameters such as the hand displacement trajectory, rotational angular velocity, and contact angle are used as the state space, and the action score is used as the reward function. By collecting a large amount of trainee operation data, the Q-value function is continuously optimized, thereby adjusting the tolerance range of each parameter. For example, the initial setting of the drawing contact angle tolerance is ±5°, and it may be adjusted to ±7° after learning.

[0052] Through the above technical solutions, the present application realizes the structured storage and intelligent optimization of ceramic manufacturing process knowledge. The structured database enables the traceability and quantification of standard process flows and key parameters. The machine learning optimization unit can dynamically adjust parameter thresholds according to actual operation data, improving the adaptability and accuracy of the system. Thus, this solution overcomes the subjectivity and inflexibility of traditional experience teaching and provides objective and dynamic knowledge support for ceramic teaching.

[0053] In some of the above solutions of the present application, the terminal device needs to coordinate the real-time operations of the visual perception, data processing, and interaction guidance modules. However, insufficient embedded hardware performance may lead to delays in multi-modal data processing; an overly small size or low touch sensitivity of the human-computer interaction interface will affect the trainees' viewing of the real-time operation screen and guidance information; a short battery life of the power supply module will limit the continuous use of the teaching system in a power-free environment. The present application further proposes that the main control unit adopts a high-performance embedded computer to coordinate the operations of each module; the human-computer interaction interface is a 10.1-inch touch screen for displaying real-time operation screens, guidance information, learning progress statistics, and historical operation playback; the power supply module adopts a rechargeable lithium battery with a battery life of no less than 8 hours. Among them, the main control unit uses Jetson AGX Orin as the operation core, whose GPU computing power reaches 200 TOPS, supporting the parallel processing of preprocessing, feature extraction, and action recognition tasks of multi-modal data; the human-computer interaction interface adopts a capacitive touch screen with a resolution of 1920×1200, supporting multi-touch operations, and the interface layout is divided into a real-time screen area, a guidance information area, and a statistical data display area; the power supply module is built-in with a 10000mAh lithium polymer battery and equipped with a PD fast charging protocol, which can be fully charged within 2 hours. Specifically, the main control unit connects to the data acquisition card of the visual perception module through the PCIe bus, receives RGB images, depth information, and infrared thermal distribution data, and allocates computing resources to the preprocessing unit, feature extraction unit, and action recognition unit through a multi-threaded scheduling algorithm to ensure that the data processing delay is less than 50ms. The human-computer interaction interface covers an anti-fingerprint coating on the surface of the touch screen, the touch sampling rate is set to 120Hz, the real-time operation screen is displayed at a frame rate of 30fps, and the guidance information area uses a scrolling update method to display voice prompt texts and AR annotation screenshots. The power supply module uses dynamic voltage regulation technology to reduce the working voltage of the main control unit from 12V to 9V in a low-load state, thereby extending the battery life from 6 hours to 8.5 hours and realizing the function of charging while in use through a Type-C interface.

[0054] As a preferred embodiment, the solution of the present application is specifically implemented as follows: The terminal device includes a main control unit, a human-machine interaction interface, and a power supply module. The main control unit uses a high-performance embedded computer Jetson AGX Orin to coordinate the operations of each module. The human-machine interaction interface is a 10.1-inch touch screen, which is used to display real-time operation screens, guidance information, learning progress statistics, and historical operation playback. The power supply module uses a rechargeable lithium battery with a battery life of no less than 8 hours.

[0055] Specifically, the main control unit Jetson AGX Orin is configured with a 12-core ARM Cortex-A78AE CPU and a 2048-core CUDA GPU, with 32GB LPDDR5 memory and 64GB eMMC storage, capable of simultaneously processing multiple high-definition video streams and performing real-time AI inference. The human-machine interaction interface uses a 10.1-inch IPS touch screen with a resolution of 1920x1200, supporting 10-point touch and a brightness of 400nits. The power supply module uses a 10000mAh lithium battery, supporting fast charging technology, and the charging time is about 2 hours.

[0056] Through the above technical solutions, the present application realizes high-performance computing, intuitive interaction, and long-time operation of the terminal device. The powerful computing power of the main control unit supports the real-time operation of complex vision algorithms. The human-machine interaction interface provides clear and intuitive operation guidance, and the long-lasting battery ensures the sustainable operation of the system. These features together improve the efficiency and quality of manual ceramic teaching and reduce the teaching cost.

[0057] In some of the above solutions of the present application, when the infrared thermal imager detects the surface temperature distribution of the clay material, changes in the ambient temperature and humidity will cause deviations in the detection data, affecting the accuracy of the process compliance analysis. The present application further proposes a system including an environmental perception unit. The environmental perception unit collects the temperature and humidity data of the operation environment, with a temperature range of 20°C - 40°C and a humidity range of 30% - 70%, and inputs the environmental parameters into the data processing module to correct the temperature detection error of the infrared thermal imager. Among them, the environmental perception unit integrates temperature and humidity sensors. The temperature and humidity sensors are connected to the data processing module through a wired or wireless communication protocol; the temperature and humidity sensors are installed on the terminal device or independently deployed in the operation area, and their sampling frequency is synchronized with the data acquisition frame rate of the infrared thermal imager; the data processing module has an error correction algorithm built-in. The error correction algorithm establishes a temperature compensation model based on the environmental temperature and humidity data. The compensation model maps the original infrared temperature data to the corrected temperature value through linear regression or look-up table method. Specifically, when the infrared thermal imager collects the surface temperature of the clay material, the change of ambient temperature and humidity will cause fluctuations in the infrared radiation absorption rate. For example, when the ambient humidity is higher than 70%, water vapor in the air absorbs the radiation in the infrared band of 8-14 μm, resulting in the detected temperature being lower than the actual value. The ambient parameters are obtained in real time through the temperature and humidity sensor, and the data processing module calls the compensation model to dynamically correct the infrared temperature data. The compensation model is constructed based on experimental calibration data, and the calibration data contains the mapping relationship between the infrared detection value and the true value of the standard blackbody radiation source under different temperature and humidity combinations. During operation, when the ambient temperature reaches 30 °C and the humidity reaches 60%, the system automatically increases the temperature value detected by the infrared thermal imager by 1.2 °C - 1.8 °C to offset the measurement deviation caused by environmental factors. Thus, the process compliance analysis module can accurately judge the state of the clay material based on the corrected temperature data. For example, during the throwing stage, it can identify the problem of insufficient plasticity of the clay material caused by abnormal temperature.

[0058] As a preferred embodiment, the solution of the present application is specifically implemented as follows: The environmental perception unit includes a temperature and humidity sensor, which is installed on the shell of the terminal device. The temperature and humidity sensor uses the DHT22 model, with a temperature detection range of -40 °C to 80 °C and an accuracy of ±0.5 °C; the relative humidity detection range is 0% to 100% with an accuracy of ±2%. The temperature and humidity data are transmitted to the main control unit through the I2C interface, and the sampling frequency is 1 Hz.

[0059] The data processing module receives the temperature and humidity data, combines the original temperature readings of the infrared thermal imager, and applies the temperature compensation algorithm. The compensation algorithm is based on a polynomial fitting model, considering the influence of ambient temperature and relative humidity on infrared radiation, and dynamically adjusts the emissivity parameter of the thermal imager.

[0060] The corrected temperature data is used for the analysis of the surface temperature distribution of the clay material to support the judgment of process compliance. For example, during the throwing process, if it is detected that the local temperature of the clay material rises abnormally, the system will prompt the trainee to adjust the hand strength or increase lubrication.

[0061] Through the above technical solutions, the present application improves the temperature detection accuracy of the infrared thermal imager under different environmental conditions, providing reliable data support for temperature control in the ceramic manufacturing process. The real-time monitoring and compensation of environmental parameters enable the system to adapt to different working scenarios, improving the accuracy and reliability of teaching guidance.

[0062] In some of the above solutions of the present application, the data processing module separately performs operation step classification and action quality evaluation through independent models, resulting in increased computational resource occupancy and difficulty in establishing the correlation mapping between action features and quality scores, affecting the real-time feedback efficiency and the consistency of evaluation results. The present application further proposes that the data processing module adopts a multi-task learning model during action recognition, and simultaneously outputs the operation step classification results and action quality scores, and the action quality scores are used to quantify the standardization of operations. Among them, the multi-task learning model adopts the architecture of shared underlying feature extraction layer and independent task branch layer. The shared layer uses ResNet-50 network to extract spatiotemporal features. The operation step classification branch uses a fully connected layer and Softmax function to output step category probabilities. The action quality scoring branch uses a regression layer to output a continuous value of 0-100 points. The two task branches use weighted loss functions during training, in which the step classification loss uses the cross entropy function, and the quality scoring loss uses the mean square error function, with a weight ratio of 1:0.8. The action quality score is calculated based on the three indicators of hand trajectory smoothness, angle deviation root mean square value, and contact distance fluctuation amplitude, with weights of 0.4, 0.3, and 0.3 respectively. Specifically, multimodal visual data is preprocessed and input into the shared layer to extract spatiotemporal features, and the features are simultaneously input into two task branches. The step classification branch determines the current stage of throwing, trimming or shaping according to the feature sequence, and the quality scoring branch simultaneously calculates the degree of deviation of the hand movement from the standard parameters and converts it into 0-100 points. For example, when the hand trajectory smoothness is 85 points, the root mean square value of the angle deviation is 75 points, and the contact distance fluctuation amplitude is 80 points, the comprehensive score is 0.4×85+0.3×75+0.3×80=81 points. The scoring results and step classification are transmitted to the interactive guidance module together. The guidance module dynamically adjusts the feedback intensity according to the scoring value, for example, triggering a high-frequency vibration prompt when the score is lower than 70 points. The multi-task model reduces repeated calculations by sharing feature layers, so that step recognition and quality assessment maintain feature consistency, and the weighted loss function balances the requirements of classification accuracy and scoring accuracy.

[0063] As a preferred embodiment, the solution of the present application is specifically implemented as follows: When deploying a multi-task learning model in a data processing module, a dual-branch architecture based on a deep neural network is adopted. The backbone network uses ResNet-50 for feature extraction, the first branch implements the classification of operation steps through a fully connected layer and a Softmax activation function, and the second branch outputs the action quality score through a fully connected layer and a Sigmoid activation function. A joint loss function is used for model training, in which the cross entropy loss function is used for operation step classification, and the mean square error loss function is used for action quality scoring. The weight coefficients of the two loss terms are set to 0.6 and 0.4. During the wheel throwing operation stage, the model simultaneously outputs the "wheel throwing" classification result and an action quality score of 82 points. The scoring basis includes the smoothness of the hand trajectory and the stability of the contact pressure. Through the above technical solutions, the present application realizes the real-time quantitative evaluation of the operation standardization of trainees, and solves the problem of inconsistent feedback caused by relying on subjective experience judgment in traditional teaching. Through the parallel processing ability of the multi-task model, the step classification and quality score are synchronously output in a single inference, reducing the consumption of computing resources. The action quality score, as an objective indicator, can provide trainees with quantifiable improvement directions. For example, the improvement of hand movement stability can be reflected by the change in the score, thereby reducing the material waste caused by operation deviation.

[0064] In some of the above solutions of the present application, traditional manual ceramic teaching relies on the subjective experience judgment of the master, the operation standardization of trainees lacks quantitative evaluation, and the error types are difficult to be systematically classified, resulting in the lack of pertinence of improvement suggestions. The present application further proposes that when the interactive guidance module generates a learning report, it statistically calculates the operation deviation frequency, average deviation value and typical error types of each step of the trainee, and provides improvement suggestions based on the process knowledge base. Among them, the operation deviation frequency is calculated by statistically calculating the proportion of the number of times that the trainee exceeds the tolerance threshold in steps such as drawing, trimming, and plasticine kneading; the average deviation value is obtained by calculating the average difference between the action characteristics such as the hand displacement trajectory and the rotational angular velocity and the standard parameter threshold; the typical error type is identified by matching the pre-stored error case feature data in the process knowledge base. The improvement suggestions are generated based on the error type associated with the optimization strategy in the knowledge base. For example, the hand trajectory deviation corresponds to the hand stability training suggestion. Specifically, when the learning report is generated, the action quality score and the operation step classification result output by the data processing module are transmitted to the interactive guidance module, and the standard parameter threshold and the error case feature data in the process knowledge base are dynamically updated through the machine learning optimization unit. If the frequency of hand trajectory deviation of the trainee in the drawing step is higher than the preset threshold, the system triggers the improvement suggestion generation logic by matching the error case features corresponding to "hand trajectory deviation" in the knowledge base, and outputs the guidance information of "it is recommended to strengthen hand stability training". Through the quantitative statistics and case matching mechanism, this process converts the trainee's operation data into a structured improvement strategy, realizing the objectivity and pertinence of teaching feedback.

[0065] As a preferred embodiment, the solution of the present application is specifically implemented as follows: During the generation of the learning report, the operation deviation frequency is quantified by statistically calculating the number of times of hand trajectory deviation of the trainee in the drawing step, and the average deviation value is obtained by taking the average of the absolute difference between the hand contact angle and the standard parameter. The typical error types are classified into categories such as "uneven contact pressure" and "excessive rotation speed" by clustering algorithms, and the improvement suggestions are generated based on the "hand stability training video link" and "pressure regulation simulation program access entrance" stored in the knowledge base. The learning report shows the corresponding relationship between the error distribution heat map of each step and the suggestion module in tabular form. Through the above technical solutions, the present application realizes the quantitative evaluation and targeted guidance of the operation behaviors of trainees, and solves the problem that traditional teaching lacks objective evaluation criteria. Through the automatic classification of error types and the matching with the knowledge base, the root causes of operation defects can be accurately identified, specific executable training suggestions can be provided, effectively reducing the material waste caused by non-standard actions and shortening the skill mastery cycle.

[0066] In some of the above solutions of the present application, a system solution for manual ceramic teaching guidance through multi-modal data collection and processing is proposed. However, in the actual operation process, the operation process of trainees lacks systematic execution steps and feedback optimization mechanism, resulting in insufficient coherence of operation steps, limited accuracy of real-time guidance, and inability to continuously optimize teaching parameters through historical data.

[0067] The present application further proposes a method for using a manual ceramic teaching system applying a vision sensor, including the following steps: S1: System initialization, collecting multi-modal data of the operation scene through the vision perception module, and the sensor calibration unit completes the spatio-temporal calibration; S2: The trainee starts to operate, and the data processing module collects and preprocesses multi-modal data in real time, and extracts the key features of the hand, clay, and tools; S3: Action recognition and compliance analysis, identifying the current operation step through 3DCNN, analyzing the time series features of continuous actions through LSTM, comparing the action features with the standard parameters in the process knowledge base, and calculating the deviation value; S4: Multi-modal feedback output, if the deviation value ≤ the tolerance threshold, the interactive guidance module outputs a positive prompt; if the deviation value > the tolerance threshold, corrective guidance is output through AR display, voice prompt, and tactile vibration; S5: Learning data recording and optimization, storing the operation data of this time into the process knowledge base, and the machine learning optimization unit updates the standard parameter threshold; S6: After all operation steps are completed, a learning report is generated, showing the scores of each step, common error types, and improvement suggestions. Among them, the spatio-temporal calibration in step S1 is realized through timestamp synchronization and spatial coordinate alignment to ensure the consistency of multi-modal data in the time and space dimensions; the key feature extraction in step S2 includes the hand key point coordinates, the surface temperature distribution of the clay, and the tool displacement trajectory; when comparing the action features with the standard parameters in step S3, a dynamic threshold matching algorithm is adopted, and the tolerance threshold is dynamically adjusted according to the operation proficiency of the trainee; the positive prompt in step S4 includes a green trajectory line in the AR interface or a voice broadcast of "the current step is correct"; the standard parameter threshold update in step S5 adopts an incremental learning algorithm, and the historical data is stored in a time series database; the learning report in step S6 shows the correlation between the deviation frequency of each step and the improvement suggestions through a visualization chart. Specifically, during the system initialization phase, the RGB camera, depth camera, and infrared thermal imager in the visual perception module achieve timestamp synchronization and spatial coordinate alignment through the calibration unit. For example, the timestamp synchronization accuracy is controlled within 10 ms, and the spatial coordinate alignment error does not exceed 2 mm. During the operation of the trainee, the data processing module extracts the hand key point coordinates through the HRNet network and estimates the three-dimensional hand pose through 3DCNN. For example, the pitch angle detection error ≤ 1.5°, and the yaw angle detection error ≤ 2°. The action recognition unit uses the LSTM network to model continuous actions. For example, the time window of the drawing action is set to 5 seconds, and the sampling interval is 0.1 second. During the compliance analysis phase, the parameter matching unit performs dynamic time warping (DTW) comparison between the hand displacement trajectory and the standard trajectory in the process knowledge base, and the deviation value is calculated using the Euclidean distance metric. When the deviation value exceeds the tolerance threshold, the tactile feedback unit adjusts the vibration intensity according to the degree of deviation. For example, when the deviation value is between 10% - 20%, the intermittent vibration mode is adopted, and when the deviation value > 20%, the continuous vibration mode is adopted. During the learning data optimization phase, the machine learning optimization unit updates the standard parameter threshold through the reinforcement learning algorithm. For example, the hand angle tolerance range for the drawing step is adjusted from ±5° to ±7°. When generating the learning report, the typical error types are classified through the clustering algorithm. For example, "hand trajectory deviation" is subdivided into horizontal deviation and vertical deviation categories, and the corresponding improvement training plan is associated.

[0068] As a preferred embodiment, the solution of the present application is specifically implemented as follows: In the ceramic teaching laboratory, the trainee wears a wearable device integrated with an AR glasses, and the operation table is configured with an RGB camera, a depth camera, and an infrared thermal imager array. After the system is started, the visual perception module automatically collects the color image, depth point cloud, and temperature distribution data of the operation area. The sensor calibration unit completes the spatial coordinate alignment of multiple sensors through the checkerboard calibration board and uses the NTP protocol to achieve millisecond-level time synchronization. When the trainee starts the drawing operation, the data processing module extracts the three-dimensional coordinates of 21 key points of the hand through the HRNet network, calculates the pitch angle, yaw angle, and roll angle of the hand using the 3DCNN model, and the LSTM network analyzes the action sequence features of 5 consecutive seconds to identify the current stage as "centering the drawing". The process knowledge base retrieves the standard parameters of this stage, matches the curvature radius and angular velocity threshold of the hand displacement trajectory, and calculates that the trajectory offset reaches 12 mm. The interactive guidance module superimposes a green reference trajectory line on the surface of the billet through the AR glasses, and at the same time plays a voice prompt "The right hand needs to be translated 10 mm outward", and the smart bracelet outputs two short vibrations. After the trainee adjusts the action, the system records the corrected trajectory data and stores it in the knowledge base, and the reinforcement learning model updates the trajectory tolerance threshold to ±8 mm. After completing all the operation steps, the touch screen generates a learning report, showing that the deviation frequency of the "centering the drawing" step has decreased by 35%, and it is recommended to strengthen the wrist stability training. Through the above technical solutions, the present application realizes the real-time action quantification evaluation and dynamic feedback in the ceramic production process, and solves the problem of feedback delay caused by the traditional teaching relying on subjective experience judgment. Through multi-modal data fusion and deep learning models, the deviation of the hand movement trajectory is accurately identified, and combined with AR visual guidance and tactile reminder, so that the trainees can immediately correct incorrect operations. The system automatically records the operation data and optimizes the standard parameter thresholds to form a personalized teaching closed loop, effectively reducing the material waste caused by non-standard operations. The quantitative analysis function of the learning report helps the trainees focus on the weak links and significantly improves the systematicness and repeatability of skill training.

[0069] The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

[0070] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A manual ceramic teaching system applying a vision sensor, characterized in that: It includes a visual perception module, a data processing module, an interaction guidance module, a process knowledge base, and a terminal device integrating the above modules; The visual perception module is used to collect multi-modal visual data of the operation scene of the trainee in real time, The multi-modal visual data includes RGB images, depth information, and infrared thermal distribution data; The data processing module is used to preprocess, extract features, recognize actions, and analyze process compliance of the multi-modal visual data to generate operation deviation data; The interaction guidance module is used to output guidance information through multi-modal feedback according to the operation deviation data; The process knowledge base is used to store standard ceramic production processes, key action specifications, and typical error cases, and support learning and optimization based on operation data; The terminal device is used to integrate each module and provide a human-computer interaction interface to support the operation of the trainee and the interaction with the system.

2. The manual ceramic teaching system applying a vision sensor according to claim 1, characterized in that: The visual perception module includes: An RGB camera, used to collect color images of the operation scene, with a resolution of not less than 1080P and a frame rate of not less than 30fps; A depth camera, used to collect depth information of the operation scene, with an accuracy of not less than ±2cm; An infrared thermal imager, used to collect the surface temperature distribution of the clay material, with a wavelength range of 8-14μm and a temperature detection range of -20°C to 1200°C; A sensor calibration unit, used to synchronize timestamps and calibrate spatial coordinates of the RGB camera, depth camera, and infrared thermal imager to ensure the spatio-temporal consistency of multi-modal data.

3. The manual ceramic teaching system applying a vision sensor according to claim 1, wherein: The data processing module includes: A preprocessing unit, used to denoise, segment the background, and extract regions of interest from multi-modal visual data. The background segmentation uses a U-Net network model; A feature extraction unit, used to extract hand key point coordinates through an HRNet network and estimate the 3D pose of the hand through a 3D convolutional neural network; An action recognition unit, used to model the time series features of continuous actions through an LSTM network to recognize the current operation step; A parameter matching unit, used to compare action features with standard parameter thresholds in the process knowledge base and calculate deviation values.

4. The manual ceramic teaching system applying a vision sensor according to claim 1, wherein: The interaction guidance module includes: An AR display unit, used to overlay standard action trajectory lines or key point reference diagrams in the real operation scene through AR glasses or a mobile phone AR engine; A voice interaction unit, used to generate natural language guidance information through TTS technology, including operation correction prompts or step guidance; A tactile feedback unit, used to output vibration reminders through wearable devices when the operation deviation exceeds the threshold, and the vibration mode is positively correlated with the deviation degree.

5. A manual ceramic teaching system using a vision sensor according to claim 1, characterized in that: The process knowledge base includes: A structured database, storing standard process flows, key action specifications, and typical error cases; A machine learning optimization unit, used to update standard parameter thresholds based on trainee operation data.

6. The manual ceramic teaching system applying a vision sensor according to claim 1, wherein: The terminal device includes: A main control unit, using a high-performance embedded computer, used to coordinate the operations of each module; A human-computer interaction interface, used to display real-time operation screens, guidance information, learning progress statistics, and historical operation replays; A power supply module, using a rechargeable lithium battery, with a battery life of not less than 8 hours.

7. A manual ceramic teaching system applying a vision sensor according to any one of claims 1-6, characterized in that: The described system further includes an environment perception unit, which is used to collect the temperature and humidity data of the operating environment and input the environmental parameters into the data processing module for correcting the temperature detection error of the infrared thermal imager.

8. A manual ceramic teaching system using a vision sensor according to any one of claims 1-6, characterized in that: When performing action recognition, the data processing module adopts a multi-task learning model to simultaneously output the classification result of the operation steps and the action quality score, and the action quality score is used to quantify the operation standardization.

9. A manual ceramic teaching system applying a vision sensor according to any one of claims 1-6, characterized in that: When generating a learning report, the interaction guidance module counts the operation deviation frequency, average deviation value and typical error types of each step of the trainee, and provides improvement suggestions based on the process knowledge base.

10. A method for using a manual ceramic teaching system applying a vision sensor according to any one of claims 1-9, characterized in that: The steps are as follows: (1) System initialization, collecting multi-modal data of the operation scene through the visual perception module, and the sensor calibration unit completes the spatio-temporal calibration; (2) The trainee starts to operate, and the data processing module collects and preprocesses the multi-modal data in real time, and extracts the key features of the hand, the clay material and the tool; (3) Action recognition and compliance analysis, identifying the current operation step through 3DCNN, analyzing the time series features of continuous actions through LSTM, comparing the action features with the standard parameters in the process knowledge base, and calculating the deviation value; (4) Multi-modal feedback output, if the deviation value ≤ the tolerance threshold, the interaction guidance module outputs a positive prompt; if the deviation value > the tolerance threshold, corrective guidance is output through AR display, voice prompt and tactile vibration; (5) Learning data recording and optimization, storing the operation data of this time into the process knowledge base, and the machine learning optimization unit updates the standard parameter threshold; (6) After all operation steps are completed, a learning report is generated, showing the scores of each step, common error types and improvement suggestions.

Citation Information

Cited By

  • Interactive mechanical interaction deduction method and system applied to exhibition hall

    CN120891932A

  • Emergency medical rescue drill teaching system based on virtual reality

    CN121053839A

  • Virtual reality-based emergency medical rescue drill teaching system

    CN121053839B