Practical training monitoring system and method based on machine vision and deep learning

The training and monitoring system based on machine vision and deep learning, combined with multimodal data fusion and adaptive threshold adjustment, solves the shortcomings of traditional monitoring systems in terms of adaptability and security in multiple scenarios, and achieves accurate prediction of training safety risks and objective evaluation of learning effects.

CN120954093AInactive Publication Date: 2025-11-14CHONGQING VOCATIONAL COLLEGE OF IND & INFORMATION TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511092573.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing training and monitoring systems suffer from problems such as fixed and limited monitoring range of single sensors and poor environmental adaptability of traditional image processing algorithms, making it difficult to achieve multi-scenario adaptive safety risk prediction and learning effect evaluation.

Method used

A training and monitoring system based on machine vision and deep learning is adopted. Data is acquired simultaneously through image acquisition, sound acquisition and multiple types of sensors. Multimodal data fusion is performed by combining feature alignment and spatiotemporal correlation algorithms. Real-time analysis is performed using an improved convolutional neural network model. Safety judgment and learning evaluation are carried out through adaptive threshold adjustment and a safety baseline database.

Benefits of technology

It enables accurate prediction of practical training safety risks and objective evaluation of learning outcomes, improves the system's adaptability to multiple scenarios, avoids misjudgments and omissions, and enhances teaching safety and the fairness of evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954093A_ABST
    Figure CN120954093A_ABST
Patent Text Reader

Abstract

The invention relates to the field of teaching monitoring and recognition, and particularly discloses a practical training monitoring system based on machine vision and deep learning, and the system comprises a data collection module which comprises an image collection unit, a sound collection unit, and a multi-type sensor unit, and is used for synchronously obtaining the image data, sound data, and environment parameter data of a practical training scene; the data processing module is used for preprocessing the multi-modal data acquired by the data acquisition module, realizing multi-modal data fusion through feature alignment and a space-time association algorithm, and generating structured practical training data; the deep learning analysis module adopts an improved convolutional neural network model and is used for performing real-time analysis on the fused structured practical training data and identifying personnel actions, equipment states and environment changes in a practical training scene; by adopting the technical scheme of the invention, the defects of dimension limitation of a single sensor and environment robustness of traditional image processing can be broken through, and accurate pre-judgment of practical training safety risks and objective evaluation of learning effects are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of teaching monitoring and identification, and in particular to a training monitoring system and method based on machine vision and deep learning. Background Technology

[0002] In various practical training scenarios, whether it's industrial skills training, vocational education hands-on practice, or scientific research experiments, ensuring safety and evaluating the effectiveness of learning remain core requirements. As training content becomes increasingly complex, equipment becomes more sophisticated, and the number of participants grows, traditional monitoring methods are no longer sufficient to meet practical needs. Currently, practical training monitoring mainly relies on two types of technical solutions: one is a monitoring system based on a single sensor. For example, in machining training, vibration sensors are deployed to monitor equipment anomalies; in chemical engineering training, gas sensors are installed to detect leak risks; or infrared sensors are used to determine whether personnel have entered hazardous areas. While such systems can achieve real-time acquisition of specific parameters, they have significant limitations: the sensor's monitoring range is fixed and its dimensions are singular. For instance, vibration sensors cannot identify operators' failure to wear protective equipment properly, and gas sensors struggle to distinguish between leak sources and environmental interference factors, resulting in incomplete coverage of safety hazards. Furthermore, in terms of evaluating learning effectiveness, a single sensor can only record equipment operating data and cannot correlate the operator's actions with the operational results, making it difficult to form an objective basis for skills assessment. The second approach is a monitoring scheme based on traditional image processing. This involves deploying cameras to collect images of the training scene and using algorithms such as edge detection and feature matching to identify abnormal states. For example, in electrical engineering training, image comparison is used to determine whether wiring steps meet standards; in welding training, color threshold analysis is used to determine whether the flame temperature meets the standard. However, traditional image processing algorithms have extremely poor environmental adaptability: when the lighting in the training scene changes, equipment is obstructed, or the background is complex (such as a cluttered workbench), the algorithm is prone to feature extraction failures, leading to a sharp drop in real-time performance and a surge in false alarm rates. In safety monitoring, false alarms can cause a crisis of trust in alarms among teachers and students, while missed alarms may lead to safety accidents; in skills assessment, insufficient algorithm stability can lead to scoring biases, affecting the fairness of teaching.

[0003] Furthermore, existing technologies generally lack multi-scenario adaptability. The safety risks and skill evaluation indicators differ significantly across different types of practical training (such as mechanical, electronic, and chemical engineering): mechanical training requires a focus on monitoring equipment linkage protection and operational procedures; electronic training needs to pay attention to short-circuit risks and wiring processes; and chemical engineering training requires vigilance regarding reagent leaks and compliance with operational procedures. Traditional systems are often custom-developed for single scenarios, making it difficult to adapt to the needs of multi-domain training through parameter adjustments. This results in universities or enterprises having to repeatedly invest in deploying multiple monitoring systems, leading to resource waste. Therefore, how to overcome the dimensional limitations of a single sensor and the environmental robustness deficiencies of traditional image processing, and build a multi-scenario adaptive monitoring system that integrates deep learning and machine vision technologies to achieve accurate prediction of training safety risks and objective evaluation of learning effects has become a pressing technical challenge in the field of training monitoring. Summary of the Invention

[0004] This invention provides a training monitoring system based on machine vision and deep learning, which can overcome the dimensional limitations of single sensors and the environmental robustness defects of traditional image processing, and achieve accurate prediction of training safety risks and objective evaluation of learning effects.

[0005] To solve the above-mentioned technical problems, this application provides the following technical solution: A training and monitoring system based on machine vision and deep learning includes: The data acquisition module includes an image acquisition unit, a sound acquisition unit, and multiple types of sensor units, which are used to synchronously acquire image data, sound data, and environmental parameter data of the training scenario. The data processing module is used to preprocess the multimodal data acquired by the data acquisition module, and to achieve multimodal data fusion through feature alignment and spatiotemporal correlation algorithms to generate structured training data. The deep learning analysis module uses an improved convolutional neural network model to perform real-time analysis on the fused structured training data, identify personnel actions, equipment status and environmental changes in the training scenario, and obtain analysis results. The analysis and adjustment module is used to receive analysis results and implement adaptive threshold adjustment based on the analysis results. When a teacher's teaching action is detected, it automatically retrieves the preset teacher action feature library and dynamically corrects the threshold parameters according to the teacher's action amplitude, speed and operation area to avoid misjudgment when students imitate the teacher's actions. The evaluation and early warning module has a pre-set safety baseline database. Based on the personnel and equipment contact time, contact speed and risk level parameters output by the deep learning analysis module, and combined with threshold parameters, it makes a safety judgment. When a safety risk is determined, a graded early warning is triggered, and a risk trend prediction result is generated through a behavior prediction algorithm. Based on the analysis results and threshold parameters, a learning evaluation result is generated according to the preset evaluation rules.

[0006] The basic principle and beneficial effects of the solution are as follows: In this invention, the data acquisition module integrates multimodal data from images, sounds, and multiple types of sensors, breaking through the monitoring dimension limitations of a single sensor and capturing training scene information from multiple levels such as vision, hearing, and environmental parameters; the data processing module transforms heterogeneous data into structured training data through feature alignment and spatiotemporal correlation algorithms, solving the problem of traditional data being scattered and difficult to analyze collaboratively, and laying the foundation for subsequent in-depth analysis.

[0007] The deep learning analysis module employs an improved convolutional neural network model, optimizing the feature extraction layer and classifier structure for the dynamic characteristics of practical training scenarios. This enables accurate identification of personnel action details (such as the angle of hand gestures and the force of equipment contact), equipment operating status (such as abnormal vibration frequency and sudden temperature changes), and environmental changes (such as light fluctuations and dust concentration), overcoming the sensitivity of traditional image processing to environmental interference. The analysis and adjustment module dynamically adjusts thresholds based on real-time analysis results and uses a teacher action feature database for differentiated threshold correction. This ensures strict monitoring of student violations while avoiding misjudgments caused by imitation of teacher demonstrations, solving the problem of insufficient adaptability of traditional fixed thresholds in teaching scenarios.

[0008] The evaluation and early warning module combines a safety baseline database with behavior prediction algorithms to achieve graded early warning of safety risks based on parameters such as contact time, speed, and risk level. At the same time, it generates learning effect evaluation results based on preset evaluation rules.

[0009] In achieving accurate prediction and objective evaluation, the adaptive threshold of the analysis and adjustment module is not a simple numerical fluctuation, but rather a model built based on historical data and teacher teaching actions to suit the current scenario. When a novice student's operation is detected, and their imitation of the teacher's actions poses no risk, the threshold parameter is automatically relaxed to reduce unnecessary interference. For high-risk training projects (such as high-voltage electrical work), due to the existence of safety assessments, if a risk exists, the threshold parameter is tightened to the safety baseline, achieving intelligent matching between risk level and threshold sensitivity. This dynamic adjustment mechanism improves the accuracy of safety prediction, avoiding teaching interruptions caused by overly strict thresholds while eliminating safety loopholes caused by overly lenient thresholds.

[0010] In summary, this invention achieves the effects of accurate prediction of practical training safety risks and objective evaluation of learning outcomes.

[0011] Furthermore, the image acquisition unit includes three sets of distributed infrared cameras, two of which are deployed at a 45° angle above the training platform to capture stereoscopic images of the operation, and one set is deployed directly in front of the equipment to capture detailed states of the equipment interface; the sound acquisition unit uses a dual-microphone array to directionally acquire operation command sounds and equipment operation sounds within a 3-meter range using beamforming technology; the multi-type sensor unit includes a temperature and humidity sensor, a vibration sensor, and an infrared ranging sensor, which are respectively installed on the surface of the equipment, the bottom of the operating platform, and the boundary of the danger zone; The data processing module's preprocessing includes dehazing and noise reduction of image data, echo cancellation and feature extraction of sound data, and filtering and calibration of multi-type sensor data. The feature alignment algorithm achieves spatiotemporal matching of multimodal data through timestamp synchronization. The spatiotemporal correlation algorithm uses an improved dynamic time warping algorithm to correlate action frames in the image with the sound features and sensor value change curves of the same period, generating structured training data containing action type, sound feature value, and environmental parameter change rate.

[0012] Furthermore, the improved convolutional neural network model includes a feature enhancement layer, a multi-branch recognition layer, and a result fusion layer. The feature enhancement layer adopts a structure combining residual networks and attention mechanisms to weight and fuse the stereo image features, directional sound features, and sensor parameter change features in the structured training data. Specifically, the edge features of the operation actions collected by the infrared camera are given a first preset weight, the abnormal soundprint features of the device captured by the dual-microphone array are given a second preset weight, and the high-frequency vibration features of the vibration sensor are given a third preset weight. The multi-branch recognition layer contains three parallel sub-networks. The first sub-network extracts the spatiotemporal sequence features of personnel actions through 3D convolution to identify dangerous actions. The second sub-network uses a temporal convolutional network to analyze equipment status parameters to determine whether there is an equipment malfunction. The third sub-network combines temperature, humidity and infrared ranging data and outputs whether the environment is abnormal through a fully connected layer. The result fusion layer integrates the output results of the three sub-networks through a confidence weighting algorithm. When the confidence of a certain recognition result exceeds the preset maximum value, it is directly output as the analysis result. When the confidence is in the middle preset range, the original data of the corresponding modality is called for secondary verification.

[0013] Furthermore, the adaptive threshold adjustment of the analysis and adjustment module includes a dynamic risk assessment model and a real-time guidance triggering mechanism; the dynamic risk assessment model is constructed based on a Markov decision process and calculates the risk index RI of the current operation using the following formula: in, For the first Real-time values ​​of several monitoring parameters, including contact speed and operating force. Based on the threshold, The risk weights for this parameter are derived through training with historical accident data. The scene correction coefficient is dynamically adjusted by the teacher's action feature database; when A level one warning is triggered at this time. A level-two warning is triggered at this time; The real-time guidance triggering mechanism generates a personalized guidance video and sends it to a preset address when it detects that the distance between the student's action and the action sequence calculated by the teacher's standard action using a dynamic time warping algorithm exceeds a threshold.

[0014] Furthermore, the teacher action feature database is constructed using a hierarchical clustering algorithm, which decomposes teacher actions into basic action tuples. Each action tuple contains a spatial coordinate trajectory, a velocity change curve, and a force distribution map; when a teacher's demonstration action is detected, the system matches the corresponding feature cluster using the following similarity formula: in, Hausdorff distance is used to calculate the similarity of location trajectories. The similarity of the velocity curves is calculated using a dynamic time warping algorithm. For force distribution similarity, cosine similarity is used for calculation. These are weighting coefficients, which are dynamically adjusted based on the type of training. when When this action is deemed a teacher demonstration, the corresponding parameter threshold is adjusted as follows: ,in This is the adjustment coefficient.

[0015] Furthermore, the analysis and adjustment module also includes a real-time operation quality evaluation submodule, which calculates the student's operation quality score MQ using the following comprehensive scoring formula: in, For the first A standard sequence of actions, For the actual sequence of student actions, The distance between action sequences is calculated using the dynamic time warping algorithm. The standard action sequence length, This refers to the percentage by which the action takes longer than the standard time. Assign a weight to the importance of the action; when When the error occurs, an operation analysis report is generated, which includes error location markers and improvement suggestions, and sent to a preset address.

[0016] Furthermore, the safety assessment of the evaluation and early warning module adopts a spatiotemporal risk accumulation model, and calculates the comprehensive risk index (CRI) using the following formula: in, Let be a time variable, representing the t-th time unit after the start of monitoring. The total duration of the set risk assessment time window, Let t be the time of contact between the personnel and the equipment. For safe contact time threshold, For contact speed, For the safe speed threshold, The risk level parameter is output through the improved convolutional neural network model. This is a time decay coefficient, used to reflect the degree of impact of risks on the overall risk at different times, and is set according to the safety importance of the training scenario. and These are weighting coefficients, representing the importance of contact time and contact speed in risk assessment, derived from historical safety accident data. when When a red alert is triggered, the device power is automatically cut off and the operating area is locked; A yellow alert is triggered, and a warning message is sent to a preset address.

[0017] Furthermore, the behavior prediction algorithm employs an improved long short-term memory network combined with an attention mechanism, using device state vectors. Vector data derived from device states identified by the deep learning analysis module, and personnel motion vectors. Vector data derived from human actions identified by the deep learning analysis module, and environmental parameter vectors. The vector data derived from environmental parameter data acquired by the data acquisition module is used to construct a three-dimensional prediction model. in, The historical time step variable represents the i-th time unit before the current time t, used to extract historical data. The sigmoid function is used as the activation function to map the summation result to the risk probability range of 0-1, thereby achieving a quantitative output of the risk trend. The result is a prediction of the risk trend in the next k steps. The time step weights reflect the degree of influence of data from different historical moments on future predictions and are obtained through training with historical accident data; LSTM is a Long Short-Term Memory network. The attention mechanism refers to a mechanism that dynamically assigns weights to different features based on their correlation with risk during model training and prediction: for abnormal equipment vibration frequency, the weight increases when the vibration frequency approaches the equipment failure threshold to highlight the early warning role of this feature; for the curvature of personnel hand trajectories, the weight increases accordingly when the trajectory deviates from the standard operating path to focus on monitoring the risk of violations; for the rate of change of environmental temperature and humidity, the weight increases when the rate of change exceeds the safety threshold to strengthen the prediction of safety accidents caused by sudden changes in temperature and humidity. The prediction results are also used to optimize the training process using the following formula: in, The optimized training process time allocation coefficient represents the relative proportion of time that should be allocated to each operation step, taking risk factors into account. The standard time for the j-th operation step is based on the pre-set operation time standard in the practical training syllabus. This is the risk sensitivity coefficient, which reflects the degree of impact of different risk levels on process optimization. These are the operation step numbers, used to distinguish different operation stages in the training process; The predicted risk trend value at time t+j is In specific steps, t represents the current time and j represents the time offset corresponding to the operation step. This is used to associate the operation step with the risk prediction result and realize the risk assessment of a specific step.

[0018] Furthermore, when generating learning evaluation results, the evaluation and early warning module employs a multi-dimensional ability assessment model and calculates the comprehensive learning evaluation index (LEI) using the following formula: in, The distance between the k-th action and the standard action sequence is calculated using the dynamic time warping algorithm. The length of the k-th standard action sequence. This represents the total number of operations. This represents the average value of the comprehensive risk index during the practical training process. The maximum permissible comprehensive risk index set in the safety baseline database; This is the total standard operating time. This represents the total actual operation time; These are weighting coefficients, which respectively reflect the importance of action standardization, safety control capability, and operational efficiency, and are adjusted according to the type of training. The preset evaluation rules include: when When the evaluation result is excellent, it automatically generates reinforcement training suggestions including analysis of the advantageous movements and sends them to a preset address; when When the evaluation result is good, a 3D animation demonstration of the actions that need improvement is pushed and sent to a preset address; when When the evaluation result is "needs improvement", a detailed error analysis report is generated and sent to the preset address. Attached Figure Description

[0019] Figure 1 This is a logical block diagram of a training and monitoring system based on machine vision and deep learning. Detailed Implementation

[0020] The following detailed description illustrates the specific implementation method: A training and monitoring system based on machine vision and deep learning includes: The data acquisition module includes an image acquisition unit, a sound acquisition unit, and multiple types of sensor units, which are used to synchronously acquire image data, sound data, and environmental parameter data of the training scenario. The data processing module is used to preprocess the multimodal data acquired by the data acquisition module, and to achieve multimodal data fusion through feature alignment and spatiotemporal correlation algorithms to generate structured training data. The deep learning analysis module uses an improved convolutional neural network model to perform real-time analysis on the fused structured training data, identify personnel actions, equipment status and environmental changes in the training scenario, and obtain analysis results. The analysis and adjustment module is used to receive analysis results and implement adaptive threshold adjustment based on the analysis results. When a teacher's teaching action is detected, it automatically retrieves the preset teacher action feature library and dynamically corrects the threshold parameters according to the teacher's action amplitude, speed and operation area to avoid misjudgment when students imitate the teacher's actions. The evaluation and early warning module has a pre-set safety baseline database. Based on the personnel and equipment contact time, contact speed and risk level parameters output by the deep learning analysis module, and combined with threshold parameters, it makes a safety judgment. When a safety risk is determined, a graded early warning is triggered, and a risk trend prediction result is generated through a behavior prediction algorithm. Based on the analysis results and threshold parameters, a learning evaluation result is generated according to the preset evaluation rules.

[0021] The image acquisition unit includes three sets of distributed infrared cameras, two of which are deployed at a 45° angle above the training platform to capture stereoscopic images of the operation, and one set is deployed directly in front of the equipment to capture detailed states of the equipment interface; the sound acquisition unit uses a dual-microphone array to directionally acquire operation command sounds and equipment operation sounds within a 3-meter range using beamforming technology; the multi-type sensor unit includes a temperature and humidity sensor, a vibration sensor, and an infrared ranging sensor, which are respectively installed on the surface of the equipment, the bottom of the operating platform, and the boundary of the danger zone. The data processing module's preprocessing includes dehazing and noise reduction of image data, echo cancellation and feature extraction of sound data, and filtering and calibration of multi-type sensor data. The feature alignment algorithm achieves spatiotemporal matching of multimodal data through timestamp synchronization. The spatiotemporal correlation algorithm adopts an improved dynamic time warping algorithm to correlate action frames in the image with the sound features and sensor value change curves of the same period, generating structured training data containing action type, sound feature value, and environmental parameter change rate.

[0022] In practical application: Taking a vocational school's machining training workshop as an example, the system achieves full-process monitoring of lathe operation training (including workpiece clamping, tool feeding, and dimensional measurement). The data acquisition module synchronously collects images, sounds, and environmental parameters of the training scene; the data processing module fuses and processes multimodal data to generate structured training data; the deep learning analysis module identifies operator actions, lathe status, and environmental changes; the analysis and adjustment module dynamically adjusts monitoring thresholds, identifies teacher demonstration actions and avoids misjudgments when students imitate them, while providing real-time guidance; the evaluation and early warning module judges safety risks and triggers early warnings, generating student operation quality evaluation results.

[0023] Specifically, the data acquisition module is configured as follows: The image acquisition unit consists of three sets of infrared cameras or 4K cameras (selected according to budget), deployed at a 45° angle above the lathe training table (2 sets) and directly in front of the lathe spindle (1 set). The angled camera captures stereoscopic images of the operator's hand movements (such as tool grip posture and feed trajectory), while the directly in front camera captures details of the contact between the spindle and the workpiece (such as whether the clamping is secure).

[0024] The sound acquisition unit consists of a dual-microphone array mounted on both sides of the lathe, which uses beamforming technology to directionally acquire sounds within a 3-meter radius, including teacher commands (such as "start the spindle"), lathe operation sounds (such as abnormal spindle noises), and metal cutting sounds (such as abnormal friction sounds).

[0025] The various types of sensor units include: Temperature and humidity sensor (mounted on lathe motor housing): monitors motor temperature (range 20℃-125℃) and workshop humidity (range 0-100%RH); Vibration sensor (mounted at the bottom of the control panel): monitors the vibration frequency of the lathe (range 0~10kHz); Infrared ranging sensor (installed on the edge of the lathe guardrail): monitors the distance between personnel and rotating parts (range 0~5 meters).

[0026] The specific operations of the data processing module are as follows: First, preprocessing is performed on the image data to remove fog and reduce noise (eliminating image blurring caused by workshop dust); echo cancellation is performed on the sound data (filtering noise from other equipment in the workshop) and soundprint features are extracted (e.g., the soundprint of the main shaft during normal operation is 1000Hz±50Hz); and the sensor data is filtered and calibrated (removing high-frequency interference signals from the vibration sensor).

[0027] Then, multimodal data fusion is performed. First, feature alignment is achieved by synchronizing image frames (30 frames / second), sound features (10ms / frame), and sensor data (100ms / time) to the same time axis using timestamp synchronization (accuracy ±10ms). Then, spatiotemporal correlation is performed. An improved dynamic time warping algorithm is used to correlate the image frame of "operator's hand approaching the rotating workpiece" with the infrared ranging sensor value <30cm and vibration sensor frequency change data from the same period to generate structured data (e.g., "t=10.2s, action type: dangerous approach, environmental parameter change rate: vibration frequency +20%)".

[0028] When in use, the improved convolutional neural network model includes a feature enhancement layer, a multi-branch recognition layer, and a result fusion layer. The feature enhancement layer adopts a structure that combines residual networks with attention mechanisms to perform weighted fusion of image stereo features, directional sound features, and sensor parameter change features in the structured training data. Specifically, the edge features of the operation actions captured by the infrared camera are given a first preset weight, the abnormal soundprint features of the device captured by the dual-microphone array are given a second preset weight, and the high-frequency vibration features of the vibration sensor are given a third preset weight. The multi-branch recognition layer contains three parallel sub-networks. The first sub-network extracts the spatiotemporal sequence features of personnel actions through 3D convolution to identify dangerous actions. The second sub-network uses a temporal convolutional network to analyze equipment status parameters to determine whether there is an equipment malfunction. The third sub-network combines temperature, humidity and infrared ranging data and outputs whether the environment is abnormal through a fully connected layer. The result fusion layer integrates the output results of the three sub-networks through a confidence weighting algorithm. When the confidence of a certain recognition result exceeds the preset maximum value, it is directly output as the analysis result. When the confidence is in the middle preset range, the original data of the corresponding modality is called for secondary verification.

[0029] The adaptive threshold adjustment of the analysis and regulation module includes a dynamic risk assessment model and a real-time guidance triggering mechanism; the dynamic risk assessment model is based on a Markov decision process and calculates the risk index RI of the current operation using the following formula: in, For the first Real-time values ​​of several monitoring parameters, including contact speed and operating force. Based on the threshold, The risk weights for this parameter are derived through training with historical accident data. The scene correction coefficient is dynamically adjusted by the teacher's action feature database; when A level one warning is triggered at this time. A level-two warning is triggered at this time; The real-time guidance triggering mechanism generates a personalized guidance video and sends it to a preset address when it detects that the distance between the student's action and the action sequence calculated by the teacher's standard action using a dynamic time warping algorithm exceeds a threshold.

[0030] The teacher action feature database is constructed using a hierarchical clustering algorithm, which decomposes teacher actions into basic action tuples. Each action tuple contains a spatial coordinate trajectory, a velocity change curve, and a force distribution map; when a teacher's demonstration action is detected, the system matches the corresponding feature cluster using the following similarity formula: in, Hausdorff distance is used to calculate the similarity of location trajectories. The similarity of the velocity curves is calculated using a dynamic time warping algorithm. For force distribution similarity, cosine similarity is used for calculation. These are weighting coefficients, which are dynamically adjusted based on the type of training. when When this action is deemed a teacher demonstration, the corresponding parameter threshold is adjusted as follows: ,in This is the adjustment coefficient.

[0031] The analysis and adjustment module also includes a real-time evaluation submodule for operational quality, which calculates the student's operational quality score (MQ) using the following comprehensive scoring formula: in, For the first A standard sequence of actions, For the actual sequence of student actions, The distance between action sequences is calculated using the dynamic time warping algorithm. The standard action sequence length, This refers to the percentage by which the action takes longer than the standard time. Assign a weight to the importance of the action; when When the error occurs, an operation analysis report is generated, which includes error location markers and improvement suggestions, and sent to a preset address.

[0032] The safety assessment of the early warning module adopts a spatiotemporal risk accumulation model, and the comprehensive risk index (CRI) is calculated using the following formula: in, Let be a time variable, representing the t-th time unit after the start of monitoring. The total duration of the set risk assessment time window, Let t be the time of contact between the personnel and the equipment. For safe contact time threshold, For contact speed, For the safe speed threshold, The risk level parameter is output through the improved convolutional neural network model. This is a time decay coefficient, used to reflect the degree of impact of risks on the overall risk at different times, and is set according to the safety importance of the training scenario. and These are weighting coefficients, representing the importance of contact time and contact speed in risk assessment, derived from historical safety accident data. when When a red alert is triggered, the device power is automatically cut off and the operating area is locked; A yellow alert is triggered, and a warning message is sent to a preset address.

[0033] The behavior prediction algorithm employs an improved long short-term memory network combined with an attention mechanism, using device state vectors. Vector data derived from device states identified by the deep learning analysis module, and personnel motion vectors. Vector data derived from human actions identified by the deep learning analysis module, and environmental parameter vectors. The vector data derived from environmental parameter data acquired by the data acquisition module is used to construct a three-dimensional prediction model. in, The historical time step variable represents the i-th time unit before the current time t, used to extract historical data. The sigmoid function is used as the activation function to map the summation result to the risk probability range of 0-1, thereby achieving a quantitative output of the risk trend. The result is a prediction of the risk trend in the next k steps. The time step weights reflect the degree of influence of data from different historical moments on future predictions and are obtained through training with historical accident data; LSTM is a Long Short-Term Memory network. The attention mechanism refers to a mechanism that dynamically assigns weights to different features based on their correlation with risk during model training and prediction: for abnormal equipment vibration frequency, the weight increases when the vibration frequency approaches the equipment failure threshold to highlight the early warning role of this feature; for the curvature of personnel hand trajectories, the weight increases accordingly when the trajectory deviates from the standard operating path to focus on monitoring the risk of violations; for the rate of change of environmental temperature and humidity, the weight increases when the rate of change exceeds the safety threshold to strengthen the prediction of safety accidents caused by sudden changes in temperature and humidity. The prediction results are also used to optimize the training process using the following formula: in, The optimized training process time allocation coefficient represents the relative proportion of time that should be allocated to each operation step, taking risk factors into account. The standard time for the j-th operation step is based on the pre-set operation time standard in the practical training syllabus. This is the risk sensitivity coefficient, which reflects the degree of impact of different risk levels on process optimization. These are the operation step numbers, used to distinguish different operation stages in the training process; The predicted risk trend value at time t+j is In specific steps, t represents the current time and j represents the time offset corresponding to the operation step. This is used to associate the operation step with the risk prediction result and realize the risk assessment of a specific step.

[0034] When generating learning evaluation results, the evaluation and early warning module uses a multi-dimensional ability assessment model to calculate the comprehensive learning evaluation index (LEI) using the following formula: in, The distance between the k-th action and the standard action sequence is calculated using the dynamic time warping algorithm. The length of the k-th standard action sequence. This represents the total number of operations. This represents the average value of the comprehensive risk index during the practical training process. The maximum permissible comprehensive risk index set in the safety baseline database; This is the total standard operating time. This represents the total actual operation time; These are weighting coefficients, which respectively reflect the importance of action standardization, safety control capability, and operational efficiency, and are adjusted according to the type of training. The preset evaluation rules include: when When the evaluation result is excellent, it automatically generates reinforcement training suggestions including analysis of the advantageous movements and sends them to a preset address; when When the evaluation result is good, a 3D animation demonstration of the actions that need improvement is pushed and sent to a preset address; when When the evaluation result is "needs improvement", a detailed error analysis report is generated and sent to the preset address.

[0035] In practical application, an improved convolutional neural network model is used to analyze the fused structured data. Specifically, the feature enhancement layer fuses and weights multimodal features, such as the edge features of the tool feed trajectory captured by the infrared camera (weight 60%), the abnormal soundprint features of the lathe collected by the microphone (weight 25%), and the high-frequency vibration features of the vibration sensor (weight 15%).

[0036] In the multi-branch recognition layer, the first sub-network (personnel action recognition) extracts spatiotemporal sequence features through 3D convolution to identify dangerous actions such as operating without protective gloves or crossing the safety red line; the second sub-network (equipment status recognition) analyzes vibration frequency and temperature data through temporal convolutional networks to determine faults such as abnormal spindle speed and tool wear (e.g., when the vibration frequency is continuously >2000Hz and the temperature is >60℃, it is determined to be tool wear); the third sub-network (environmental anomaly recognition) combines temperature, humidity and infrared ranging data to output environmental anomalies such as workshop humidity >80% (which can easily lead to equipment corrosion) and personnel intrusion into the protected area.

[0037] In the result fusion layer, if the confidence level of a certain recognition result is >90% (such as a hand crossing the safety red line), it is output directly; if the confidence level is between 60% and 90% (such as suspected tool wear), the original image (tool close-up) and sound data (cutting soundprint) are called for secondary verification, and the accuracy is finally confirmed to be improved.

[0038] In the adaptive threshold adjustment process, a risk index RI model is first constructed based on a Markov decision process: in: This refers to the contact speed (such as the speed at which the hand approaches the workpiece, measured value 5m / s). The safety threshold is 3m / s. ; This refers to the operating force (such as the clamping force for clamping a workpiece, measured value 800N). The safety threshold is 500N. ; The vibration frequency (measured value 2500Hz). The safe threshold is 2000Hz. ; The scene correction factor is 1.5 for teacher demonstrations and 1.0 for student operations.

[0039] when (If the calculated value is 1.3), a Level 1 warning is triggered (audio-visual alarm + lathe pause); when This triggered a Level II warning.

[0040] Then, teacher action recognition and threshold correction are performed. The teacher action feature database uses a hierarchical clustering algorithm to store basic action tuples for demonstrating workpiece clamping (e.g., spatial trajectory is the coordinate sequence from workpiece placement to chuck clamping; velocity curve is uniform motion within 0-5 seconds; force distribution: chuck clamping force 300-400N). The logic for similarity matching is that when the teacher demonstrates, the similarity with the feature database is calculated using a formula. (>0.85), judged as a demonstration action, threshold correction: lathe training adjustment coefficient: For example, the contact speed threshold has been relaxed from 3m / s to 3.096m / s to avoid misjudgment when students imitate it.

[0041] Then, real-time guidance and quality evaluation are carried out. When the dynamic time normalization distance between the student's action and the teacher's standard action is greater than the threshold (such as the tool feed trajectory deviation is greater than 5mm), a personalized guidance video is generated (such as the feed angle should be kept at 30°) and sent to the preset address.

[0042] Operation Quality Score (MQ) Calculation: in Step 5 corresponds to the five steps of "workpiece clamping" and "tool setting". The values ​​are 0.2, 0.3, 0.2, 0.1, and 0.2 respectively; if a student's clamping process takes 20% more time than the standard time ( ), and trajectory deviation distance If this happens, the score for that step will decrease by 15%. The system generates a report (e.g., if the clamping force is insufficient, it is recommended to increase the clamping pressure) and sends it to a preset address, such as the teacher's terminal.

[0043] Finally, a safety risk assessment and early warning are conducted, and the Comprehensive Risk Index (CRI) is calculated (within a time window). ): in Let t be the contact time between the hand and the rotating workpiece. For contact speed, Risk level (1-5). When (For example, if the contact time is continuous for 8 seconds and the speed is 4 m / s), a red warning is triggered (automatically cutting off the lathe power and locking the operating table); when This triggers a yellow alert (a message indicating a high risk of student actions being pushed to the teacher's terminal).

[0044] The risk trend prediction and process optimization specifically involves constructing a predictive model using an improved LSTM combined with an attention mechanism. Inputs include device state vectors (e.g., spindle speed), personnel motion vectors (e.g., hand trajectories), and environmental vectors (e.g., temperature), to predict risk trends over the next 3 seconds. For example, if it is predicted that the tool will collide with the workpiece at t+3s, a deceleration command can be triggered in advance. Training process optimization: Through... Adjust the time allocation for each step, such as for high-risk tool setting steps ( Increases operation time by 20% and reduces collision risk.

[0045] Comprehensive learning assessment is achieved through the calculation of the Learning Assessment Index (LEI): in The average risk index throughout the entire training process. This is the standard duration. If a student's LEI = 0.88 (≥ 0.85), the evaluation is excellent, high feed accuracy is generated, and suggestions to strengthen high-speed cutting practice are recommended; if LEI = 0.65 (< 0.7), a "3D error correction animation of clamping steps" will be pushed. The preset address can be the teacher's email address or server address, making it convenient for students or teachers to access via smart terminals and for the overall display through a visual interface.

[0046] The above are merely embodiments of the present invention. The invention is not limited to the fields covered by these embodiments. Commonly known structures and characteristics in the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are able to access all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. A training and monitoring system based on machine vision and deep learning, characterized in that, include: The data acquisition module includes an image acquisition unit, a sound acquisition unit, and multiple types of sensor units, which are used to synchronously acquire image data, sound data, and environmental parameter data of the training scenario. The data processing module is used to preprocess the multimodal data acquired by the data acquisition module, and to achieve multimodal data fusion through feature alignment and spatiotemporal correlation algorithms to generate structured training data. The deep learning analysis module uses an improved convolutional neural network model to perform real-time analysis on the fused structured training data, identify personnel actions, equipment status and environmental changes in the training scenario, and obtain analysis results. The analysis and adjustment module is used to receive analysis results and implement adaptive threshold adjustment based on the analysis results. When a teacher's teaching action is detected, it automatically retrieves the preset teacher action feature library and dynamically corrects the threshold parameters according to the teacher's action amplitude, speed and operation area to avoid misjudgment when students imitate the teacher's actions. The evaluation and early warning module has a pre-set safety baseline database. Based on the personnel and equipment contact time, contact speed and risk level parameters output by the deep learning analysis module, combined with threshold parameters, it makes a safety judgment. When a safety risk is determined to exist, a graded early warning is triggered, and a risk trend prediction result is generated through a behavior prediction algorithm. Based on the analysis results and threshold parameters, learning evaluation results are generated according to preset evaluation rules.

2. The training and monitoring system based on machine vision and deep learning according to claim 1, characterized in that, The image acquisition unit includes three sets of distributed infrared cameras, two of which are deployed at a 45° angle above the training platform to capture stereoscopic images of the operation, and one set is deployed directly in front of the equipment to capture detailed states of the equipment interface; the sound acquisition unit uses a dual-microphone array to directionally acquire operation command sounds and equipment operation sounds within a 3-meter range using beamforming technology; the multi-type sensor unit includes a temperature and humidity sensor, a vibration sensor, and an infrared ranging sensor, which are respectively installed on the surface of the equipment, the bottom of the operating platform, and the boundary of the danger zone. The data processing module's preprocessing includes dehazing and noise reduction of image data, echo cancellation and feature extraction of sound data, and filtering and calibration of multi-type sensor data. The feature alignment algorithm achieves spatiotemporal matching of multimodal data through timestamp synchronization. The spatiotemporal correlation algorithm uses an improved dynamic time warping algorithm to correlate action frames in the image with the sound features and sensor value change curves of the same period, generating structured training data containing action type, sound feature value, and environmental parameter change rate.

3. The training and monitoring system based on machine vision and deep learning according to claim 2, characterized in that, The improved convolutional neural network model includes a feature enhancement layer, a multi-branch recognition layer, and a result fusion layer. The feature enhancement layer adopts a structure that combines residual networks and attention mechanisms to perform weighted fusion of image stereo features, directional sound features, and sensor parameter change features in the structured training data. Specifically, the edge features of operation actions captured by the infrared camera are given a first preset weight, the abnormal soundprint features of the device captured by the dual-microphone array are given a second preset weight, and the high-frequency vibration features of the vibration sensor are given a third preset weight. The multi-branch recognition layer contains three parallel sub-networks. The first sub-network extracts the spatiotemporal sequence features of personnel actions through 3D convolution to identify dangerous actions. The second sub-network uses a temporal convolutional network to analyze equipment status parameters to determine whether there is an equipment malfunction. The third sub-network combines temperature, humidity and infrared ranging data and outputs whether the environment is abnormal through a fully connected layer. The result fusion layer integrates the output results of the three sub-networks through a confidence weighting algorithm. When the confidence of a certain recognition result exceeds the preset maximum value, it is directly output as the analysis result. When the confidence is in the middle preset range, the original data of the corresponding modality is called for secondary verification.

4. The training and monitoring system based on machine vision and deep learning according to claim 3, characterized in that, The adaptive threshold adjustment of the analysis and adjustment module includes a dynamic risk assessment model and a real-time guidance triggering mechanism; the dynamic risk assessment model is based on a Markov decision process and calculates the risk index RI of the current operation using the following formula: in, For the first Real-time values ​​of several monitoring parameters, including contact speed and operating force. Based on the threshold, The risk weights for this parameter are derived through training with historical accident data. The scene correction coefficient is dynamically adjusted by the teacher's action feature database; when A level one warning is triggered at this time. A level-two warning is triggered at this time; The real-time guidance triggering mechanism generates a personalized guidance video and sends it to a preset address when it detects that the distance between the student's action and the action sequence calculated by the teacher's standard action using a dynamic time warping algorithm exceeds a threshold.

5. The training and monitoring system based on machine vision and deep learning according to claim 4, characterized in that, The teacher action feature database is constructed using a hierarchical clustering algorithm, which decomposes teacher actions into basic action tuples. Each action tuple contains a spatial coordinate trajectory, a velocity change curve, and a force distribution map; when a teacher's demonstration action is detected, the system matches the corresponding feature cluster using the following similarity formula: in, Hausdorff distance is used to calculate the similarity of location trajectories. The similarity of the velocity curves is calculated using a dynamic time warping algorithm. For force distribution similarity, cosine similarity is used for calculation. These are weighting coefficients, which are dynamically adjusted based on the type of training. when When this action is deemed a teacher demonstration, the corresponding parameter threshold is adjusted as follows: ,in This is the adjustment coefficient.

6. The training and monitoring system based on machine vision and deep learning according to claim 5, characterized in that, The analysis and adjustment module also includes a real-time operation quality evaluation submodule, which calculates the student's real-time operation quality score (MQ) using the following comprehensive scoring formula: in, For the first A standard sequence of actions, For the actual sequence of student actions, The distance between action sequences is calculated using the dynamic time warping algorithm. The standard action sequence length, This refers to the percentage by which the action takes longer than the standard time. Assign a weight to the importance of the action; when When the error occurs, an operation analysis report is generated, which includes error location markers and improvement suggestions, and sent to a preset address.

7. The training and monitoring system based on machine vision and deep learning according to claim 6, characterized in that, The safety assessment module uses a spatiotemporal risk accumulation model to calculate the comprehensive risk index (CRI) using the following formula: in, Let be a time variable, representing the t-th time unit after the start of monitoring. The total duration of the set risk assessment time window, Let t be the time of contact between the personnel and the equipment. For safe contact time threshold, For contact speed, For the safe speed threshold, The risk level parameter is output through the improved convolutional neural network model. This is a time decay coefficient, used to reflect the degree of impact of risks on the overall risk at different times, and is set according to the safety importance of the training scenario. and These are weighting coefficients, representing the importance of contact time and contact speed in risk assessment, derived from historical safety accident data. when When a red alert is triggered, the device power is automatically cut off and the operating area is locked; A yellow alert is triggered, and a warning message is sent to a preset address.

8. The training and monitoring system based on machine vision and deep learning according to claim 7, characterized in that, The behavior prediction algorithm employs an improved long short-term memory network combined with an attention mechanism, using device state vectors. Vector data derived from device states identified by the deep learning analysis module, and personnel motion vectors. Vector data derived from human actions identified by the deep learning analysis module, and environmental parameter vectors. The vector data derived from environmental parameter data acquired by the data acquisition module is used to construct a three-dimensional prediction model. in, The historical time step variable represents the i-th time unit before the current time t, used to extract historical data. The sigmoid function is used as the activation function to map the summation result to the risk probability range of 0-1, thereby achieving a quantitative output of the risk trend. The result is a prediction of the risk trend in the next k steps. The time step weights reflect the degree of influence of data from different historical moments on future predictions and are obtained through training with historical accident data; LSTM is a Long Short-Term Memory network. The attention mechanism refers to a mechanism that dynamically assigns weights to different features based on their correlation with risk during model training and prediction: for abnormal equipment vibration frequency, the weight increases when the vibration frequency approaches the equipment failure threshold to highlight the early warning role of this feature; for the curvature of personnel hand trajectories, the weight increases accordingly when the trajectory deviates from the standard operating path to focus on monitoring the risk of violations; for the rate of change of environmental temperature and humidity, the weight increases when the rate of change exceeds the safety threshold to strengthen the prediction of safety accidents caused by sudden changes in temperature and humidity. The prediction results are also used to optimize the training process using the following formula: in, The optimized training process time allocation coefficient represents the relative proportion of time that should be allocated to each operation step, taking risk factors into account. The standard time for the j-th operation step is based on the pre-set operation time standard in the practical training syllabus. This is the risk sensitivity coefficient, which reflects the degree of impact of different risk levels on process optimization. These are the operation step numbers, used to distinguish different operation stages in the training process; The predicted risk trend value at time t+j is In specific steps, t represents the current time and j represents the time offset corresponding to the operation step. This is used to associate the operation step with the risk prediction result and realize the risk assessment of a specific step.

9. The training and monitoring system based on machine vision and deep learning according to claim 8, characterized in that, When generating learning evaluation results, the evaluation and early warning module uses a multi-dimensional ability assessment model to calculate the comprehensive learning evaluation index (LEI) using the following formula: in, The distance between the k-th action and the standard action sequence is calculated using the dynamic time warping algorithm. The length of the k-th standard action sequence. This represents the total number of operations. This represents the average value of the comprehensive risk index during the practical training process. The maximum permissible comprehensive risk index set in the safety baseline database; This is the total standard operating time. This represents the total actual operation time. These are weighting coefficients, which respectively reflect the importance of action standardization, safety control capability, and operational efficiency, and are adjusted according to the type of training. The preset evaluation rules include: when When the evaluation result is excellent, it automatically generates reinforcement training suggestions including analysis of the advantageous movements and sends them to a preset address; when When the evaluation result is good, a 3D animation demonstration of the actions that need improvement is pushed and sent to a preset address; when When the evaluation result is "needs improvement", a detailed error analysis report is generated and sent to the preset address.

10. A training and monitoring method based on machine vision and deep learning, characterized in that, The system described in any one of claims 1-9 is employed.

Citation Information

Cited By

  • Practical training evaluation method and system based on multi-modal data fusion

    CN121707797A

  • A virtual torque-driven motor disassembly and assembly intelligent evaluation system and method

    CN122347890A

  • A virtual torque-driven motor disassembly and assembly intelligent evaluation system and method

    CN122347890B