Woodworking skill training method based on deep reinforcement learning

By introducing deep reinforcement learning and cassia group optimization algorithms into the woodworking skills training system, the training paths are dynamically generated and the students' operating behavior is evaluated in real time, and the problems of insufficient dynamic adaptability and insufficient evaluation system in the existing system are solved, and efficient and personalized woodworking skills training and evaluation are achieved.

CN120107034APending Publication Date: 2025-06-06GUANGZHOU KUMUKU CULTURAL HERITAGE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510182061.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing woodworking skills training system lacks dynamic adaptability, the evaluation system is not comprehensive enough, and the optimization ability is limited, making it difficult to meet the real-time and objectivity of personalized training and operational evaluation, as well as the dynamic optimization needs of training paths.

Method used

Using a method based on deep reinforcement learning, combining virtual reality technology and squid group optimization algorithm, we dynamically generate training paths, evaluate students' operating behaviors in real time, provide quantitative feedback, and optimize training strategies through reward mechanisms and fitness functions.

Benefits of technology

It realizes dynamic optimization of the training path, improves training efficiency and standardization of evaluation, provides a personalized training experience, and enhances the objectivity and real-timeness of skill evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107034A_ABST
    Figure CN120107034A_ABST
Patent Text Reader

Abstract

The invention discloses a woodworking skill training method based on deep reinforcement learning. The woodworking skill training method comprises the following steps: S1, establishing a high-simulation woodworking skill training environment by utilizing a virtual reality technology; s2, collecting operation behavior data of the trainee in the woodworking skill training environment through a motion capture device; s3, constructing a deep reinforcement learning model based on the collected operation behavior data; s4, evaluating the training strategy according to the fitness function; s5, dynamically generating a training path adaptive to the current ability level of the trainee; s6, providing real-time feedback of operation behaviors according to a reward mechanism in the deep reinforcement learning model; s7, generating quantitative scores of the task completion rate, the task continuity and the deviation rate; and S8, when the training is finished, generating a comprehensive training report including a trainee training process, a skill evaluation result and an improvement suggestion. According to the method, multiple improvements of dynamic path generation, quantitative evaluation and strategy optimization are realized in woodworking skill training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of carpentry skills, and in particular to a carpentry skills training method based on deep reinforcement learning. Background Art

[0002] With the development of artificial intelligence technology, intelligent training systems are gradually applied to the field of vocational skills training to improve the efficiency and quality of skills training. Carpentry skills are a highly practical and technical vocational skill. Traditional carpentry skills training methods mainly rely on offline teacher guidance and repeated practice. The teaching effect and evaluation accuracy are limited by many factors in actual application.

[0003] At present, most woodworking skills training adopts a unified teaching curriculum and standardized operating procedures. During the training process, students gradually master the use of tools and craftsmanship skills through observation, imitation and practice. The existing training methods have obvious limitations: on the one hand, due to the large differences in students' skill levels and learning abilities, a unified teaching model is difficult to meet individual needs, resulting in inefficiency in training for some students; on the other hand, traditional woodworking skills training lacks real-time feedback and scientific evaluation of the operation process. Students usually rely on the subjective judgment of teachers for corrections. The evaluation results are easily affected by human factors, and objectivity and real-time performance are difficult to guarantee.

[0004] In recent years, virtual reality technology and artificial intelligence technology have begun to emerge in the field of woodworking skills training. By simulating real operation scenarios in virtual environments and recording trainees' operation behaviors with the help of data collection and analysis technologies, the application of existing technologies has improved training efficiency and evaluation standardization to a certain extent. However, the existing intelligent training systems still have the following significant problems:

[0005] Lack of dynamic adaptability: Existing virtual training systems are usually based on fixed training paths and task designs. It is difficult to dynamically adjust the training paths and task difficulty according to the trainees' real-time performance, resulting in rigid training methods and the inability to give full play to the advantages of personalized training.

[0006] The assessment system is not comprehensive enough: Most existing assessment methods focus on static evaluation of task completion results, lack quantitative analysis and real-time feedback of the operation process, and are difficult to accurately reflect the trainees' skill levels and weaknesses.

[0007] Limited optimization capabilities: Existing technologies mainly rely on simple rule definitions or manual settings in training path and task design, lacking path generation and task allocation mechanisms based on algorithm optimization, resulting in insufficient rationality and pertinence of training content.

[0008] To sum up, the existing technology has significant deficiencies in personalized training, real-time and objectivity of operation evaluation, and dynamic optimization capability of training paths, and it is difficult to meet the needs of modern woodworking skill training for intelligence, efficiency and personalization. The defects of the existing technology directly affect the training effect and skill improvement efficiency of students. A new method is urgently needed to solve the above problems. Summary of the invention

[0009] One purpose of the present invention is to propose a carpentry skill training method based on deep reinforcement learning. The present invention realizes multiple improvements in dynamic path generation, quantitative evaluation and strategy optimization in carpentry skill training.

[0010] A woodworking skill training method based on deep reinforcement learning according to an embodiment of the present invention comprises the following steps:

[0011] S1. Use virtual reality technology to establish a highly simulated woodworking skills training environment;

[0012] S2. Collect the trainees’ operation behavior data in the woodworking skills training environment through motion capture equipment;

[0013] S3. Build a deep reinforcement learning model based on the collected operation behavior data;

[0014] S4. Based on the deep reinforcement learning model, the salp swarm optimization algorithm is used to initialize multiple training strategies, and the training strategies are evaluated according to the fitness function;

[0015] S5. Combine the trial-and-error mechanism of the deep reinforcement learning model and the strategy optimization function of the salp swarm optimization algorithm to dynamically generate training paths that adapt to the current ability level of the trainees;

[0016] S6. Present the optimized training path to the trainees in real time through virtual reality equipment, capture the trainees’ real-time operation behavior data using virtual reality interactive equipment, and provide real-time feedback on the operation behavior based on the reward mechanism in the deep reinforcement learning model;

[0017] S7. Based on the operational behavior data collected during the real-time guidance and feedback process, the reward function of the deep reinforcement learning model and the fitness function of the salp swarm optimization algorithm are used to comprehensively evaluate the trainees' operational behaviors and generate quantitative scores for task completion rate, task continuity, and deviation rate;

[0018] S8. At the end of the training, a comprehensive training report is generated that includes the trainees’ training process, skill assessment results, and improvement suggestions.

[0019] Optionally, the S1 includes the following steps:

[0020] S11. Conduct three-dimensional modeling of woodworking tools required for training, where the parameters of the three-dimensional modeling include the geometric shape, surface characteristics and operation force feedback model of the tools;

[0021] S12. Constructing a virtual model of wood material properties according to the physical property parameters of wood, wherein the physical property parameters include density, hardness, elastic modulus and friction coefficient, and combining data simulation technology in materials science to make the virtual properties of wood materials consistent with the actual physical properties;

[0022] S13. Establish a virtual model of the woodworking skill training operation scene, the operation scene includes an operation table, a tool storage area and a virtual task work area, the modeling parameters of the operation table include size, surface roughness and boundary constraints, the modeling parameters of the tool storage area include tool arrangement and access order, and the modeling parameters of the virtual task work area include task position accuracy and task sequence rules;

[0023] S14. Using a virtual reality device, the virtual model of the woodworking tools, wood material properties and operation scenes is interactively mapped with the trainees' operation behaviors in real time;

[0024] S15. In the virtual woodworking skill training environment, a real-time interactive experience is provided to trainees through tactile feedback devices, visual feedback devices and audio prompt systems. The tactile feedback devices are used to simulate the force feedback of operating tools, the visual feedback devices are used to present virtual scenes and operating task status, and the audio prompt system is used to provide operating suggestions and task completion prompts.

[0025] Optionally, S3 includes the following steps:

[0026] S31. Based on the collected operation behavior data D, define the state space S of the deep reinforcement learning model, wherein the state space includes the contact point position T between the trainee's tool and the material, the real-time trajectory L of the operation path, and the operation force F applied by the trainee;

[0027] S32. Combined with the operational requirements of the woodworking skill training task, define the action space A of the deep reinforcement learning model. The action space includes the following three types of operations:

[0028] Tool Selection Operation A tool , used to characterize trainees’ tool selection behavior in a specific training task;

[0029] Select Operation Type A type , used to describe the types of tasks that trainees perform during training, including cutting, splicing, and grinding;

[0030] Parameter adjustment operation A param , used to indicate the trainees’ optimization adjustment of operation force and path parameters during training;

[0031] S33. Combined with the goal of woodworking skills training, the reward function R(s,a) of the deep reinforcement learning model is designed. The reward function evaluates the trainee's operation performance based on the collected operation behavior data D:

[0032] R(s,a)=w 1 R acc +w 2 R eff -w 3 R safe ;

[0033] Among them, R acc represents the accuracy reward, which quantifies the accuracy of the operation according to the tool trajectory deviation and the target point position deviation. eff represents efficiency reward, which evaluates the operation efficiency based on task completion time and resource usage, R safe represents safety penalty, which is imposed based on the behavior that the operation intensity exceeds the threshold or the path deviates from the safety range. 1 ,w 2 ,w 3 are the weight factors for accuracy, efficiency, and safety, respectively;

[0034] S34. Multiple expert modules are embedded in the deep reinforcement learning model. Each expert module focuses on different skill levels and operation characteristics. The beginner expert module guides students to complete basic tool operation tasks, the intermediate skill expert module optimizes students' multi-task switching and path planning capabilities, and the advanced skill expert module improves students' accuracy and efficiency in operation scenarios.

[0035] S35. Use the trial-and-error mechanism of the deep reinforcement learning model and the results of the reward function R(s,a) to conduct real-time evaluation of the student’s current skill level, and dynamically select the most appropriate expert module to participate in the training task based on the evaluation results.

[0036] Optionally, S4 includes the following steps:

[0037] S41. Based on the woodworking skill training environment and trainee operation behavior data D, the initial training strategy population is randomly generated using the Salp Swarm Optimization Algorithm

[0038] S42. Based on the deep reinforcement learning model, the fitness function F(X i ) for training strategy X i Conduct a comprehensive assessment:

[0039] F(X i )=α·Ψ(acc(X i ))+β·Ξ(eff(X i ))-γ·Ω(safe(Xi ));

[0040] Among them, acc(X i ) represents the training strategy X i The accuracy of matching the tool trajectory with the target point position, eff(X i ) represents the training strategy X i Efficiency indicators in terms of task completion time and operation resource consumption, safe(X i ) represents the training strategy X i The degree of violation of the operation intensity and the path deviation from the safety range, α, β, γ are the importance coefficients of accuracy, efficiency and safety, Ψ(·), Ξ(·) and Ω(·) are nonlinear conversion functions for different evaluation dimensions respectively;

[0041] S43. Using the Salp Swarm Optimization Algorithm to Optimize the Initial Training Strategy Population {X i 0}For multiple iterations, the collaborative update process of the leading salp and the followers includes:

[0042]

[0043] in, is the i-th training strategy of the t+1th generation, Λ(·) represents the strategy update function, is the corresponding training strategy for the tth generation, is the training strategy with the highest fitness value in the current iteration, c 1 is the exploration factor, rand(0,1) is a random number in the interval [0,1], ⊕ represents the strategy merging operation, and Υ(D) is the adaptive offset term extracted based on the operation behavior data set D, which is used to balance the strategy differences in different woodworking operation scenarios;

[0044] S44. Calculate all training strategies at the end of each iteration The fitness value of Compare the highest fitness value with the preset threshold Θ. If the following conditions are met, the iteration ends, otherwise the next generation of iteration continues:

[0045]

[0046] Among them, T max is the maximum number of iterations, Θ is the fitness threshold set for the woodworking skill training process, is the fitness function value;

[0047] S45. When the termination condition is met, the optimal training strategy will be obtained As the final result:

[0048]

[0049] The final output of the optimal training strategy includes the optimized order of task paths and suggestions for adjusting operating parameters.

[0050] Optionally, S5 includes the following steps:

[0051] S51. Combine the trial-and-error mechanism of the deep reinforcement learning model and the reward function R(s,a) to set the training path optimization goal. The optimization goals include dynamic adjustment of task difficulty, optimal arrangement of task sequence, and real-time optimization of operation accuracy:

[0052]

[0053] Among them, G is the total score of the path optimization target, T diff (P i ) represents task P i The difficulty score, Seq(P i ) is task P i The score of the order in the current path, Δ prec (P i ) is task P i The optimization amount of operation accuracy, w 4 ,w 5 ,w 6 is the weight coefficient, which is used to adjust the proportion of different objectives;

[0054] S52. Using the optimal training strategy As the initial input of the optimization path, the attribute set of the parsing task {P 1 ,P 2 ,…,P n Assign initial path attributes to each task, including task difficulty, task order, and initial accuracy threshold;

[0055] S53. Optimize the target G according to the training path and dynamically adjust the task difficulty T using the exploration behavior of the leading salp diff And the task sequence Seq:

[0056] T diff (P i ) t+1 =T diff (P i ) t +c 2 ·rand(0,1)·(T opt -T diff (P i ) t );

[0057] Seq(P i )t+1 =Seq(P i ) t +c 3 ·rand(0,1)·(Seq opt -Seq(P i ) t );

[0058] Among them, T diff (P i ) t and Seq(P i ) t They are respectively the task P at the tth iteration i The difficulty and order value, T opt and Seq opt is the ideal target value, c 2 and c 3 To explore the factors;

[0059] S54. Operation accuracy Δ in the optimized path prec (P i ) to make local adjustments and ensure that the operation accuracy gradually approaches the ideal value through following behavior:

[0060]

[0061] Among them, Δ prec (P i ) t is the operation accuracy value at the tth iteration, Δ prec (P i ) t-1 is the operation accuracy value at the t-1th iteration, κ is the fine-tuning factor, ∈ 1 is the noise term;

[0062] S55. Repeat the exploration and following process of S53 and S54, calculate the path optimization target score G of each iteration, and terminate the iteration if the following termination condition is met:

[0063] max(G t )≥G threshold ort≥T max ;

[0064] Among them, G t is the optimization target score of the tth iteration, G threshold is the preset score threshold, T max is the maximum number of iterations;

[0065] S56. When the termination condition is met, the final training path including the optimization results of task difficulty, task sequence and operation accuracy is output to guide the trainees to carry out the next stage of woodworking skills training.

[0066] Optionally, the S7 includes the following steps:

[0067] S71. During the real-time guidance and feedback process, use virtual reality equipment and sensor systems to collect trainees’ operational behavior data sets D t ;

[0068] S72. Based on the accuracy and effectiveness of students completing tasks, using real-time operation behavior data set D t Define the task completion rate C rate :

[0069]

[0070] Where n is the total number of tasks, I(P i ) represents task P i The completion status of the target is based on the operation trajectory and the target deviation setting;

[0071] S73. Combine the trainees’ task execution sequence and operation path fluency, and use the real-time operation behavior data set D t Defining task continuity C cont :

[0072]

[0073] Among them, T t,i+1 and T t,i denote the target positions of task i+1 and task i respectively, |T t,i+1 -T t,i | represents the position deviation when students switch tasks in their operation path, T t,opt represents the ideal task position switching path;

[0074] S74. Based on the deviation of the trainees’ trajectory and the deviation of the task goal in the operation behavior, the real-time operation behavior data set D t Define the deviation rate C dev :

[0075]

[0076] Among them, L t,i is the actual operation trajectory of task i, L t,opt,i is the ideal operation trajectory of task i, |L t,i -L t,opt,i |Indicates the deviation between the student's actual trajectory and the ideal trajectory;

[0077] S75. Combining the reward function R(s,a) of the deep reinforcement learning model and the fitness function F(X i) Use the quantitative values ​​of task completion rate, task continuity and deviation rate to generate the comprehensive evaluation score of the trainees C total :

[0078]

[0079] Among them, C rate,i , C cont,i , C dev,i Represents tasks P i The task completion rate, task continuity and deviation rate, w 1i 、w 2i 、w 3i For each task P i Weighting factors on completion rate, continuity and deviation rate, α 1 , β 1 , γ 1 is the exponential factor of each indicator, which is used to enhance the impact of different indicators on the evaluation results. ζ is the fitness adjustment coefficient, which is adjusted based on the dynamic performance of the overall operation behavior of the trainees. Λ(F(X i )) represents the fitness function F(X) generated by the salp swarm optimization algorithm i )'s contribution to the score:

[0080]

[0081] Among them, F(X i ) is the fitness value of the current task path strategy, max(F(X i )) and min(F(X i )) are the maximum and minimum fitness values ​​in the current iteration, η 1 Adjusts the smoothing factor for the fitness to avoid large fluctuations in the score.

[0082] The beneficial effects of the present invention are:

[0083] (1) The present invention adopts the trial-and-error mechanism of the deep reinforcement learning model combined with the salp swarm optimization algorithm to realize the dynamic generation and real-time optimization of the training path. By introducing the training path optimization objective function based on the trainees' real-time operation behavior data, the task difficulty, task sequence and operation accuracy can be dynamically adjusted. The trainees' operation ability can be evaluated in real time during the training process, and the training tasks can be dynamically adjusted, so that the trainees can learn at a pace that suits their own skill level.

[0084] (2) The present invention combines the reward function of the deep reinforcement learning model with the fitness function of the salp swarm optimization algorithm to propose three-dimensional quantitative indicators of task completion rate, task continuity and deviation rate, and designs a comprehensive scoring model, which can comprehensively analyze the trainees' operating behaviors, including task completion status, continuity of the operating process, and deviation between actual operations and targets, thereby providing high-precision skill assessment results and effectively avoiding misjudgments caused by human factors.

[0085] (3) The present invention combines the advantages of the salp swarm optimization algorithm in strategy search and global optimization. By dynamically adjusting the training strategy population and updating the optimal strategy in real time, it ensures that trainees can obtain the optimal task allocation and operation guidance at different skill stages. It also adopts the collaborative mechanism of leading and following salps to achieve global optimization and local fine-tuning of task paths for different operation goals, thereby improving the rationality of task allocation and the accuracy of operation parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0087] Figure 1 A flowchart of a woodworking skill training method based on deep reinforcement learning proposed by the present invention;

[0088] Figure 2 A schematic diagram of the training path optimization process of a woodworking skill training method based on deep reinforcement learning proposed in the present invention based on a deep reinforcement learning model and a salp swarm optimization algorithm. DETAILED DESCRIPTION

[0089] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.

[0090] refer to Figure 1-Figure 2 , a woodworking skill training method based on deep reinforcement learning, comprising the following steps:

[0091] S1. Use virtual reality technology to establish a highly simulated woodworking skills training environment;

[0092] S2. Collect the trainees’ operation behavior data in the woodworking skills training environment through motion capture equipment;

[0093] S3. Build a deep reinforcement learning model based on the collected operation behavior data;

[0094] S4. Based on the deep reinforcement learning model, the salp swarm optimization algorithm is used to initialize multiple training strategies, and the training strategies are evaluated according to the fitness function;

[0095] S5. Combine the trial-and-error mechanism of the deep reinforcement learning model and the strategy optimization function of the salp swarm optimization algorithm to dynamically generate training paths that adapt to the current ability level of the trainees;

[0096] S6. Present the optimized training path to the trainees in real time through virtual reality equipment, capture the trainees’ real-time operation behavior data using virtual reality interactive equipment, and provide real-time feedback on the operation behavior based on the reward mechanism in the deep reinforcement learning model;

[0097] S7. Based on the operational behavior data collected during the real-time guidance and feedback process, the reward function of the deep reinforcement learning model and the fitness function of the salp swarm optimization algorithm are used to comprehensively evaluate the trainees' operational behaviors and generate quantitative scores for task completion rate, task continuity, and deviation rate;

[0098] S8. At the end of the training, a comprehensive training report is generated that includes the trainees’ training process, skill assessment results, and improvement suggestions.

[0099] In this implementation, S1 includes the following steps:

[0100] S11. Conduct three-dimensional modeling of woodworking tools required for training, where the parameters of the three-dimensional modeling include the geometric shape, surface characteristics and operation force feedback model of the tools;

[0101] S12. construct a virtual model of wood material properties based on the physical property parameters of wood, the physical property parameters including density, hardness, elastic modulus and friction coefficient, and combine the data simulation technology in materials science to make the virtual properties of wood materials consistent with the actual physical properties;

[0102] S13. Establish a virtual model of the woodworking skill training operation scene, the operation scene includes an operation table, a tool storage area, and a virtual task work area. The modeling parameters of the operation table include size, surface roughness, and boundary constraints. The modeling parameters of the tool storage area include tool arrangement and access sequence. The modeling parameters of the virtual task work area include task position accuracy and task sequence rules;

[0103] S14. Use virtual reality equipment to interactively map the virtual models of woodworking tools, wood material characteristics and operation scenarios with the trainees’ operation behaviors in real time;

[0104] S15. In the virtual woodworking skill training environment, a real-time interactive experience is provided to trainees through tactile feedback devices, visual feedback devices and audio prompt systems. The tactile feedback devices are used to simulate the force feedback of operating tools, the visual feedback devices are used to present virtual scenes and operating task status, and the audio prompt system is used to provide operating suggestions and task completion prompts.

[0105] In this implementation, S3 includes the following steps:

[0106] S31. Based on the collected operation behavior data D, define the state space S of the deep reinforcement learning model, where the state space includes the contact point position T between the trainee's tool and the material, the real-time trajectory L of the operation path, and the operation force F applied by the trainee;

[0107] S32. Combined with the operational requirements of the woodworking skill training task, define the action space A of the deep reinforcement learning model. The action space includes the following three types of operations:

[0108] Tool Selection Operation A tool , used to characterize trainees’ tool selection behavior in a specific training task;

[0109] Select Operation Type A type , used to describe the types of tasks that trainees perform during training, including cutting, splicing, and grinding;

[0110] Parameter adjustment operation A param , used to indicate the trainees’ optimization adjustment of operation force and path parameters during training;

[0111] S33. Combined with the goal of woodworking skills training, the reward function R(s,a) of the deep reinforcement learning model is designed. The reward function evaluates the trainee's operation performance based on the collected operation behavior data D:

[0112] R(s,a)=w 1 R acc +w 2 R eff -w 3 R safe ;

[0113] Among them, R acc represents the accuracy reward, which quantifies the accuracy of the operation according to the tool trajectory deviation and the target point position deviation. eff represents efficiency reward, which evaluates the operation efficiency based on task completion time and resource usage, R safe represents safety penalty, which is imposed based on the behavior that the operation intensity exceeds the threshold or the path deviates from the safety range. 1 ,w 2 ,w 3 are the weight factors for accuracy, efficiency, and safety, respectively;

[0114] S34. Multiple expert modules are embedded in the deep reinforcement learning model. Each expert module focuses on different skill levels and operation characteristics. The beginner expert module guides students to complete basic tool operation tasks, the intermediate skill expert module optimizes students' multi-task switching and path planning capabilities, and the advanced skill expert module improves students' accuracy and efficiency in operation scenarios.

[0115] S35. Use the trial-and-error mechanism of the deep reinforcement learning model and the results of the reward function R(s,a) to conduct real-time evaluation of the student’s current skill level, and dynamically select the most appropriate expert module to participate in the training task based on the evaluation results.

[0116] In this implementation, S4 includes the following steps:

[0117] S41. Based on the woodworking skill training environment and trainee operation behavior data D, the initial training strategy population is randomly generated using the salp swarm optimization algorithm

[0118] S42. Based on the deep reinforcement learning model, the fitness function F(X i ) for training strategy X i Conduct a comprehensive assessment:

[0119] F(X i )=α·Ψ(acc(X i ))+β·Ξ(eff(X i ))-γ·Ω(safe(X i ));

[0120] Among them, acc(X i ) represents the training strategy X i The accuracy of matching the tool trajectory with the target point position, eff(X i ) represents the training strategy X i Efficiency indicators in terms of task completion time and operation resource consumption, safe(X i ) represents the training strategy X i The degree of violation of the operation intensity and the path deviation from the safety range, α, β, γ are the importance coefficients of accuracy, efficiency and safety, Ψ(·), Ξ(·) and Ω(·) are nonlinear conversion functions for different evaluation dimensions respectively;

[0121] S43. Using the Salp Swarm Optimization Algorithm to Optimize the Initial Training Strategy Population After multiple iterations, the collaborative update process of the leading salp and the followers includes:

[0122]

[0123] in, is the i-th training strategy of the t+1th generation, Λ(·) represents the strategy update function, is the corresponding training strategy for the tth generation, is the training strategy with the highest fitness value in the current iteration, c 1 is the exploration factor, rand(0,1) is a random number in the interval [0,1], ⊕ represents the strategy merging operation, and Υ(D) is the adaptive offset term extracted based on the operation behavior data set D, which is used to balance the strategy differences in different woodworking operation scenarios;

[0124] S44. Calculate all training strategies at the end of each iteration The fitness value of Compare the highest fitness value with the preset threshold Θ. If the following conditions are met, the iteration ends, otherwise the next generation of iteration continues:

[0125]

[0126] Among them, T max is the maximum number of iterations, Θ is the fitness threshold set for the woodworking skill training process, is the fitness function value;

[0127] S45. When the termination condition is met, the optimal training strategy will be obtained As the final result:

[0128]

[0129] The final output of the optimal training strategy includes the optimized order of task paths and suggestions for adjusting operating parameters.

[0130] In this implementation, S5 includes the following steps:

[0131] S51. Combine the trial-and-error mechanism of the deep reinforcement learning model and the reward function R(s,a) to set the training path optimization goal. The optimization goals include dynamic adjustment of task difficulty, optimal arrangement of task sequence, and real-time optimization of operation accuracy:

[0132]

[0133] Among them, G is the total score of the path optimization target, T diff (P i ) represents task P i The difficulty score, Seq(P i ) is task P i The score of the order in the current path, Δ prec (P i ) is task P iThe optimization amount of operation accuracy, w 4 ,w 5 ,w 6 is the weight coefficient, which is used to adjust the proportion of different objectives;

[0134] S52. Using the optimal training strategy As the initial input of the optimization path, the attribute set of the parsing task {P 1 ,P 2 ,…,P n Assign initial path attributes to each task, including task difficulty, task order, and initial accuracy threshold;

[0135] S53. Optimize the target G according to the training path and dynamically adjust the task difficulty T using the exploration behavior of the leading salp diff And the task sequence Seq:

[0136] T diff (P i ) t+1 =T diff (P i ) t +c 2 ·rand(0,1)·(T opt -T diff (P i ) t );

[0137] Seq(P i ) t+1 =Seq(P i ) t +c 3 ·rand(0,1)·(Seq opt -Seq(P i ) t );

[0138] Among them, T diff (P i ) t and Seq(P i ) t They are respectively the task P at the tth iteration i The difficulty and order value, T opt and Seq opt is the ideal target value, c 2 and c 3 To explore the factors;

[0139] S54. Operation accuracy Δ in the optimized path prec (P i ) to make local adjustments and ensure that the operation accuracy gradually approaches the ideal value through following behavior:

[0140]

[0141] Among them, Δ prec (P i ) t is the operation accuracy value at the tth iteration, Δ prec (P i ) t-1 is the operation accuracy value at the t-1th iteration, κ is the fine-tuning factor, ∈ 1 is the noise term;

[0142] S55. Repeat the exploration and following process of S53 and S54, calculate the path optimization target score G of each iteration, and terminate the iteration if the following termination condition is met:

[0143] max(G t )≥G threshold ort≥T max ;

[0144] Among them, G t is the optimization target score of the tth iteration, G threshold is the preset score threshold, T max is the maximum number of iterations;

[0145] S56. When the termination condition is met, the final training path including the optimization results of task difficulty, task sequence and operation accuracy is output to guide the trainees to carry out the next stage of woodworking skills training.

[0146] In this implementation, S7 includes the following steps:

[0147] S71. During the real-time guidance and feedback process, use virtual reality equipment and sensor systems to collect trainees’ operational behavior data sets D t ;

[0148] S72. Based on the accuracy and effectiveness of students completing tasks, using real-time operation behavior data set D t Define the task completion rate C rate :

[0149]

[0150] Where n is the total number of tasks, I(P i ) represents task P i The completion status of the target is based on the operation trajectory and the target deviation setting;

[0151] S73. Combine the trainees’ task execution sequence and operation path fluency, and use the real-time operation behavior data set D t Defining task continuity C cont :

[0152]

[0153] Among them, T t,i+1 and T t,i denote the target positions of task i+1 and task i respectively, |T t,i+1 -T t,i | represents the position deviation when students switch tasks in their operation path, T t,opt represents the ideal task position switching path;

[0154] S74. Based on the deviation of the trainees’ trajectory and the deviation of the task goal in the operation behavior, the real-time operation behavior data set D t Define the deviation rate C dev :

[0155]

[0156] Among them, L t,i is the actual operation trajectory of task i, L t,opt,i is the ideal operation trajectory of task i, |L t,i -L t,opt,i |Indicates the deviation between the student's actual trajectory and the ideal trajectory;

[0157] S75. Combining the reward function R(s,a) of the deep reinforcement learning model and the fitness function F(X i ) Use the quantitative values ​​of task completion rate, task continuity and deviation rate to generate the comprehensive evaluation score of the trainees C total :

[0158]

[0159] Among them, C rate,i , C cont,i , C dev,i Represents tasks P i The task completion rate, task continuity and deviation rate, w 1i 、w 2i 、w 3i For each task P i Weighting factors on completion rate, continuity and deviation rate, α 1 , β 1 , γ 1 is the exponential factor of each indicator, which is used to enhance the impact of different indicators on the evaluation results. ζ is the fitness adjustment coefficient, which is adjusted based on the dynamic performance of the overall operation behavior of the trainees. Λ(F(X i )) represents the fitness function F(X) generated by the salp swarm optimization algorithm i )'s contribution to the score:

[0160]

[0161] Among them, F(X i ) is the fitness value of the current task path strategy, max(F(X i )) and min(F(X i )) are the maximum and minimum fitness values ​​in the current iteration, η 1 Adjusts the smoothing factor for the fitness to avoid large fluctuations in the score.

[0162] Embodiment 1:

[0163] In December 2024, a vocational skills training center introduced the intelligent woodworking skills training system based on deep reinforcement learning and salp swarm optimization algorithm described in the present invention to improve the efficiency of woodworking skills training and the skill level of trainees. The laboratory is equipped with a set of high-performance virtual reality equipment, force feedback operation devices and motion capture systems to simulate a real woodworking scene. The system collects trainees' operation behavior data in real time, optimizes the training path through algorithms, and dynamically adjusts the task difficulty. The following is a specific student case to demonstrate the practical application and effect of the method of the present invention.

[0164] On December 5, 2024, student Zhang Ming (a woodworking beginner) started his first woodworking skills training. After wearing the virtual reality device, he entered a simulated woodworking scene. The system assigned him the first task: use a saw to cut a wooden board in a straight line. During the real-time data collection, the system recorded Zhang Ming's multiple trajectory deviations during the operation, and the cutting line showed obvious bending. Specific data showed: the average trajectory deviation was 8 mm, and the cutting completion time was 5 minutes and 25 seconds, which was much higher than the system-recommended standard time of 3 minutes. The task completion rate was only 65%.

[0165] After evaluating Zhang Ming's operating performance through the reward function of the deep reinforcement learning model, the system adjusted the training path, reduced the task difficulty, and enabled the "assisted cutting" mode. In this mode, a clearer cutting path indicator line was provided in the virtual environment, and force feedback prompts were added to help Zhang Ming maintain the stability of the tool during the cutting process. After the second training, real-time data showed that Zhang Ming's trajectory deviation was reduced to 5 mm, the cutting time was reduced to 4 minutes and 10 seconds, and the task completion rate was increased to 82%.

[0166] On December 6, 2024, the system assigned Zhang Ming a second task: splicing a simple mortise and tenon structure. The task required the trainee to use mortise and tenon splicing tools to make precise grooves and splice two pieces of wood. After the training began, real-time collected data showed that the average deviation of Zhang Ming's mortise and tenon groove depth was 3 mm, and the splicing time was 12 minutes and 30 seconds, which was much higher than the recommended time of 8 minutes. The task completion rate was only 70%.

[0167] Through the deep reinforcement learning model and the salp swarm optimization algorithm, the system dynamically adjusted the task path and divided the task into two sub-steps: first, the virtual guidance device prompted Zhang Ming how to use the tool to make the groove; second, after the groove was completed, the real-time feedback mechanism monitored the force and position deviation applied during the splicing process, and gave improvement suggestions. After this adjustment, Zhang Ming reduced the groove depth deviation to 1.5 mm in the second attempt, shortened the splicing time to 9 minutes and 45 seconds, and increased the task completion rate to 88%.

[0168] After the training, the system generated a detailed training report, recording Zhang Ming's task completion rate, task continuity and operation deviation rate data indicators. The report showed that during the training from December 5th to 6th, Zhang Ming's task completion rate increased from 65% in the initial training to 88%, task continuity increased from 74% to 90%, and the operation deviation rate decreased from 18% to 8%.

[0169] In order to verify the effectiveness of the method of the present invention, the training center compared Zhang Ming's training data with the training data of students using the traditional woodworking teaching method:

[0170] project Method of the present invention Traditional methods Initial task completion rate 65% 60% Task completion rate after training 88% 75% Initial mission continuity 74% 68% Post-training task continuity 90% 78% Initial operation deviation rate 18% 22% Operation deviation rate after training 8% 15%

[0171] This example verifies the practical application value of the method of the present invention in woodworking skill training by recording the operation data, task adjustment process and training effect of the trainee Zhang Ming in different tasks. Through dynamic path optimization, real-time feedback and comprehensive evaluation mechanism, the method of the present invention significantly improves the training efficiency and skill mastery level of the trainees, while reducing the operation deviation rate and task completion time, providing a new solution for modern vocational skills training.

[0172] The present invention adopts the trial-and-error mechanism of the deep reinforcement learning model combined with the salp swarm optimization algorithm to realize the dynamic generation and real-time optimization of the training path. By introducing the training path optimization objective function based on the trainees' real-time operation behavior data, the task difficulty, task sequence and operation accuracy can be dynamically adjusted. The trainees' operation ability can be evaluated in real time during the training process, and the training tasks can be dynamically adjusted, so that the trainees can learn at a pace that suits their own skill level.

[0173] By combining the reward function of the deep reinforcement learning model and the fitness function of the salp swarm optimization algorithm, the present invention proposes three-dimensional quantitative indicators of task completion rate, task continuity and deviation rate, and designs a comprehensive scoring model, which can comprehensively analyze the trainees' operating behaviors, including task completion status, continuity of operating procedures, and deviations between actual operations and targets, thereby providing high-precision skill assessment results and effectively avoiding misjudgments caused by human factors.

[0174] The present invention combines the advantages of the salp swarm optimization algorithm in strategy search and global optimization. By dynamically adjusting the training strategy population and updating the optimal strategy in real time, it ensures that trainees at different skill stages can obtain the optimal task allocation and operation guidance. It also adopts the collaborative mechanism of leading and following salps to achieve global optimization and local fine-tuning of task paths for different operation goals, thereby improving the rationality of task allocation and the accuracy of operation parameters.

[0175] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A woodworking skill training method based on deep reinforcement learning, characterized in that: The steps include: S1. Use virtual reality technology to establish a highly simulated woodworking skills training environment; S2. Collect the trainees’ operation behavior data in the woodworking skills training environment through motion capture equipment; S3. Build a deep reinforcement learning model based on the collected operation behavior data; S4. Based on the deep reinforcement learning model, the salp swarm optimization algorithm is used to initialize multiple training strategies, and the training strategies are evaluated according to the fitness function; S5. Combine the trial-and-error mechanism of the deep reinforcement learning model and the strategy optimization function of the salp swarm optimization algorithm to dynamically generate training paths that adapt to the current ability level of the trainees; S6. Present the optimized training path to the trainees in real time through virtual reality equipment, capture the trainees’ real-time operation behavior data using virtual reality interactive equipment, and provide real-time feedback on the operation behavior based on the reward mechanism in the deep reinforcement learning model; S7. Based on the operational behavior data collected during the real-time guidance and feedback process, the reward function of the deep reinforcement learning model and the fitness function of the salp swarm optimization algorithm are used to comprehensively evaluate the trainees' operational behaviors and generate quantitative scores for task completion rate, task continuity, and deviation rate; S8. At the end of the training, a comprehensive training report is generated that includes the trainees’ training process, skill assessment results, and improvement suggestions.

2. A woodworking skill training method based on deep reinforcement learning according to claim 1, characterized in that: The S1 comprises the following steps: S11. Conduct three-dimensional modeling of woodworking tools required for training, where the parameters of the three-dimensional modeling include the geometric shape, surface characteristics and operation force feedback model of the tools; S12. Constructing a virtual model of wood material properties according to the physical property parameters of wood, wherein the physical property parameters include density, hardness, elastic modulus and friction coefficient, and combining data simulation technology in materials science to make the virtual properties of wood materials consistent with the actual physical properties; S13. Establish a virtual model of the woodworking skill training operation scene, the operation scene includes an operation table, a tool storage area and a virtual task work area, the modeling parameters of the operation table include size, surface roughness and boundary constraints, the modeling parameters of the tool storage area include tool arrangement and access order, and the modeling parameters of the virtual task work area include task position accuracy and task sequence rules; S14. Using a virtual reality device, the virtual model of the woodworking tools, wood material properties and operation scenes is interactively mapped with the trainees' operation behaviors in real time; S15. In the virtual woodworking skill training environment, a real-time interactive experience is provided to trainees through tactile feedback devices, visual feedback devices and audio prompt systems. The tactile feedback devices are used to simulate the force feedback of operating tools, the visual feedback devices are used to present virtual scenes and operating task status, and the audio prompt system is used to provide operating suggestions and task completion prompts.

3. A woodworking skill training method based on deep reinforcement learning according to claim 1, characterized in that: The S3 comprises the following steps: S31. Based on the collected operation behavior data D, define the state space S of the deep reinforcement learning model, wherein the state space includes the contact point position T between the trainee's tool and the material, the real-time trajectory L of the operation path, and the operation force F applied by the trainee; S32. Combined with the operational requirements of the woodworking skill training task, define the action space A of the deep reinforcement learning model. The action space includes the following three types of operations: Tool Selection Operation A tool , used to characterize trainees’ tool selection behavior in a specific training task; Select Operation Type A type , used to describe the types of tasks that trainees perform during training, including cutting, splicing, and grinding; Parameter adjustment operation A param , used to indicate the trainees’ optimization adjustment of operation force and path parameters during training; S33. Combined with the goal of woodworking skills training, the reward function R(s,a) of the deep reinforcement learning model is designed. The reward function evaluates the trainee's operation performance based on the collected operation behavior data D: R(s,a)=w1R acc +w2R eff -w3R safe ; Among them, R acc represents the accuracy reward, which quantifies the accuracy of the operation according to the tool trajectory deviation and the target point position deviation. eff represents efficiency reward, which evaluates the operation efficiency based on task completion time and resource usage, R safe represents safety penalty, which is imposed based on the behavior that the operation intensity exceeds the threshold or the path deviates from the safety range. w1, w2, and w3 are the weight factors of accuracy, efficiency, and safety respectively; S34. Multiple expert modules are embedded in the deep reinforcement learning model. Each expert module focuses on different skill levels and operation characteristics. The beginner expert module guides students to complete basic tool operation tasks, the intermediate skill expert module optimizes students' multi-task switching and path planning capabilities, and the advanced skill expert module improves students' accuracy and efficiency in operation scenarios. S35. Use the trial-and-error mechanism of the deep reinforcement learning model and the results of the reward function R(s,a) to conduct real-time evaluation of the student’s current skill level, and dynamically select the most appropriate expert module to participate in the training task based on the evaluation results.

4. A woodworking skill training method based on deep reinforcement learning according to claim 1, characterized in that: The S4 comprises the following steps: S41. Based on the woodworking skill training environment and trainee operation behavior data D, the initial training strategy population is randomly generated using the salp swarm optimization algorithm S42. Based on the deep reinforcement learning model, the fitness function F(X i ) for training strategy X i Conduct a comprehensive assessment: F(X i )=α·Ψ(acc(X i ))+β·Ξ(eff(X i ))-γ·Ω(safe(X i ));; Among them, acc(X i ) represents the training strategy X i The accuracy of matching the tool trajectory with the target point position, eff(X i ) represents the training strategy X i Efficiency indicators in terms of task completion time and operation resource consumption, safe(X i ) represents the training strategy X i The degree of violation of the operation intensity and the path deviation from the safety range, α, β, γ are the importance coefficients of accuracy, efficiency and safety, Ψ(·), Ξ(·) and Ω(·) are nonlinear conversion functions for different evaluation dimensions respectively; S43. Using the Salp Swarm Optimization Algorithm to Optimize the Initial Training Strategy Population After multiple iterations, the collaborative update process of the leading salp and the followers includes: in, is the i-th training strategy of the t+1th generation, Λ(·) represents the strategy update function, is the corresponding training strategy for the tth generation, is the training strategy with the highest fitness value in the current iteration, c1 is the exploration factor, rand(0,1) is a random number in the interval [0,1], represents the strategy merging operation, Υ(D) is the adaptive offset term extracted based on the operation behavior data set D, which is used to balance the strategy differences in different woodworking operation scenarios; S44. Calculate all training strategies at the end of each iteration The fitness value of Compare the highest fitness value with the preset threshold Θ. If the following conditions are met, the iteration ends, otherwise the next generation of iteration continues: Among them, T max is the maximum number of iterations, Θ is the fitness threshold set for the woodworking skill training process, is the fitness function value; S45. When the termination condition is met, the optimal training strategy will be obtained As the final result: The final output of the optimal training strategy includes the optimized order of task paths and suggestions for adjusting operating parameters.

5. A woodworking skill training method based on deep reinforcement learning according to claim 1, characterized in that: The S5 comprises the following steps: S51. Combine the trial-and-error mechanism of the deep reinforcement learning model and the reward function R(s,a) to set the training path optimization goal. The optimization goals include dynamic adjustment of task difficulty, optimal arrangement of task sequence, and real-time optimization of operation accuracy: Among them, G is the total score of the path optimization target, T diff (P i ) represents task P i The difficulty score, Seq(P i ) is task P i The score of the order in the current path, Δ prec (P i ) is task P i The optimization amount of operation accuracy, w4, w5, w6 are weight coefficients used to adjust the proportion of different objectives; S52. Using the optimal training strategy As the initial input of the optimization path, the attribute set of the parsing task {P1, P2, …, P n Assign initial path attributes to each task, including task difficulty, task order, and initial accuracy threshold; S53. Optimize the target G according to the training path and dynamically adjust the task difficulty T using the exploration behavior of the leading salp diff And the task sequence Seq: T diff (P i ) t+1 =T diff (P i ) t +c2·rand(0,1)·(T opt -T diff (P i ) t ); Seq(P i ) t+1 =Seq(P i ) t +c3·rand(0,1)·(Seq opt -Seq(P i ) t ); Among them, T diff (P i ) t and Seq(P i ) t They are respectively the task P at the tth iteration i The difficulty and order value, T opt and Seq opt is the ideal target value, c2 and c3 are exploration factors; S54. Operation accuracy Δ in the optimized path prec (P i ) to make local adjustments and ensure that the operation accuracy gradually approaches the ideal value through following behavior: Among them, Δ prec (P i ) t is the operation accuracy value at the tth iteration, Δ prec (P i ) t-1 is the operation accuracy value at the t-1th iteration, κ is the fine-tuning factor, and ∈1 is the noise term; S55. Repeat the exploration and following process of S53 and S54, calculate the path optimization target score G of each iteration, and terminate the iteration if the following termination condition is met: max(G t )≥G threshold or t≥T max ; Among them, G t is the optimization target score of the tth iteration, G threshold is the preset score threshold, T max is the maximum number of iterations; S56. When the termination condition is met, the final training path including the optimization results of task difficulty, task sequence and operation accuracy is output to guide the trainees to carry out the next stage of woodworking skills training.

6. A woodworking skill training method based on deep reinforcement learning according to claim 1, characterized in that: The S7 comprises the following steps: S71. During the real-time guidance and feedback process, use virtual reality equipment and sensor systems to collect trainees’ operational behavior data sets D t ; S72. Based on the accuracy and effectiveness of students completing tasks, using real-time operation behavior data set D t Define the task completion rate C rate : Where n is the total number of tasks, I(P i ) represents task P i The completion status of the target is based on the operation trajectory and the target deviation setting; S73. Combine the trainees’ task execution sequence and operation path fluency, and use the real-time operation behavior data set D t Defining task continuity C cont : Among them, T t,i+1 and T t,i denote the target positions of task i+1 and task i respectively, |T t,i+1 -T t,i | represents the position deviation when students switch tasks in their operation path, T t,opt represents the ideal task position switching path; S74. Based on the deviation of the trainees’ trajectory and the deviation of the task goal in the operation behavior, the real-time operation behavior data set D t Define the deviation rate C dev : Among them, L t,i is the actual operation trajectory of task i, L t,opt,i is the ideal operation trajectory of task i, |L t,i -L t,opt,i |Indicates the deviation between the student's actual trajectory and the ideal trajectory; S75. Combining the reward function R(s,a) of the deep reinforcement learning model and the fitness function F(X i ) Use the quantitative values ​​of task completion rate, task continuity and deviation rate to generate the comprehensive evaluation score of the trainees C total : Among them, C rate,i , C cont,i , C dev,i Represents tasks P i The task completion rate, task continuity and deviation rate, w 1i 、w 2i 、w 3i For each task P i The weight factors on completion rate, continuity and deviation rate, α1, β1, γ1 are the exponential factors of each indicator, which are used to enhance the influence of different indicators on the evaluation results, ζ is the fitness adjustment coefficient, which is adjusted based on the dynamic performance of the overall operation behavior of the trainees, Λ(F(X i )) represents the fitness function F(X) generated by the salp swarm optimization algorithm i )'s contribution to the score: Among them, F(X i ) is the fitness value of the current task path strategy, max(F(X i )) and min(F(X i )) are the maximum and minimum fitness values ​​in the current iteration respectively, and η1 is the fitness adjustment smoothing factor, which is used to avoid excessive fluctuations in the score.

Citation Information

Cited By

  • Method for optimizing transmission performance of AOC active optical cable for high-speed interconnection of intelligent computing center

    CN122457481A