Virtual simulation-based aircraft composite material repair teaching interaction method and system
By constructing a virtual simulation environment and using deep learning algorithms to identify trainees' operations, and combining this with reinforcement learning to generate personalized strategies, the problem of virtual teaching systems being unable to identify operational norms and adjust difficulty has been solved, thus achieving efficient training in composite material repair.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING AEROSPACE POLYTECHNIC COLLEGE
- Filing Date
- 2026-03-24
- Publication Date
- 2026-04-24
AI Technical Summary
Existing virtual teaching systems cannot accurately identify students' operational procedures, lack adaptive learning strategy adjustment mechanisms, and cannot effectively demonstrate the physical evolution of composite material repair processes, resulting in poor learning outcomes for students.
A virtual simulation environment for damage to composite materials in aircraft is constructed. Deep learning algorithms are used to identify trainee operations in real time. Reinforcement learning algorithms are combined to dynamically generate personalized learning strategies. Spatiotemporal features of operations are captured through spatiotemporal graph convolutional networks. A physics engine is used to simulate the repair process and establish a multi-dimensional evaluation system.
It enables accurate identification and quantitative consistency verification of trainees' operations, dynamically adjusts teaching difficulty and content, improves the depth and efficiency of training, reduces training costs, and enhances learning outcomes.
Smart Images

Figure CN121922017A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of teaching and artificial intelligence technology, specifically relating to an interactive teaching method and system for aircraft composite repair based on virtual simulation. Background Technology
[0002] Composite materials for aircraft are widely used in the aerospace field due to their superior properties such as high specific strength and high specific modulus. However, composite materials are prone to damage such as delamination, cracks, and voids during manufacturing and service. High-quality repair is a crucial aspect of ensuring flight safety. Because composite material repair processes are complex, involving multiple precision steps such as grinding, layup, and curing, traditional teaching methods rely on physical teaching aids, which suffer from high costs, significant waste, limited learning scenarios, and the inability to record the entire process of student operations.
[0003] In recent years, virtual simulation technology has been introduced into the field of education. However, existing virtual teaching systems mostly focus on scene roaming or simple interaction, lacking in-depth perception and intelligent evaluation of students' operational behavior. Specifically: First, they cannot accurately identify the spatiotemporal characteristics of students' hand movements, making it difficult to judge whether the operation is standardized; second, they lack an adaptive learning strategy adjustment mechanism, making it impossible to dynamically adjust the teaching difficulty according to the students' mastery level; and third, they lack an intuitive demonstration of the physical evolution process during repair (such as resin flow and residual stress), resulting in students being unable to learn effectively. Summary of the Invention
[0004] The purpose of this application is to provide a virtual simulation-based interactive teaching method and system for aircraft composite repair, in order to solve the problems of existing technologies being unable to identify operating procedures and adjust the teaching difficulty in a timely manner.
[0005] On the one hand, this application provides an interactive teaching method for aircraft composite material repair based on virtual simulation, including: A virtual simulation environment for composite material damage in aircraft is constructed, which includes a 3D aircraft model library, composite laminate structure models, a damage morphology database, and a virtual repair toolset. When trainees operate in the virtual simulation environment of aircraft composite material damage, deep learning algorithms are used to identify and analyze the trainees' operations in real time, establish a mapping relationship between the feature vector of the operation behavior and the standard repair process, and obtain consistency verification data. Based on the consistency verification data and the student's historical operation data, a reinforcement learning algorithm is used to dynamically generate a personalized learning strategy, and the personalized learning strategy is transmitted to a pre-specified device to achieve adaptive adjustment of teaching difficulty and content.
[0006] Furthermore, the method also includes: The physics engine is used to simulate the material mechanical behavior during the composite material repair process, including physical phenomena such as resin flow, curing shrinkage and / or residual stress distribution. A multi-dimensional evaluation system was constructed to comprehensively assess trainees’ performance from three dimensions: operational standardization, process integrity, and repair quality prediction, and a visual feedback report was generated.
[0007] Furthermore, deep learning algorithms are used to perform real-time recognition and analysis of student actions, including: The operation time-space graph is obtained by establishing the positions of the student's hand joints as nodes and constructing the skeletal connections and temporal connections between joints. A deep learning algorithm is used to identify the spatiotemporal graph of the operation to determine the student's operation identification result; the deep learning algorithm is set as a spatiotemporal graph convolutional network.
[0008] Furthermore, a mapping relationship is established between operational behavior feature vectors and standard repair processes to obtain consistency verification data, including: The student operation recognition results are constructed into an operation behavior feature vector in chronological order, and the cosine similarity between the operation behavior feature vector and the standard repair process is calculated to obtain consistency verification data; the standard repair process includes reference operation behaviors arranged in sequence.
[0009] Furthermore, the real-time recognition and analysis of student actions using deep learning algorithms also includes: A spatiotemporal graph convolutional network is constructed, and the network parameters of the spatiotemporal graph convolutional network are initialized and encoded to obtain the network parameter encoding; Obtain the historical operation spatiotemporal graph and the corresponding manual operation labels, and obtain the loss function value encoded for each network parameter based on the historical operation spatiotemporal graph and the corresponding manual operation labels; Based on the loss function value of the network parameter encoding, obtain the best and worst parameter encodings among all network parameter encodings; Based on the optimal parameter encoding, an adaptive encoding displacement strategy is used to perform an oscillating back-and-forth search on the network parameter encoding to obtain the first target network parameter encoding. Based on the optimal parameter encoding, a segmented hybrid strategy is used to perform a segmented hybrid search on the first target network parameter encoding to obtain the second target network parameter encoding; Based on the worst-case parameter encoding, a global development strategy is used to globally develop the second target network parameter encoding to obtain the third target network parameter encoding. Determine whether the training termination condition has been met. If so, determine the final network parameters of the spatiotemporal graph convolutional network based on the third target network parameter encoding to obtain the trained deep learning algorithm. Otherwise, based on the third target network parameter encoding, return to the step of obtaining the loss function value and perform the next training. The trained deep learning algorithm is used to identify the spatiotemporal graph of the operation to determine the student's operation identification result.
[0010] Furthermore, based on the optimal parameter encoding, an adaptive encoding shift strategy is used to perform an oscillating back-and-forth search on the network parameter encoding to obtain the first target network parameter encoding as follows: ; ; ; ; ; In the formula, For the first k During the training process, the first i Network parameter encoding, For the first i The first target network parameter encoding, i =1,2,...,NP, where NP represents the total number of network parameter codes. For the first k+ During the first training session i The oscillating back-and-forth search speed of each network parameter encoding For the first k During the training process, the first i The oscillating back-and-forth search speed of each network parameter encoding For inertial weights, A random direction control factor that is either 0 or 1. As the first learning factor, As the second learning factor, The first random number between (0,1) The second random number between (0,1) for The historical best value, Encode the optimal parameters. Pi The oscillation factor is a random oscillation between [0,1]. This represents the maximum value of the inertia weight. This represents the minimum value of the inertia weight. This is the first weighted control coefficient. Here, is the second weighting control coefficient, and exp is an exponential function with the natural constant e as its base. The maximum value of the learning factor. K represents the minimum learning factor, and K represents the maximum number of training iterations.
[0011] Further, based on the optimal parameter encoding, a segmented hybrid search is performed on the first target network parameter encoding using a segmented hybrid strategy to obtain the second target network parameter encoding, including: Obtain the loss function value corresponding to the first target network parameter encoding, and determine the loss degree ranking of each first target network parameter encoding in descending order based on the loss function value corresponding to the first target network parameter encoding. The loss ranking is normalized to the [0,1] interval to obtain the loss position parameter; Based on the optimal parameter encoding and the loss position parameter, a segmented hybrid search is performed on the first target network parameter encoding to obtain the second target network parameter encoding: ; ; In the formula, For the first k During the training process, the first j The historical best value of the first target network parameter encoding. For the first j The second target network parameter encoding, For the location parameter of the loss degree, Encode the optimal parameters. As a disturbance factor, The first normally distributed random number is generated according to a normal distribution. The gravitational coefficient, For random matching with the first j The first target network parameter encoding is different from the historical best value of other first target network parameter encodings. The second normally distributed random number is generated according to the normal distribution. Reverse learning encoding for optimal parameter encoding, This is the basic value of the gravitational coefficient. For natural parameters, This is the gravitational control coefficient. For the first j The Cartesian distance between each first target network parameter encoding and its corresponding other first target network parameter encodings.
[0012] Furthermore, based on the worst-case parameter encoding, a global development strategy is used to globally develop the second target network parameter encoding, resulting in the third target network parameter encoding as follows: ; ; ; In the formula, Encoding the nth second target network parameter during the kth training iteration. Encode the nth third target network parameter. For random numbers that follow a normal distribution, Encode the worst-case parameter. The third random number between (0,1) For Levi's flight factor, The fourth random number between (0,1) The fifth random number between (0,1) For intermediate parameters, This is an adjustable coefficient. This is the symbol for the gamma function.
[0013] Furthermore, based on the consistency verification data and the student's historical operation data, a reinforcement learning algorithm is used to dynamically generate a personalized learning strategy, including: Based on the consistency verification data and the student's historical operation data, a state space is constructed; Based on the state space, a proximal policy optimization algorithm is used to dynamically generate personalized learning strategies.
[0014] On the other hand, this application provides a virtual simulation-based interactive teaching system for aircraft composite repair, comprising: The virtual simulation module is used to build a virtual simulation environment for damage to composite materials in aircraft. The virtual simulation environment includes a 3D aircraft model library, composite laminate structure models, a damage morphology database, and a virtual repair toolset. The consistency verification module is used to identify and analyze the trainee's operations in real time through deep learning algorithms when the trainee operates in the virtual simulation environment of composite material damage of the aircraft, establish the mapping relationship between the operation behavior feature vector and the standard repair process, and obtain consistency verification data. The teaching feedback module is used to dynamically generate personalized learning strategies based on the consistency verification data and the student's historical operation data using reinforcement learning algorithms, and transmit the personalized learning strategies to a pre-specified device to achieve adaptive adjustment of teaching difficulty and content.
[0015] The beneficial effects of this application are as follows: This application provides a virtual simulation-based interactive teaching method and system for aircraft composite material repair. By constructing a virtual simulation environment for aircraft composite material damage, trainees can intuitively understand the composite material repair process and achieve virtual training, thus enhancing the depth and breadth of training. During operation in the virtual simulation environment, a spatiotemporal graph convolutional network is used to identify the spatiotemporal graph of the trainee's hand joints, accurately capturing the spatiotemporal features of the operational movements. Combined with cosine similarity calculation, quantitative consistency verification between the trainee's operation and the standard process flow is achieved, ensuring objective and accurate evaluation results. Finally, personalized learning strategies are dynamically generated based on reinforcement learning algorithms, enabling adaptive adjustment of teaching content and difficulty, thereby improving training efficiency. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0017] Figure 1 A flowchart illustrating an interactive teaching method for aircraft composite material repair based on virtual simulation, provided in this application; Figure 2 A schematic diagram of a virtual simulation-based interactive teaching system for aircraft composite repair provided in this application; Among them, 201-Virtual simulation module; 202-Conformance verification module; 203-Teaching feedback module.
[0018] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0019] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0020] The embodiments of this application are described in detail below with reference to the accompanying drawings.
[0021] like Figure 1 As shown in the figure, this application provides an interactive teaching method for aircraft composite material repair based on virtual simulation, including: S101. Construct a virtual simulation environment for damage to composite materials in aircraft. The virtual simulation environment includes a three-dimensional aircraft model library, a composite laminate structure model, a damage morphology database, and a virtual repair toolset. The virtual simulation environment is the foundational platform of the entire interactive teaching system, and its construction quality directly affects the authenticity and effectiveness of the training results.
[0022] Parametric modeling techniques can be used to construct digital twin models of composite material structures for aircraft. The core idea of parametric modeling is to abstract the model's geometry, material properties, boundary conditions, and other elements into adjustable parameters, driving model changes through these parameters. In practice, parametric design in CATIA software is used to establish a parametric model of the composite laminate. Parameters include: layup sequence (e.g., [0 / 45 / -45 / 90]s), single-layer thickness (e.g., 0.125mm), fiber volume content (e.g., 60%), and resin system type (e.g., epoxy resin). Through parametric modeling, composite material structure models with different configurations can be quickly generated to meet diverse training needs.
[0023] Then, a database of typical damage morphologies was established. The database includes various damage types such as impact damage, delamination, debonding, cracks, and hole edge damage, with each type further subdivided into different severity levels. Damage data comes from three sources: first, damage detection data from real aircraft structures, obtained through ultrasonic testing, X-ray inspection, and other methods; second, accelerated aging test data from laboratories, obtained by simulating service environments; and third, numerical simulation data, used to predict damage evolution through finite element analysis. The database is stored using a relational database management system (such as MySQL) and supports retrieval by damage type, location, size, and other attributes.
[0024] Secondly, a virtual repair toolset was developed. This toolset includes damage detection tools (ultrasonic flaw detectors, thermal imagers, impact detectors, etc.), surface treatment tools (sanders, cleaners, release agents, etc.), repair materials (prepregs, resins, honeycomb cores, etc.), and curing equipment (autoclaves, ovens, vacuum bags, etc.). Each tool can be built with a detailed 3D model and operational logic, allowing trainees to operate them in a virtual environment as if using real tools. The tool's operational feedback is calculated through a physics engine; for example, sanding produces a virtual dust effect, and resin coating simulates flow and wetting processes.
[0025] Finally, a virtual simulation of nondestructive testing methods is implemented. Taking ultrasonic testing as an example, the system simulates the propagation process of ultrasonic waves in composite materials, considering the influence of the material's anisotropy on sound velocity and attenuation. Trainees can adjust parameters such as probe frequency, angle, and scanning speed to observe the test results under different parameter settings. The system also simulates interference factors that may be encountered in actual testing, such as surface roughness and coupling agent thickness, improving the realism of the training.
[0026] After constructing the virtual simulation environment for aircraft composite material damage, trainees can select the corresponding 3D aircraft model, composite laminate structure, and damage morphology from the 3D aircraft model library, composite laminate structure model, and damage morphology database, respectively. It's important to note that these data must adhere to pre-set combination rules. For example, aircraft A can only be combined with composite laminate structures b and c; therefore, after selecting A, the selection must be made from b and c. By constructing the virtual simulation environment for aircraft composite material damage, trainees can also select models and loss morphologies according to pre-defined learning strategies, and centrally select and operate tools through the virtual repair tool, thus completing the interactive learning process.
[0027] S102. When trainees operate in the virtual simulation environment of composite material damage of the aircraft, the deep learning algorithm is used to identify and analyze the trainees' operations in real time, establish the mapping relationship between the operation behavior feature vector and the standard repair process, and obtain consistency verification data. Accurately identifying students' operational behaviors is a key prerequisite for achieving intelligent teaching. This application uses a spatiotemporal graph convolutional network (ST-GCN) to achieve this goal.
[0028] Spatiotemporal graph convolutional networks (SPCRCs) are a deep learning architecture specifically designed for processing spatiotemporal sequence data, achieving significant results in action recognition. This application applies SPCRC to action recognition, specifically as follows: First, the joint position information of the trainee's hand is collected using a data glove (sensors are set to collect the displacement of each joint; for example, during the movement of the data glove, the corresponding virtual hand in the virtual environment is operated synchronously, thus achieving interaction) or a depth camera, constructing a spatiotemporal graph representation of the hand skeleton. The nodes of the spatiotemporal graph correspond to 21 joints of the hand (5 joints per finger, totaling 15, plus 6 key points of the palm), and the node features include 3D coordinates, velocity, acceleration, etc. The edges of the spatiotemporal graph are divided into two categories: spatial edges connect adjacent joints at the same time, reflecting the skeletal structure of the hand; temporal edges connect the same joint at adjacent time points, reflecting the temporal continuity of the action.
[0029] The network architecture employs a stacked multi-layer graph convolutional architecture, with each layer containing both spatial and temporal branches. Spatial graph convolutions aggregate information at the graph structure level, extracting spatial features of hand poses at the same moment; temporal convolutions perform convolution operations along the temporal dimension, capturing the temporal features of the action. Through this multi-layer stacking, the network learns feature representations ranging from low-level joint motion to high-level semantic actions. To focus on key operational steps, an attention mechanism is introduced after the graph convolutional layers, adaptively assigning weights to different joints and time frames to highlight features relevant to the current operation.
[0030] The model is trained using supervised learning, with training data derived from standard operation demonstrations by senior engineers and trainees' historical operation records. Annotated information includes operation type (e.g., sanding, applying adhesive, layup), operation quality (correct / incorrect), and operation stage (preparation / execution / finishing). The trained model can recognize the trainee's current operation in real time and determine whether it conforms to the standard process flow. The recognition results are output as feature vectors for subsequent adaptive teaching decisions.
[0031] S103. Based on the consistency verification data and the student's historical operation data, a reinforcement learning algorithm is used to dynamically generate a personalized learning strategy, and the personalized learning strategy is transmitted to a pre-specified device to achieve adaptive adjustment of teaching difficulty and content.
[0032] Personalized instruction is key to improving training efficiency. This application employs the Proximal Policy Optimization (PPO) algorithm to dynamically adjust the content and difficulty of instruction.
[0033] Reinforcement learning is a machine learning method that learns optimal policies through interaction with the environment. This application models the teaching process as a Markov Decision Process (MDP), defined as a tuple (S, A, P, R, γ), where S is the state space, A is the action space, P is the state transition probability, R is the reward function, and γ is the discount factor.
[0034] The state space S includes: a student knowledge mastery vector, with the same dimension as the number of knowledge points, where each element represents the degree of mastery of the corresponding knowledge point (a continuous value between 0 and 1); an operation proficiency matrix, recording the student's proficiency in various operation types; and learning progress indicators, including the number of completed training courses, cumulative training time, and average score. State information is updated based on the student's historical operation data and real-time performance.
[0035] Action space A includes: content selection (choosing the next knowledge point or training scenario to learn from the knowledge base); difficulty level adjustment (adjusting the difficulty coefficient of the current training, such as damage complexity, time limit, accuracy requirements, etc.); and training scenario switching (switching to different aircraft parts or damage types for training). Action selection is output by the policy network, which takes the current state as input and outputs the probability distribution of each action.
[0036] The reward function R comprehensively considers three dimensions: improved learning efficiency, measured by the change in knowledge mastery before and after training; knowledge retention rate, measured by interval test scores; and student satisfaction, measured by students' subjective evaluations. The design of the reward function follows the principle of moderate challenge, meaning that the reward is maximized when students are given tasks slightly above their current ability level, avoiding tasks that are too easy or too difficult.
[0037] The core idea of the PPO algorithm is to introduce constraints on the policy gradient method, limiting the magnitude of each policy update to ensure training stability. In its implementation, both the policy network and the value network employ a multilayer perceptron structure, optimizing network parameters through extensive interaction with the environment. After training, the system can automatically select the optimal teaching strategy based on the student's real-time status, achieving truly personalized instruction.
[0038] This application utilizes virtual simulation technology to replace physical training, significantly reducing the cost of consumables such as composite material test pieces, repair materials, and specialized tools. It also avoids the wear and tear associated with real aircraft structures, reducing training costs by over 60%. The virtual simulation environment can simulate various complex damage morphologies and extreme conditions, including impact damage, delamination, debonding, cracks, and other damage types, as well as repair scenarios for different aircraft parts and under varying environmental conditions, providing trainees with comprehensive practical experience. Operations in the virtual environment eliminate real safety hazards, allowing trainees to safely perform various high-risk operations, such as high-temperature curing and chemical treatments, without worrying about personal injury or equipment damage. The reinforcement learning-based adaptive teaching engine dynamically adjusts the teaching content and difficulty based on trainees' real-time performance, enabling personalized instruction and improving training efficiency and learning outcomes.
[0039] In some embodiments, the method further includes: The physics engine is used to simulate the material mechanical behavior during the composite material repair process, including physical phenomena such as resin flow, curing shrinkage and / or residual stress distribution. A multi-dimensional evaluation system was constructed to comprehensively assess trainees’ performance from three dimensions: operational standardization, process integrity, and repair quality prediction, and a visual feedback report was generated.
[0040] High-precision physical simulation is key to ensuring the authenticity of training, while multi-dimensional evaluation is the foundation for achieving objective assessment.
[0041] The physical simulation employs a coupled solution strategy of finite element method (FEM) and computational fluid dynamics (CFD). For the resin impregnation process, a porous media flow model is established, treating the fiber reinforcement as a porous medium, with the resin flow following Darcy's law. The governing equations include the continuity equation and the momentum equation, and the boundary conditions consider the resin inlet pressure, outlet pressure, and permeability tensor of the fiber. The numerical solution uses the finite volume method, with implicit time discretization and unstructured meshes for spatial discretization.
[0042] For the curing process, a thermo-chemical-mechanical coupled model is established. The heat conduction equation describes the evolution of the temperature field, the chemical reaction kinetic equation describes the change in the degree of curing, and the mechanical equilibrium equation describes the development of stress and deformation. The three equations are interconnected through coupling terms: the exothermic chemical reaction affects the temperature field, temperature and degree of curing affect material properties, and material properties and temperature gradient affect stress distribution. The numerical solution adopts a sequential coupling method, solving the three equations sequentially at each time step, iterating until convergence.
[0043] The multi-dimensional evaluation system employs the Analytic Hierarchy Process (AHP) to determine the weights of each evaluation indicator. First, a hierarchical model is constructed: the target layer represents the comprehensive assessment of trainees' abilities; the criteria layer comprises three dimensions: operational standardization, process integrity, and repair quality prediction; and the indicator layer contains the specific indicators for each dimension. Then, a judgment matrix is constructed through expert questionnaires, and the weight vectors of each indicator are calculated. The final weight allocation is: operational standardization 30%, process integrity 40%, and repair quality prediction 30%.
[0044] The evaluation of operational standardization employs a rule-matching method, comparing trainees' operational sequences with standard process flows and calculating scores for each indicator. The evaluation of process integrity uses a knowledge graph reasoning method, judging the rationality and completeness of trainees' repair solutions. The evaluation of repair quality prediction uses a convolutional neural network, taking a 3D model of the virtual repair result as input and outputting a quality level prediction (Excellent, Good, Average, Poor). The comprehensive evaluation results are presented to trainees in the form of radar charts and detailed reports, helping them understand their strengths and areas for improvement.
[0045] The physics engine realistically simulates the mechanical behavior of materials during the repair process, providing trainees with tactile feedback and a visual experience close to real-world operation, thus enhancing the immersion and effectiveness of the training. The multi-dimensional evaluation system enables fine-grained analysis and quantitative assessment of trainees' operations, ensuring objective, fair, and traceable results, providing a reliable basis for trainee competency certification.
[0046] In some embodiments, deep learning algorithms are used to identify and analyze student actions in real time, including: The operation time-space graph is obtained by establishing the positions of the student's hand joints as nodes and constructing the skeletal connections and temporal connections between joints. A deep learning algorithm is used to identify the spatiotemporal graph of the operation to determine the student's operation identification result; the deep learning algorithm is set as a spatiotemporal graph convolutional network.
[0047] In some embodiments, a mapping relationship is established between operational behavior feature vectors and standard repair processes to obtain consistency verification data, including: The student operation recognition results are constructed into an operation behavior feature vector in chronological order, and the cosine similarity between the operation behavior feature vector and the standard repair process is calculated to obtain consistency verification data; the standard repair process includes reference operation behaviors arranged in sequence.
[0048] In some embodiments, the real-time identification and analysis of student actions using deep learning algorithms further includes: A spatiotemporal graph convolutional network is constructed, and the network parameters of the spatiotemporal graph convolutional network are initialized and encoded to obtain network parameter encodings. For example, random initialization and encoding into vectors can obtain network parameter encodings. Then, multiple different network parameter encodings can be obtained repeatedly to achieve initialization.
[0049] Obtain the historical operation spatiotemporal graph and the corresponding manual operation labels, and obtain the loss function value encoded for each network parameter based on the historical operation spatiotemporal graph and the corresponding manual operation labels; such as the cross-entropy loss function value or the root mean square loss function value.
[0050] Based on the loss function value of the network parameter encoding, obtain the best and worst parameter encodings among all network parameter encodings; Based on the optimal parameter encoding, an adaptive encoding displacement strategy is used to perform an oscillating back-and-forth search on the network parameter encoding to obtain the first target network parameter encoding. Based on the optimal parameter encoding, a segmented hybrid strategy is used to perform a segmented hybrid search on the first target network parameter encoding to obtain the second target network parameter encoding; Based on the worst-case parameter encoding, a global development strategy is used to globally develop the second target network parameter encoding to obtain the third target network parameter encoding. Determine whether the training termination condition has been reached (e.g., whether the number of training iterations has been reached). If so, determine the final network parameters of the spatiotemporal graph convolutional network based on the third target network parameter encoding to obtain the trained deep learning algorithm. Otherwise, based on the third target network parameter encoding, return to the step of obtaining the loss function value and perform the next training. The trained deep learning algorithm is used to identify the spatiotemporal graph of the operation to determine the student's operation identification result.
[0051] To address the challenges of training deep learning models, this application proposes a parameter optimization method that integrates adaptive encoding shift, segmented hybrid search, and a global exploitation strategy. This method enhances local exploitation capabilities through oscillatory search, balances exploration and exploitation through a segmented hybrid strategy, and avoids getting trapped in local optima using a global exploitation strategy, significantly improving the training efficiency and recognition accuracy of spatiotemporal graph convolutional networks.
[0052] In some embodiments, based on the optimal parameter encoding, an adaptive encoding shift strategy is used to perform an oscillating back-and-forth search on the network parameter encoding to obtain the first target network parameter encoding as follows: ; ; ; ; ; In the formula, For the first k During the training process, the first i Network parameter encoding, For the first i The first target network parameter encoding, i =1,2,...,NP, where NP represents the total number of network parameter codes. For the first k+ During the first training session i The oscillating back-and-forth search speed of each network parameter encoding For the first k During the training process, the first i The oscillating back-and-forth search speed of each network parameter encoding For inertial weights, A random direction control factor that is either 0 or 1. As the first learning factor, As the second learning factor, The first random number between (0,1) The second random number between (0,1) for The historical best value, Encode the optimal parameters. Pi The oscillation factor is a random oscillation between [0,1]. The maximum value for the inertia weight can be set to 0.65; The minimum value for the inertia weight can be set to 0.40; This is the first weighted control coefficient. Here, is the second weighting control coefficient, and exp is an exponential function with the natural constant e as its base. The maximum value for the learning factor can be set to 0.5; The minimum learning factor can be set to 0.01; K is the maximum number of training iterations.
[0053] An oscillation factor is introduced into the formula, giving the parameter update direction a probability of random reversal. This mechanism allows the algorithm to oscillate back and forth around the current optimal solution, similar to a sawtooth search path, enabling a more refined scan of local regions and effectively overcoming the problem of traditional algorithms easily "skipping" extreme points during local searches. The inertia weights and learning factor adopt nonlinear adaptive changes. In the early stages of training, larger inertia weights are beneficial for global, large-scale searches; as training progresses, the weights decrease and the learning factor is adjusted, causing the algorithm to gradually shift towards refined local mining, improving convergence accuracy.
[0054] In some embodiments, based on the optimal parameter encoding, a segmented hybrid search is performed on the first target network parameter encoding using a segmented hybrid strategy to obtain the second target network parameter encoding, including: Obtain the loss function value corresponding to the first target network parameter encoding, and determine the loss degree ranking of each first target network parameter encoding in descending order based on the loss function value corresponding to the first target network parameter encoding. The loss ranking is normalized to the [0,1] interval to obtain the loss position parameter; Based on the optimal parameter encoding and the loss position parameter, a segmented hybrid search is performed on the first target network parameter encoding to obtain the second target network parameter encoding: ; ; In the formula, For the first k During the training process, the first j The historical best value of the first target network parameter encoding. For the first j The second target network parameter encoding, For the location parameter of the loss degree, Encode the optimal parameters. This is the perturbation factor, which can be set to 0.01; The first normally distributed random number is generated according to a normal distribution. The gravitational coefficient, For random matching with the firstj The first target network parameter encoding is different from the historical best value of other first target network parameter encodings. The second normally distributed random number is generated according to the normal distribution. Reverse learning encoding for optimal parameter encoding, The basic value for the gravitational coefficient can be set as a constant term between [0.01, 0.1]. For natural parameters, This is the gravitational control coefficient, which can be set as a constant term between [0.1, 10]. For the first j The Cartesian distance between each first target network parameter encoding and its corresponding other first target network parameter encodings.
[0055] A differentiated update strategy is adopted for network parameter encodings of different quality. For poorly performing encodings, direct learning from the best encoding accelerates their convergence; for moderately performing encodings, a gravity mechanism and random perturbation are introduced to maintain population diversity and prevent premature convergence; for excellent-performing encodings, a back-learning strategy is used to explore the symmetric position of the best encoding, which may discover a better solution. By processing in segments, the convergence speed of the algorithm is guaranteed while maintaining population diversity, solving the problem of search stagnation caused by "all encodings converging" in traditional algorithms.
[0056] In some embodiments, based on the worst-case parameter encoding, a global development strategy is used to globally develop the second target network parameter encoding to obtain the third target network parameter encoding as follows: ; ; ; In the formula, Encoding the nth second target network parameter during the kth training iteration. Encode the nth third target network parameter. For random numbers that follow a normal distribution, Encode the worst-case parameter. The third random number between (0,1) For Levi's flight factor, The fourth random number between (0,1) The fifth random number between (0,1) For intermediate parameters, This is an adjustable coefficient and can be set to 1.5; This is the symbol for the gamma function.
[0057] Traditional algorithms often focus only on the best individual, ignoring information about the "worst individual." This implementation uses the worst parameter encoding as a reference to guide other individuals away from that area, thus guiding the population to search for more promising regions. A Levy flight strategy is introduced, characterized by alternating short-distance walks and long-distance jumps. When the model gets trapped in a local optimum, the long jump mechanism of Levy flight helps the parameters escape the local trap instantly, greatly enhancing the algorithm's global search capability.
[0058] Optionally, a greedy strategy can be introduced during the global development process to improve the training speed of the algorithm. After each encoding search, out-of-bounds handling can be performed on the encoding to ensure the effectiveness of the algorithm.
[0059] In some embodiments, a personalized learning strategy is dynamically generated using a reinforcement learning algorithm based on the consistency verification data and the student's historical operation data, including: Based on the consistency verification data and the student's historical operation data, a state space is constructed; Based on the state space, a proximal policy optimization algorithm is used to dynamically generate personalized learning strategies.
[0060] In this interactive learning process, the goal of the reinforcement learning agent is to generate the optimal sequence of teaching strategies, enabling learners to master aircraft composite repair skills through the shortest learning path. Therefore, the design of the reward function needs to comprehensively consider operational consistency, learning efficiency, skill mastery, and the matching degree of teaching difficulty.
[0061] Set at time t The state observed by the agent is s t The teaching strategies and actions adopted are as follows a t (Teaching strategy actions can be set as codes corresponding to specific courses. All courses are coded in natural numerical order to form an action space. Selecting a number indicates that the student will learn the course corresponding to that number. Course instruction videos or images will be distributed, and the corresponding models, materials, and damage morphologies will be invoked in the virtual environment, allowing students to learn and interact.) The feedback status after the student executes this strategy is as follows: s t+1 Then the total reward function R t Defined as: ; in, , , , , These are hyperparameters used to balance the weight of each reward in the total reward, and their sum is 1. The specific calculation formulas for each sub-indicator are as follows: 1. Operational consistency reward ; This award aims to encourage teaching strategies that guide trainees to produce operations that conform to standard process procedures. Based on consistency verification data (i.e., cosine similarity), the formula is: ; in, Let Δt be the cosine similarity between the student's operational behavior feature vector at time t and the standard repair process (i.e., consistency verification data); Δt is the time step. The preset consistency qualification threshold (e.g., 0.85); α , β All are weighting coefficients; sgn(·) is the sign function, which takes 1 when the value in the parentheses is greater than 0, -1 when it is less than 0, and 0 when it is equal to 0.
[0062] 2. Skill mastery progress rewards ; This award assesses the trainee's change in mastery of each sub-task of the repair process, using the following formula: ; Where N is the total number of sub-tasks in the repair process; For state The next student in i Mastery score for each subtask (value range [0,1]); The threshold is used to determine the degree; I(·) is the indicator function, which takes the value 1 when the condition is met and 0 otherwise; Σ is the summation symbol.
[0063] Optionally, the student in the first i Mastery scores for individual sub-tasks can be obtained through a comprehensive evaluation model that integrates multi-source data. First, the learner's operational sequences, standardization indicators, and / or repair quality prediction results in a virtual simulation environment are collected. Then, using Bayesian Knowledge Tracking (BKT) or deep learning algorithms, the operational results are used as input to dynamically update the learner's posterior probability of mastery of the skill. This score comprehensively considers operational accuracy, process standardization, and result quality; it is a continuous value normalized to the [0,1] interval, reflecting the learner's current knowledge state in real time and serving as a key input to the reinforcement learning state space. Before using deep learning algorithms for evaluation, historical data and teacher ratings of that historical data can be used as training data to achieve training.
[0064] 3. Learning efficiency rewards ; To prevent the agent from generating lengthy and inefficient learning paths, constraints need to be placed on the learning time and number of steps, as shown in the formula: ; This represents the current cumulative learning time or number of steps taken. γ represents the expected standard learning time. γ is the efficiency penalty coefficient.
[0065] 4. Difficulty Matching Rewards ; Based on adaptive adjustment of teaching difficulty, this reward encourages the agent to generate personalized content that matches the learner's current ability, as shown in the formula: ; Teaching strategies a t The corresponding difficulty level can be preset for each teaching strategy; σ represents the current ability level of a student, assessed based on their historical data; σ is the scale parameter; exp(·) is an exponential function with the natural constant e as its base; |·| is the absolute value symbol.
[0066] Optionally, a student's current skill level can be obtained through a comprehensive quantitative assessment of historical learning data. The system first collects multi-dimensional features such as the student's historical operational consistency score, the average mastery of each sub-task, and the distribution of error types. Then, using a weighted scoring model (such as a deep learning model) or a pre-trained regression prediction model, the discrete historical data is mapped to a unified skill index. This index not only counts the student's past success rate but also highlights recent performance through time decay weighting, thus accurately reflecting the student's real-time proficiency in composite material repair skills and providing a benchmark for matching the difficulty of teaching strategies. The deep learning model can also be trained based on corresponding historical data and the teacher's evaluation of the student's skill level on that historical data, thereby achieving automated assessment.
[0067] 5. Penalties for violations ; For serious process errors made by trainees during operation, penalties are imposed based on the identified erroneous operation results, using the following formula: ; Let be the set of violations detected at time t. For the first j The severity weight of each type of violation is assigned, with different severity weights pre-defined for different violations; ∈ represents the "belongs to" symbol.
[0068] It is worth noting that the reward can be any of the rewards from 1 to 4, or it can be set to the reward value in other existing technologies, which can also achieve the effect of strategy customization.
[0069] like Figure 2 As shown in the embodiments of this application, an interactive teaching system for aircraft composite material repair based on virtual simulation is also provided, including: The virtual simulation module 201 is used to construct a virtual simulation environment for damage to composite materials in aircraft. The virtual simulation environment includes a three-dimensional aircraft model library, a composite laminate structure model, a damage morphology database, and a virtual repair toolset. The consistency verification module 202 is used to identify and analyze the trainee's operation in real time through deep learning algorithms when the trainee operates in the virtual simulation environment of aircraft composite material damage, establish the mapping relationship between the operation behavior feature vector and the standard repair process, and obtain consistency verification data. The teaching feedback module 203 is used to dynamically generate personalized learning strategies based on the consistency verification data and the student's historical operation data using a reinforcement learning algorithm, and transmit the personalized learning strategies to a pre-specified device to achieve adaptive adjustment of teaching difficulty and content.
[0070] Optionally, the system may also include: The repair process simulation module is used to simulate the material mechanical behavior during the repair process of composite materials using a physics engine, including physical phenomena such as resin flow, curing shrinkage and / or residual stress distribution. The teaching evaluation module is used to construct a multi-dimensional evaluation system, which comprehensively evaluates trainees’ performance from three dimensions: operational standardization, process integrity, and repair quality prediction, and generates a visual feedback report.
[0071] The above-mentioned interactive teaching system for aircraft composite material repair based on virtual simulation can execute any of the method embodiments, and its principles and beneficial effects are similar, so they will not be described again here.
[0072] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0073] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0074] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0075] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.
[0076] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A virtual simulation-based interactive teaching method for aircraft composite material repair, characterized in that, include: A virtual simulation environment for composite material damage in aircraft is constructed, which includes a 3D aircraft model library, composite laminate structure models, a damage morphology database, and a virtual repair toolset. When trainees operate in the virtual simulation environment of aircraft composite material damage, deep learning algorithms are used to identify and analyze the trainees' operations in real time, establish a mapping relationship between the feature vector of the operation behavior and the standard repair process, and obtain consistency verification data. Based on the consistency verification data and the student's historical operation data, a reinforcement learning algorithm is used to dynamically generate a personalized learning strategy, and the personalized learning strategy is transmitted to a pre-specified device to achieve adaptive adjustment of teaching difficulty and content.
2. The interactive teaching method for aircraft composite material repair based on virtual simulation according to claim 1, characterized in that, Also includes: The physics engine is used to simulate the material mechanical behavior during the composite material repair process, including physical phenomena such as resin flow, curing shrinkage and / or residual stress distribution. A multi-dimensional evaluation system was constructed to comprehensively assess trainees’ performance from three dimensions: operational standardization, process integrity, and repair quality prediction, and a visual feedback report was generated.
3. The interactive teaching method for aircraft composite material repair based on virtual simulation according to claim 1, characterized in that, Real-time recognition and analysis of student actions are performed using deep learning algorithms, including: The operation time-space graph is obtained by establishing the positions of the student's hand joints as nodes and constructing the skeletal connections and temporal connections between joints. A deep learning algorithm is used to identify the spatiotemporal graph of the operation to determine the student's operation identification result; the deep learning algorithm is set as a spatiotemporal graph convolutional network.
4. The interactive teaching method for aircraft composite material repair based on virtual simulation according to claim 3, characterized in that, Establish a mapping relationship between operational behavior feature vectors and standard repair processes to obtain consistency verification data, including: The student operation recognition results are constructed into an operation behavior feature vector in chronological order, and the cosine similarity between the operation behavior feature vector and the standard repair process is calculated to obtain consistency verification data; the standard repair process includes reference operation behaviors arranged in sequence.
5. The interactive teaching method for aircraft composite material repair based on virtual simulation according to claim 3, characterized in that, Real-time recognition and analysis of student actions using deep learning algorithms also includes: A spatiotemporal graph convolutional network is constructed, and the network parameters of the spatiotemporal graph convolutional network are initialized and encoded to obtain the network parameter encoding; Obtain the historical operation spatiotemporal graph and the corresponding manual operation labels, and obtain the loss function value encoded for each network parameter based on the historical operation spatiotemporal graph and the corresponding manual operation labels; Based on the loss function value of the network parameter encoding, obtain the best and worst parameter encodings among all network parameter encodings; Based on the optimal parameter encoding, an adaptive encoding displacement strategy is used to perform an oscillating back-and-forth search on the network parameter encoding to obtain the first target network parameter encoding. Based on the optimal parameter encoding, a segmented hybrid strategy is used to perform a segmented hybrid search on the first target network parameter encoding to obtain the second target network parameter encoding; Based on the worst-case parameter encoding, a global development strategy is used to globally develop the second target network parameter encoding to obtain the third target network parameter encoding. Determine whether the training termination condition has been met. If so, determine the final network parameters of the spatiotemporal graph convolutional network based on the third target network parameter encoding to obtain the trained deep learning algorithm. Otherwise, based on the third target network parameter encoding, return to the step of obtaining the loss function value and perform the next training. The trained deep learning algorithm is used to identify the spatiotemporal graph of the operation to determine the student's operation identification result.
6. The interactive teaching method for aircraft composite material repair based on virtual simulation according to claim 5, characterized in that, Based on the optimal parameter encoding, an adaptive encoding shift strategy is used to perform an oscillating back-and-forth search on the network parameter encoding to obtain the first target network parameter encoding as follows: ; ; ; ; ; In the formula, For the first k During the training process, the first i Network parameter encoding, For the first i The first target network parameter encoding, i =1,2,...,NP, where NP represents the total number of network parameter codes. For the first k+ During the first training session i The oscillating back-and-forth search speed of each network parameter encoding For the first k During the training process, the first i The oscillating back-and-forth search speed of each network parameter encoding For inertial weights, A random direction control factor that is either 0 or 1. As the first learning factor, As the second learning factor, The first random number between (0,1) The second random number between (0,1) for The historical best value, Encode the optimal parameters. Pi The oscillation factor is a random value between [0,1]. This represents the maximum value of the inertia weight. This represents the minimum value of the inertia weight. This is the first weighted control coefficient. Here, is the second weighting control coefficient, and exp is an exponential function with the natural constant e as its base. The maximum value of the learning factor. K represents the minimum learning factor, and K represents the maximum number of training iterations.
7. The interactive teaching method for aircraft composite material repair based on virtual simulation according to claim 6, characterized in that, Based on the optimal parameter encoding, a segmented hybrid search is performed on the first target network parameter encoding using a segmented hybrid strategy to obtain the second target network parameter encoding, including: Obtain the loss function value corresponding to the first target network parameter encoding, and determine the loss degree ranking of each first target network parameter encoding in descending order based on the loss function value corresponding to the first target network parameter encoding. The loss ranking is normalized to the [0,1] interval to obtain the loss position parameter; Based on the optimal parameter encoding and the loss position parameter, a segmented hybrid search is performed on the first target network parameter encoding to obtain the second target network parameter encoding: ; ; In the formula, For the first k During the training process, the first j The historical best value of the first target network parameter encoding. For the first j The second target network parameter encoding, For the location parameter of the loss degree, Encode the optimal parameters. As a disturbance factor, The first normally distributed random number is generated according to a normal distribution. The gravitational coefficient, For random matching with the first j The first target network parameter encoding differs from the historical best values of other first target network parameter encodings. The second normally distributed random number is generated according to the normal distribution. Reverse learning encoding for optimal parameter encoding, This is the basic value of the gravitational coefficient. For natural parameters, This is the gravitational control coefficient. For the first j The Cartesian distance between each first target network parameter encoding and its corresponding other first target network parameter encodings.
8. The interactive teaching method for aircraft composite material repair based on virtual simulation according to claim 7, characterized in that, Based on the worst-case parameter encoding, a global development strategy is used to globally develop the second target network parameter encoding to obtain the third target network parameter encoding: ; ; ; In the formula, Encoding the nth second target network parameter during the kth training iteration. Encode the nth third target network parameter. For random numbers that follow a normal distribution, Encode the worst-case parameter. The third random number between (0,1) For Levi's flight factor, The fourth random number between (0,1) The fifth random number between (0,1) For intermediate parameters, This is an adjustable coefficient. This is the symbol for the gamma function.
9. The interactive teaching method for aircraft composite material repair based on virtual simulation according to claim 1, characterized in that, Based on the consistency verification data and the student's historical operation data, a reinforcement learning algorithm is used to dynamically generate a personalized learning strategy, including: Based on the consistency verification data and the student's historical operation data, a state space is constructed; Based on the state space, a proximal policy optimization algorithm is used to dynamically generate personalized learning strategies.
10. A virtual simulation-based interactive teaching system for aircraft composite material repair, characterized in that, include: The virtual simulation module is used to build a virtual simulation environment for damage to composite materials in aircraft. The virtual simulation environment includes a 3D aircraft model library, composite laminate structure models, a damage morphology database, and a virtual repair toolset. The consistency verification module is used to identify and analyze the trainee's operations in real time through deep learning algorithms when the trainee operates in the virtual simulation environment of composite material damage of the aircraft, establish the mapping relationship between the operation behavior feature vector and the standard repair process, and obtain consistency verification data. The teaching feedback module is used to dynamically generate personalized learning strategies based on the consistency verification data and the student's historical operation data using reinforcement learning algorithms, and transmit the personalized learning strategies to a pre-specified device to achieve adaptive adjustment of teaching difficulty and content.