A method and system for aircraft skin viewpoint planning based on deep reinforcement learning
By modeling aircraft skin viewpoint planning as a Markov decision process, adopting deep reinforcement learning methods, designing a custom reward function and Gaussian mixture model, and generating a six-degree-of-freedom viewpoint planning strategy, the problems of low efficiency and poor adaptability of traditional methods in aircraft skin measurement are solved, and efficient viewpoint planning is achieved.
Patent Information
- Application Number
- CN202511020983.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-07-24
AI Technical Summary
Traditional viewpoint planning methods are difficult to adapt to the complex morphology and efficient measurement requirements of large components such as aircraft skins. They have the problems of strong reliance on manual experience and low efficiency. In addition, existing deep learning-based methods have difficulty in fully collecting data on large components.
The viewpoint planning problem is modeled as a Markov decision process. A deep reinforcement learning method based on the actor-critic (AC) network architecture is adopted. By calculating the state space and action space, designing a custom reward function and training termination conditions, and combining progressive training with a Gaussian mixture model, a viewpoint planning strategy for a large-scale continuous space with six degrees of freedom is generated.
It realizes the mathematical description of viewpoint planning in complex surface environments, improves the global optimization capability of viewpoint sequences, adapts to the changes in the topological structure of skin surfaces of different models, improves detection coverage and efficiency, and solves the problems of subjectivity and low efficiency of traditional methods.
Smart Images

Figure CN120524592B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of robot vision technology, and in particular relates to an aircraft skin viewpoint planning method and system based on deep reinforcement learning. Background Art
[0002] As a core component of the aircraft structure, the aircraft skin undertakes multiple functions of maintaining aerodynamic shape, load transfer and lightweight design. With the rapid development of aerospace technology, the size, complexity and material properties of the skin are constantly breaking through, and there is also a higher demand for the measurement of the skin. Traditional measurement relies on manual experience to scan the components step by step. It is not only difficult to adapt to the geometric characteristics and measurement requirements of complex morphology, but also requires repeated confirmation to prevent errors and omissions, which greatly affects production efficiency. Therefore, the present invention designs a six-degree-of-freedom large-scale continuous spatial viewpoint planning method for autonomous measurement of aircraft skin, which can produce the optimal viewpoint set to ensure the efficiency and coverage of skin measurement.
[0003] Viewpoint planning (VPP) aims to maximize task objectives (such as coverage, reconstruction accuracy, or target detection rate) for measuring specific objects or target areas by optimizing the sensor's position, pose, or path. To reduce the computational complexity of viewpoint planning and adapt to dynamic environments, the currently mainstream viewpoint planning strategy is NBV (Next Best View). This strategy selects the optimal viewpoint that maximizes information gain based on currently available information. However, classic rule-based NBV strategies often rely heavily on manual experience to set viewpoint criteria and manually define action spaces. For example, these strategies use the uncertainty of neural radiance fields to calculate the next viewpoint. While this approach addresses self-occlusion to some extent, the hemispherical viewpoints limit their applicability. Designed deep learning-based viewpoint planning algorithms are trained on manually measured viewpoint sequences, but they are prone to overfitting to specific scenarios. To avoid this reliance on manual experience, some studies have adopted deep reinforcement learning-based approaches that predict the optimal viewpoint by interacting with the environment based on observation sequences to obtain rewards. However, to avoid complex motion spaces, most systems use a predefined spherical space. For example, viewpoints are pre-generated from a sphere and then filtered. This minimizes manual design, but the limited motion space makes it difficult to fully capture target data. Dynamically varying the diameter of the sphere's motion space increases the flexibility of the scanning robot, but still leaves many locations unscannable due to self-occlusion.
[0004] Currently, mainstream viewpoint planning methods, either through the generation of a large number of discrete viewpoints for screening or through fixed motion space constraints, are difficult to apply to large components such as aircraft skins. A viewpoint planning method suitable for such large components is one of the key technologies for improving autonomous aerospace operations. Summary of the Invention
[0005] In response to the above technical problems, the present invention provides an aircraft skin viewpoint planning method and system based on deep reinforcement learning.
[0006] The technical solution adopted by the present invention to solve the technical problem is:
[0007] A method for aircraft skin viewpoint planning based on deep reinforcement learning, the method comprising the following steps:
[0008] S100: Model the viewpoint planning problem as a Markov decision problem, and calculate the state space and action space based on the sensor field of view constraints and the CAD model of the object being measured;
[0009] S200: Build an actor-critic AC-based network architecture, define a reward function based on the measurement target, and generate a six-degree-of-freedom simulation interaction environment required for reinforcement learning training;
[0010] S300: Define the training termination condition based on the measurement target, set the number of single theory training steps and network update parameters, and fill the experience pool;
[0011] S400: The scanner moves in space according to the network strategy, calculates the reward obtained for each action based on the reward function, obtains the learning progress of the current agent, and saves the action, state, reward and learning progress to the experience pool;
[0012] S500: Based on the updated parameters, the parameters of the AC network are updated after setting the training rounds. The Gaussian mixture model is calculated based on the learning progress to sample the tasks, a new training environment is generated, and S400 is repeated until the training termination conditions are met. The trained AC network is saved, the viewpoint planning strategy is obtained, and the six-degree-of-freedom large-scale continuous space viewpoint planning is completed.
[0013] Preferably, S100 includes:
[0014] S110: Constructing the viewpoint planning problem into a five-tuple of Markov decision processes Corresponding modeling, where Represents the state, which is the environmental information collected each time the scanner interacts with the environment; Indicates the action, that is, the action taken by the scanner each time; Represents the reward for each action taken; Represents the reward discount factor, which is used to calculate the reward discount of future actions; represents a set of transition probability functions, i.e., the probability of the scanner selecting the next action;
[0015] S120: Calculate the state space, input the CAD model of the aircraft skin, calculate the axis-aligned bounding box size of the model, record the maximum and minimum positions on the X, Y, and Z axes respectively, and the axis-aligned bounding box size is ; Define the size of a single voxel as , taking the bottom center of the model to be tested as the origin, the state space occupied by the calculation model is:
[0016] ;
[0017] in, Indicates expanding two voxel spaces outward to ensure that the bounding box still has margin at the boundary;
[0018] S130: Design a state transfer function based on the voxel occupancy state, define the voxel to be occupied, unknown, and empty, and the corresponding voxel state values are 1, 0.5, and 0 respectively. Suppose the scanner is at point Perform a scan, start from the camera center and calculate the rays along the field of view to each scanned occupied voxel, and define the coordinates of the occupied voxel center as , the maximum distance of the scanner is , the coordinates of a voxel in the state space are , then at point The state value of each voxel scanned can be calculated using the following formula:
[0019] ;
[0020] S140: Calculate the motion space based on the scanner's field of view and design a 6-DOF vector is the action parameter, Indicates that the scanner rotates along each axis without any constraints, so the range of motion is , assuming the working distance of the scanner is , combined with the size of the state space, the action range of the action space can be obtained as:
[0021] ;
[0022] S150: Setting the initial action space and state space, and selecting one tenth of the complete action space and state space as the action space and state space for initial training.
[0023] Preferably, S200 includes:
[0024] S210: Design a reinforcement learning network based on the actor-critic (AC) architecture. The network architecture consists of an actor network, a target actor network, two critic networks, and two target critic networks. The actor network consists of three hidden layers, each of which uses ReLU as the activation function. The last hidden layer uses tanh to map the action to the range (-1, 1). The critic network consists of three hidden layers, and except for the last hidden layer, which directly outputs linearly, all other hidden layers use ReLU as the activation function.
[0025] S220: Designing an objective function. In the viewpoint planning problem, a strategy needs to be learned. Maximize rewards; the goal of reinforcement learning is to maximize the expected cumulative sum of rewards , it is necessary to introduce the entropy term Expand the objective function. is a random variable The probability density function of ; in reinforcement learning, The degree of randomness is determined by the strategy Control, according to the definition of entropy, the entropy term introduced in reinforcement learning is , the objective function is:
[0026] ;
[0027] in, Indicates that in the strategy Down The distribution of Represents the entropy coefficient, which is used to control the importance of entropy;
[0028] S230: Design a reward function to calculate the reward by scanning the new addition rate and the overall coverage rate; define the overall coverage rate , single scan new rate ,in Indicates the currently occupied voxel, Indicates that new voxels have been added in this scan. Represents all occupied voxels of the model; the reward function is:
[0029] ;
[0030] in, and Indicates the weight of the two rewards. In addition, when If empty, a penalty of -1 is applied.
[0031] Preferably, S300 includes:
[0032] S310: Setting the training termination condition: ending the training when the coverage reaches 90% or when the number of training rounds reaches 5000;
[0033] S320: Set the number of single-round training steps and network update parameters. The number of single-round training steps is 30. 128 sets of data are sampled from the experience pool each time a round of action is performed. The soft update coefficient of the target network is 0.01.
[0034] S330: Fill the experience pool and calculate the viewpoint according to the greedy algorithm to interact and obtain multiple sets of action, state and reward data; first, calculate the current state with the environment origin as the initial viewpoint , then, randomly generate around the initial viewpoint From a perspective, through the benefit function Calculate the information gain for each viewpoint:
[0035] ;
[0036] in, Indicates viewpoint The uncertainty of the new information is higher, the uncertainty is lower; the action with the largest gain is selected , and calculate the reward according to the reward function , calculate the next state based on the execution action , update a set of experience data in the experience pool ,in, Indicates whether the current round has ended; the loop continues until the coverage rate reaches the target or the maximum number of steps in a single round is reached.
[0037] Preferably, S400 includes:
[0038] S410: Current status Input is sent to the Actor network, and the Actor network outputs the execution action , the scanner performs actions in the environment , calculate the reward according to the reward function;
[0039] S420: After each round of interaction, calculate the current learning progress LP based on the task and the actions in the environment:
[0040] ;
[0041] in, and Represent the weight relationship between guidance and learning progress, Controls the decay rate, Indicates the training time, represents the agent reward, represents the interaction reward of the greedy algorithm, represents the reward of the agent the last time it was trained on the task, represents the reward of the agent when training in this task;
[0042] S430: Save the action, status, reward and learning progress to the experience pool.
[0043] Preferably, S500 includes:
[0044] S510: Extract experience from the experience pool to update the Critic network. First, calculate the target of each set of sampled data. value:
[0045] ;
[0046] Minimize the loss function for the two critics respectively, , specifically:
[0047] ;
[0048] S520: Extract experience from the experience pool to update the Actor network and minimize the loss function, specifically:
[0049] ;
[0050] S530: Update the corresponding target Critic / Actor network through soft update:
[0051] ;
[0052] in, Represents the target Critic / Actor network parameters, Represents the Critic / Actor network parameters, Control the update speed;
[0053] S540: Resampling the action space at different sizes of the aircraft skin according to the current policy , Respectively represent the length, width and height of the intercepted aircraft skin different action space; for the task Rewards available and calculate learning progress , the data Save in the experience pool, obtain N groups of data from the experience pool to fit the Gaussian mixture model;
[0054] S550: Adjust the initial action space and state space size according to the fitted Gaussian mixture model, calculate the mean of LP in each Gaussian distribution, and select the distribution sample with the maximum mean , and re-sample the task according to the state space and action space of the current aircraft skin to generate a new training environment, and repeat S400 until the training termination condition is met, and the training is terminated. The trained AC network is saved, the viewpoint planning strategy is obtained, and the six-degree-of-freedom large-scale continuous space viewpoint planning is completed.
[0055] Preferably, obtaining N groups of data from the experience pool to fit the Gaussian mixture model in S540 includes:
[0056] S541: Obtain from the experience pool Group data, each group of data is , average initialization mixing weight , mean and covariance Then, calculate the posterior probability of each data:
[0057] ;
[0058] S542: Update based on posterior probability 、 、 :
[0059] ;
[0060] ;
[0061] ;
[0062] S543: Repeatedly calculate the posterior probability and update 、 、 , until the parameter change amplitude is less than the threshold, calculate the Bayesian Information Criterion BIC of the current model:
[0063] ;
[0064] in, is the maximum log-likelihood estimate of the Gaussian mixture model, Represents the number of samples used to calculate its penalty term;
[0065] S544: Repeat S541 to S543, fitting A Gaussian mixture model with different distributions is created and the one with the smallest BIC is selected as the optimal model.
[0066] An aircraft skin viewpoint planning system based on deep reinforcement learning, including a state space and action space calculation module, a network architecture building module, a parameter setting module, a training module, and a viewpoint planning module;
[0067] The state space and action space calculation module is used to model the viewpoint planning problem as a Markov decision problem and calculate the state space and action space based on the sensor field of view constraints and the CAD model of the object being measured;
[0068] The network architecture building module is used to build the actor-critic AC-based network architecture, define the reward function based on the measurement target, and generate the six-degree-of-freedom simulation interaction environment required for reinforcement learning training;
[0069] The parameter setting module is used to define the training termination conditions based on the measurement target, set the number of single-theory training steps and network update parameters, and fill the experience pool;
[0070] In the training module, the scanner performs actions in space according to the network strategy, calculates the reward obtained for each action based on the reward function, obtains the learning progress of the current agent, and saves the actions, states, rewards and learning progress to the experience pool;
[0071] The viewpoint planning module updates the parameters of the AC network after setting the training rounds based on the updated parameters, samples the tasks based on the Gaussian mixture model calculated based on the learning progress, generates a new training environment, and repeatedly executes the training module until the training termination conditions are met. The trained AC network is saved, the viewpoint planning strategy is obtained, and the viewpoint planning of a large-scale continuous space with six degrees of freedom is completed.
[0072] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the processor implements the steps of an aircraft skin viewpoint planning method based on deep reinforcement learning.
[0073] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for aircraft skin viewpoint planning based on deep reinforcement learning.
[0074] A deep reinforcement learning-based aircraft skin viewpoint planning method and system modeled the aircraft skin inspection task as a Markov decision process. By combining sensor field-of-view constraints with the geometric features of the CAD model to calculate the state and action spaces, this method mathematically describes the viewpoint planning problem in complex curved environments. This method effectively quantifies the impact of six-degree-of-freedom motion parameters on inspection coverage, addressing the subjectivity and inefficiency of traditional manual planning. An actor-critic (AC) network architecture is employed to construct a policy generation model, guiding the agent to learn an optimal path strategy through a custom reward function. Compared to traditional heuristic algorithms, this design significantly improves the global optimization capability of viewpoint sequences. By dynamically generating a progressive training environment using an experience pool replay mechanism and an asynchronous parameter update strategy, combined with a Gaussian mixture model, this approach addresses the exploration-exploitation dilemma of reinforcement learning in continuous large spaces and can adapt to the topological changes of the skin surfaces of different aircraft models. An online policy transfer mechanism allows the viewpoint planning strategy learned from virtual training to be directly applied to a physical scanning robot. This approach provides richer state information, helping the agent learn the underlying structure in the data, making more informed decisions in complex or unknown environments and improving overall performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 This is a flow chart of an aircraft skin viewpoint planning method based on deep reinforcement learning in one embodiment of the present invention;
[0076] Figure 2 This is a diagram of the overall architecture of viewpoint planning in one embodiment of the present invention;
[0077] Figure 3 A schematic diagram of voxel occupancy in one embodiment of the present invention;
[0078] Figure 4 Schematic diagram of the Actor and Critic network structure in one embodiment of the present invention, wherein (a) is a schematic diagram of the Actor network structure, and (b) is a schematic diagram of the Critic network structure;
[0079] Figure 5 This is a viewpoint planning effect diagram in one embodiment of the present invention. DETAILED DESCRIPTION
[0080] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below with reference to the accompanying drawings.
[0081] In one embodiment, Figure 1 and Figure 2 As shown, a method for aircraft skin viewpoint planning based on deep reinforcement learning, the method comprises the following steps:
[0082] S100: Model the viewpoint planning problem as a Markov decision problem, and calculate the state space and action space based on the sensor field of view constraints and the CAD model of the object being measured;
[0083] S200: Build an actor-critic AC-based network architecture, define a reward function based on the measurement target, and generate a six-degree-of-freedom simulation interaction environment required for reinforcement learning training;
[0084] S300: Define the training termination condition based on the measurement target, set the number of single theory training steps and network update parameters, and fill the experience pool;
[0085] S400: The scanner moves in space according to the network strategy, calculates the reward obtained for each action based on the reward function, obtains the learning progress of the current agent, and saves the action, state, reward and learning progress to the experience pool;
[0086] S500: Based on the updated parameters, the parameters of the AC network are updated after setting the training rounds. The Gaussian mixture model is calculated based on the learning progress to sample the tasks, a new training environment is generated, and S400 is repeated until the training termination conditions are met. The trained AC network is saved, the viewpoint planning strategy is obtained, and the six-degree-of-freedom large-scale continuous space viewpoint planning is completed.
[0087] The above-mentioned aircraft skin viewpoint planning method based on deep reinforcement learning can autonomously adjust the action space and state space according to the learning progress and skin scale, and stabilize the intelligent agent training process through the Gaussian mixture model.
[0088] In one embodiment, S100 includes:
[0089] S110: Constructing the viewpoint planning problem into a five-tuple of Markov decision processes Corresponding modeling, where Represents the state, which is the environmental information collected each time the scanner interacts with the environment; Indicates the action, that is, the action taken by the scanner each time; Represents the reward for each action taken; Represents the reward discount factor, which is used to calculate the reward discount of future actions; represents a set of transition probability functions, i.e., the probability of the scanner selecting the next action;
[0090] S120: Calculate the state space, input the CAD model of the aircraft skin, calculate the axis-aligned bounding box size of the model, record the maximum and minimum positions on the X, Y, and Z axes respectively, and the axis-aligned bounding box size is ; Define the size of a single voxel as , taking the bottom center of the model to be tested as the origin, the state space occupied by the calculation model is:
[0091] ;
[0092] in, Indicates expanding two voxel spaces outward to ensure that the bounding box still has margin at the boundary;
[0093] S130: Design a state transfer function based on the voxel occupancy state, and define the three states of voxels: occupied, unknown, and empty. Figure 3 As shown, the corresponding voxel state values are 1, 0.5 and 0 respectively. Suppose the scanner is at point Perform a scan, start from the camera center and calculate the rays along the field of view to each scanned occupied voxel, and define the coordinates of the occupied voxel center as , the maximum distance of the scanner is , the coordinates of a voxel in the state space are , then at point The state value of each voxel scanned can be calculated using the following formula:
[0094] ;
[0095] S140: Calculate the motion space based on the scanner's field of view and design a 6-DOF vector is the action parameter, Indicates that the scanner rotates along each axis without any constraints, so the range of motion is , assuming the working distance of the scanner is , combined with the size of the state space, the action range of the action space can be obtained as:
[0096] ;
[0097] S150: Setting the initial action space and state space, and selecting one tenth of the complete action space and state space as the action space and state space for initial training.
[0098] In one embodiment, S200 includes:
[0099] S210: Design a reinforcement learning network based on the actor-critic AC architecture. The network architecture includes an actor network, a target actor network, two critic networks and two target critic networks. The actor network and the critic network structure are as follows: Figure 4As shown in the figure, the Actor network consists of three hidden layers, each of which uses ReLU as the activation function. The last hidden layer uses tanh to map the action to the range (-1, 1). The Critic network consists of three hidden layers, except for the last hidden layer which directly outputs linearly, all other hidden layers use ReLU as the activation function.
[0100] S220: Designing an objective function. In the viewpoint planning problem, a strategy needs to be learned. Maximize rewards; the goal of reinforcement learning is to maximize the expected cumulative sum of rewards , it is necessary to introduce the entropy term Expand the objective function. is a random variable The probability density function of ; in reinforcement learning, The degree of randomness is determined by the strategy Control, according to the definition of entropy, the entropy term introduced in reinforcement learning is , the objective function is:
[0101] ;
[0102] in, Indicates that in the strategy Down The distribution of Represents the entropy coefficient, which is used to control the importance of entropy;
[0103] S230: Design a reward function to calculate the reward by scanning the new addition rate and the overall coverage rate; define the overall coverage rate , single scan new rate ,in Indicates the currently occupied voxel, Indicates that new voxels have been added in this scan. Represents all occupied voxels of the model; the reward function is:
[0104] ;
[0105] in, and Indicates the weight of the two rewards. In addition, when If empty, a penalty of -1 is applied.
[0106] In one embodiment, S300 includes:
[0107] S310: Setting the training termination condition: ending the training when the coverage reaches 90% or when the number of training rounds reaches 5000;
[0108] S320: Set the number of single-round training steps and network update parameters. The number of single-round training steps is 30. 128 sets of data are sampled from the experience pool each time a round of action is performed. The soft update coefficient of the target network is 0.01.
[0109] S330: Fill the experience pool and calculate the viewpoint according to the greedy algorithm to interact and obtain multiple sets of action, state and reward data; first, calculate the current state with the environment origin as the initial viewpoint , then, randomly generate around the initial viewpoint From a perspective, through the benefit function Calculate the information gain for each viewpoint:
[0110] ;
[0111] in, Indicates viewpoint The uncertainty of the new information is higher, the uncertainty is lower; the action with the largest gain is selected , and calculate the reward according to the reward function , calculate the next state based on the execution action , update a set of experience data in the experience pool ,in, Indicates whether the current round has ended; the loop continues until the coverage rate reaches the target or the maximum number of steps in a single round is reached.
[0112] In one embodiment, S400 includes:
[0113] S410: Current status Input is sent to the Actor network, and the Actor network outputs the execution action , the scanner performs actions in the environment , calculate the reward according to the reward function;
[0114] S420: After each round of interaction, calculate the current learning progress LP based on the task and the actions in the environment:
[0115] ;
[0116] in, and Represent the weight relationship between guidance and learning progress, Controls the decay rate, Indicates the training time, represents the agent reward, represents the interaction reward of the greedy algorithm, represents the reward of the agent the last time it was trained on the task, represents the reward of the agent when training in this task;
[0117] S430: Save the action, status, reward and learning progress to the experience pool.
[0118] In one embodiment, S500 includes:
[0119] S510: Extract experience from the experience pool to update the Critic network. First, calculate the target of each set of sampled data. value:
[0120] ;
[0121] Minimize the loss function for the two critics respectively, , specifically:
[0122] ;
[0123] S520: Extract experience from the experience pool to update the Actor network and minimize the loss function, specifically:
[0124] ;
[0125] S530: Update the corresponding target Critic / Actor network through soft update:
[0126] ;
[0127] in, Represents the target Critic / Actor network parameters, Represents the Critic / Actor network parameters, Control the update speed;
[0128] S540: Resampling the action space at different sizes of the aircraft skin according to the current policy , Respectively represent the length, width and height of the intercepted aircraft skin different action space; for the task Rewards available and calculate learning progress , the data Save in the experience pool, obtain N groups of data from the experience pool to fit the Gaussian mixture model;
[0129] S550: Adjust the initial action space and state space size according to the fitted Gaussian mixture model, calculate the mean of LP in each Gaussian distribution, and select the distribution sample with the maximum mean , and re-sample the task according to the state space and action space of the current aircraft skin to generate a new training environment, and repeat S400 until the training termination condition is met, and the training is terminated. The trained AC network is saved, the viewpoint planning strategy is obtained, and the six-degree-of-freedom large-scale continuous space viewpoint planning is completed.
[0130] In one embodiment, obtaining N sets of data from the experience pool and fitting the Gaussian mixture model in S540 includes:
[0131] S541: Obtain from the experience pool Group data, each group of data is , average initialization mixing weight , mean and covariance Then, calculate the posterior probability of each data:
[0132] ;
[0133] S542: Update based on posterior probability 、 、 :
[0134] ;
[0135] ;
[0136] ;
[0137] S543: Repeatedly calculate the posterior probability and update 、 、 , until the parameter change amplitude is less than the threshold, calculate the Bayesian Information Criterion BIC of the current model:
[0138] ;
[0139] in, is the maximum log-likelihood estimate of the Gaussian mixture model, Represents the number of samples used to calculate its penalty term;
[0140] S544: Repeat S541 to S543, fitting A Gaussian mixture model with different distributions is created and the one with the smallest BIC is selected as the optimal model.
[0141] Specifically, first determine the current task sampling range based on the Gaussian component weight distribution, and then obtain the length, width and height information of the new task , according to the length, width and height information, calculate the new state space and action space according to S100 to generate a new training environment. The final viewpoint planning effect diagram is as follows Figure 5 shown.
[0142] The above-mentioned deep reinforcement learning-based viewpoint planning method for aircraft skins designs a six-degree-of-freedom voxel action space. Automatically calculating the action space based on the input part size and camera field of view significantly improves the generalization of viewpoint planning. This approach avoids the hardware constraints and practical adaptability of using a spherical space, and is not limited to spherical surfaces or free-viewing orientations. To address the lack of sufficient spatial variation for flat parts, an entropy term is introduced into the objective function, increasing environmental exploration and preventing regression into local optima. This approach allows the algorithm to consider not only cumulative rewards but also the entropy of the policy during optimization, thus controlling the balance between action exploration and exploiting known optimal solutions and improving robustness. To efficiently guide the agent's learning in dynamic and complex task environments, a Gaussian mixture model is introduced to model the environment state, enabling better policy selection. This approach provides richer state information, helping the agent learn the underlying structure in the data, making more informed decisions in complex or unknown environments and improving overall performance.
[0143] In one embodiment, an aircraft skin viewpoint planning system based on deep reinforcement learning is also provided, comprising a state space and action space calculation module, a network architecture building module, a parameter setting module, a training module, and a viewpoint planning module;
[0144] The state space and action space calculation module is used to model the viewpoint planning problem as a Markov decision problem and calculate the state space and action space based on the sensor field of view constraints and the CAD model of the object being measured;
[0145] The network architecture building module is used to build the actor-critic AC-based network architecture, define the reward function based on the measurement target, and generate the six-degree-of-freedom simulation interaction environment required for reinforcement learning training;
[0146] The parameter setting module is used to define the training termination conditions based on the measurement target, set the number of single-theory training steps and network update parameters, and fill the experience pool;
[0147] In the training module, the scanner performs actions in space according to the network strategy, calculates the reward obtained for each action based on the reward function, obtains the learning progress of the current agent, and saves the actions, states, rewards and learning progress to the experience pool;
[0148] The viewpoint planning module updates the parameters of the AC network after setting the training rounds based on the updated parameters, samples the tasks based on the Gaussian mixture model calculated based on the learning progress, generates a new training environment, and repeatedly executes the training module until the training termination conditions are met. The trained AC network is saved, the viewpoint planning strategy is obtained, and the viewpoint planning of a large-scale continuous space with six degrees of freedom is completed.
[0149] Regarding the specific definition of an aircraft skin viewpoint planning system based on deep reinforcement learning, please refer to the definition of an aircraft skin viewpoint planning method based on deep reinforcement learning above, which will not be repeated here. The various modules in the above-mentioned aircraft skin viewpoint planning system based on deep reinforcement learning can be implemented in whole or in part through software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0150] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the processor implements the steps of an aircraft skin viewpoint planning method based on deep reinforcement learning.
[0151] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for aircraft skin viewpoint planning based on deep reinforcement learning.
[0152] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media used in the embodiments provided herein may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0153] The above is a detailed introduction to the aircraft skin viewpoint planning method and system based on deep reinforcement learning provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the core ideas of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A method for aircraft skin viewpoint planning based on deep reinforcement learning, characterized in that: The method comprises the following steps: S100: Model the viewpoint planning problem as a Markov decision problem, and calculate the state space and action space based on the sensor field of view constraints and the CAD model of the object being measured; S200: Build an actor-critic AC-based network architecture, define a reward function based on the measurement target, and generate a six-degree-of-freedom simulation interaction environment required for reinforcement learning training; S300: Define the training termination condition based on the measurement target, set the number of single theory training steps and network update parameters, and fill the experience pool; S400: The scanner moves in space according to the network strategy, calculates the reward obtained for each action based on the reward function, obtains the learning progress of the current agent, and saves the action, state, reward and learning progress to the experience pool; S500: Based on the updated parameters, the parameters of the AC network are updated after setting the training rounds. The Gaussian mixture model is calculated based on the learning progress to sample the tasks, a new training environment is generated, and S400 is repeated until the training termination conditions are met. The trained AC network is saved, the viewpoint planning strategy is obtained, and the six-degree-of-freedom large-scale continuous space viewpoint planning is completed.
2. The method according to claim 1, characterized in that S100 includes: S110: Constructing the viewpoint planning problem into a five-tuple of Markov decision processes Corresponding modeling, where Represents the state, which is the environmental information collected each time the scanner interacts with the environment; Indicates the action, that is, the action taken by the scanner each time; Represents the reward for each action taken; Represents the reward discount factor, which is used to calculate the reward discount of future actions; represents a set of transition probability functions, i.e., the probability of the scanner selecting the next action; S120: Calculate the state space, input the CAD model of the aircraft skin, calculate the axis-aligned bounding box size of the model, record the maximum and minimum positions on the X, Y, and Z axes respectively, and the axis-aligned bounding box size is ; Define the size of a single voxel as , taking the bottom center of the model to be tested as the origin, the state space occupied by the calculation model is: ; in, Indicates expanding two voxel spaces outward to ensure that the bounding box still has margin at the boundary; S130: Design a state transfer function based on the voxel occupancy state, define the voxel to be occupied, unknown, and empty, and the corresponding voxel state values are 1, 0.5, and 0 respectively. Suppose the scanner is at point Perform a scan, start from the camera center and calculate the rays along the field of view to each scanned occupied voxel, and define the coordinates of the occupied voxel center as , the maximum distance of the scanner is , the coordinates of a voxel in the state space are , then at point The state value of each voxel scanned can be calculated using the following formula: ; S140: Calculate the motion space based on the scanner's field of view and design a 6-DOF vector is the action parameter, Indicates that the scanner rotates along each axis without any constraints, so the range of motion is , assuming the working distance of the scanner is , combined with the size of the state space, the action range of the action space can be obtained as: ; S150: Setting the initial action space and state space, and selecting one tenth of the complete action space and state space as the action space and state space for initial training.
3. The method according to claim 2, characterized in that S200 includes: S210: Design a reinforcement learning network based on the actor-critic (AC) architecture. The network architecture consists of an actor network, a target actor network, two critic networks, and two target critic networks. The actor network consists of three hidden layers, each of which uses ReLU as the activation function. The last hidden layer uses tanh to map the action to the range (-1, 1). The critic network consists of three hidden layers, and except for the last hidden layer, which directly outputs linearly, all other hidden layers use ReLU as the activation function. S220: Designing an objective function. In the viewpoint planning problem, a strategy needs to be learned. Maximize rewards; the goal of reinforcement learning is to maximize the expected cumulative sum of rewards , it is necessary to introduce the entropy term Expand the objective function. is a random variable The probability density function of ; in reinforcement learning, The degree of randomness is determined by the strategy Control, according to the definition of entropy, the entropy term introduced in reinforcement learning is , the objective function is: ; in, Indicates that in the strategy Down The distribution of Represents the entropy coefficient, which is used to control the importance of entropy; S230: Design a reward function to calculate the reward by scanning the new addition rate and the overall coverage rate; define the overall coverage rate , single scan new rate ,in Indicates the currently occupied voxel, Indicates that new voxels have been added in this scan. Represents all occupied voxels of the model; the reward function is: ; in, and Indicates the weight of the two rewards. In addition, when If empty, a penalty of -1 is applied.
4. The method according to claim 3, characterized in that S300 includes: S310: Setting the training termination condition: ending the training when the coverage reaches 90% or when the number of training rounds reaches 5000; S320: Set the number of single-round training steps and network update parameters. The number of single-round training steps is 30. 128 sets of data are sampled from the experience pool each time a round of action is performed. The soft update coefficient of the target network is 0.
01. S330: Fill the experience pool and calculate the viewpoint according to the greedy algorithm to interact and obtain multiple sets of action, state and reward data; first, calculate the current state with the environment origin as the initial viewpoint , then, randomly generate around the initial viewpoint From a perspective, through the benefit function Calculate the information gain for each viewpoint: ; in, Indicates viewpoint The uncertainty of the new information is higher, the uncertainty is lower; select the action with the largest gain as A, and calculate the reward according to the reward function , calculate the next state based on the execution action , update a set of experience data in the experience pool ,in, Indicates whether the current round has ended; the loop continues until the coverage rate reaches the target or the maximum number of steps in a single round is reached.
5. The method according to claim 4, characterized in that S400 includes: S410: Current status Input is fed into the Actor network, the Actor network outputs action A, the scanner performs action A in the environment, and calculates the reward based on the reward function; S420: After each round of interaction, calculate the current learning progress LP based on the task and the actions in the environment: ; in, and Represent the weight relationship between guidance and learning progress, Controls the decay rate, Indicates the training time, represents the agent reward, represents the interaction reward of the greedy algorithm, represents the reward of the agent the last time it was trained on the task, represents the reward of the agent when training in this task; S430: Save the action, status, reward and learning progress to the experience pool.
6. The method according to claim 5, characterized in that S500 includes: S510: Extract experience from the experience pool to update the Critic network. First, calculate the target of each set of sampled data. value: ; Minimize the loss function for the two critics respectively, , specifically: ; S520: Extract experience from the experience pool to update the Actor network and minimize the loss function, specifically: ; S530: Update the corresponding target Critic / Actor network through soft update: ; in, Represents the target Critic / Actor network parameters, Represents the Critic / Actor network parameters, Control the update speed; S540: Resampling the action space at different sizes of the aircraft skin according to the current policy , Respectively represent the length, width and height of the intercepted aircraft skin different action space; for the task Rewards available and calculate learning progress , the data Save in the experience pool, obtain N groups of data from the experience pool to fit the Gaussian mixture model; S550: Adjust the initial action space and state space size according to the fitted Gaussian mixture model, calculate the mean of LP in each Gaussian distribution, and select the distribution sample with the maximum mean , and re-sample the task according to the state space and action space of the current aircraft skin to generate a new training environment, and repeat S400 until the training termination condition is met, and the training is terminated. The trained AC network is saved, the viewpoint planning strategy is obtained, and the six-degree-of-freedom large-scale continuous space viewpoint planning is completed.
7. The method according to claim 6, characterized in that In S540, N groups of data are obtained from the experience pool to fit the Gaussian mixture model, including: S541: Obtain from the experience pool Group data, each group of data is , average initialization mixing weight , mean and covariance Then, calculate the posterior probability of each data: ; S542: Update based on posterior probability 、 、 : ; ; ; S543: Repeatedly calculate the posterior probability and update 、 、 , until the parameter change amplitude is less than the threshold, calculate the Bayesian Information Criterion BIC of the current model: ; in, is the maximum log-likelihood estimate of the Gaussian mixture model, Represents the number of samples used to calculate its penalty term; S544: Repeat S541 to S543, fitting A Gaussian mixture model with different distributions is created and the one with the smallest BIC is selected as the optimal model.
8. An aircraft skin viewpoint planning system based on deep reinforcement learning, characterized in that: It includes state space and action space calculation module, network architecture building module, parameter setting module, training module and viewpoint planning module; The state space and action space calculation module is used to model the viewpoint planning problem as a Markov decision problem and calculate the state space and action space based on the sensor field of view constraints and the CAD model of the object being measured; The network architecture building module is used to build the actor-critic AC-based network architecture, define the reward function based on the measurement target, and generate the six-degree-of-freedom simulation interaction environment required for reinforcement learning training; The parameter setting module is used to define the training termination conditions based on the measurement target, set the number of single-theory training steps and network update parameters, and fill the experience pool; In the training module, the scanner performs actions in space according to the network strategy, calculates the reward obtained for each action based on the reward function, obtains the learning progress of the current agent, and saves the actions, states, rewards and learning progress to the experience pool; The viewpoint planning module updates the parameters of the AC network after setting the training rounds based on the updated parameters, samples the tasks based on the Gaussian mixture model calculated based on the learning progress, generates a new training environment, and repeatedly executes the training module until the training termination conditions are met. The trained AC network is saved, the viewpoint planning strategy is obtained, and the viewpoint planning of a large-scale continuous space with six degrees of freedom is completed.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Robot three-dimensional measurement path planning method based on deep reinforcement learning
CN116604571A
Robot three-dimensional reconstruction equipment viewpoint planning method and system and computer equipment
CN120198602A