PINN adaptive sampling method and equipment based on multi-agent reinforcement learning

By employing an adaptive sampling method based on multi-agent reinforcement learning, agents perform actions in the spatial dimension to identify complex regions and generate high-value sampling points. This solves the problem of training PINNs in complex partial differential equations and achieves high-precision and stable solution results.

CN121981196APending Publication Date: 2026-05-05SHAANXI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHAANXI NORMAL UNIV
Filing Date
2026-01-28
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing physical information neural networks (PINNs) suffer from fixed sampling point distribution, difficulty in adaptively identifying high-error regions, and insufficient global exploration capabilities when solving complex partial differential equations. This leads to training difficulties and decreased accuracy, especially in engineering applications such as nonlinear wave equations, where they struggle to meet the requirements for solution accuracy and stability.

Method used

An adaptive sampling method for multi-agent reinforcement learning is constructed. By introducing a fixed-time anchor mechanism in the Markov decision process, the agent performs actions in the spatial dimension. The reward is calculated based on the local change features of the pre-trained PINN. Combining regional rewards and redundancy penalties, the policy network and value network are updated using a multi-agent reinforcement learning algorithm to generate an adaptive sampling point set and perform mixed sampling training.

Benefits of technology

It significantly improves the solution accuracy and training stability of complex PDEs such as nonlinear wave equations, reduces sampling point waste, lowers the computational cost of traditional residual-driven methods, and improves sampling efficiency and model training effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121981196A_ABST
    Figure CN121981196A_ABST
Patent Text Reader

Abstract

The invention discloses a PINN adaptive sampling method and equipment based on multi-agent reinforcement learning, and belongs to the technical field of artificial intelligence and scientific calculation. According to the method, a multi-agent sampling environment is constructed, and a fixed time anchor point mechanism is introduced, so that agents concentrate on collaborative exploration of a spatial domain under a frozen time slice; a composite reward mechanism including prediction of a solution local change amplitude, area coverage reward and redundancy removal penalty is designed, intelligent agents autonomously recognize and locate a complex area with severe solution change, a sampling strategy is trained based on a multi-intelligent-agent near-end strategy optimization algorithm, a self-adaptive sampling point set is generated and fused with an initial point set, fine training is performed on a PINN, and an optimal solution is obtained. The problem that when an existing physical information neural network is used for solving a partial differential equation with strong nonlinearity or local high gradient characteristics, a fixed sampling strategy is difficult to capture a key area is solved, the solving precision and convergence speed of a model in a complex dynamic system are improved, and the sampling cost is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of scientific computing, numerical solution of partial differential equations and artificial intelligence, and specifically relates to an adaptive sampling method and device for Physics-Informed Neural Networks (PINNs) based on multi-agent reinforcement learning. Background Technology

[0002] Partial differential equations (PDEs) are important mathematical tools for describing continuously changing phenomena in nature and engineering systems. In recent years, with the rapid development of scientific computing technology, deep learning-based PDE solving methods have gradually become a research hotspot. Among them, Physical Information Neural Networks (PINNs) have gained widespread attention in fields such as electromagnetic field simulation, structural dynamics, fluid mechanics, quantum physics, and biomolecular dynamics due to their meshless and highly generalizable characteristics. PINNs embed the control PDEs, initial conditions, and boundary conditions into the loss function of the neural network in the form of residuals, enabling the network to simultaneously approximate the data and satisfy physical constraints during training, thus possessing good theoretical consistency and practical flexibility.

[0003] However, existing PINNs methods generally rely on a fixed set of sampling points pre-generated within the computational domain. For partial differential equations with nonlinear characteristics, intense local variations, or multi-scale behavior (such as nonlinear wave equations, reaction-diffusion equations, and singular perturbation equations), uniform sampling struggles to effectively capture the complex structure of solutions in local regions, leading to difficulties in network training and a significant decrease in solution accuracy. In practical engineering applications, nonlinear wave equations are prevalent in scenarios such as superconducting device simulation, nonlinear dielectric wave propagation, crystal dislocation dynamics, and large-amplitude structural vibrations. These equations often contain significant nonlinear terms and steep local regions, placing higher demands on the accuracy and stability of the solution methods, which fixed sampling strategies struggle to meet.

[0004] To improve solution accuracy, existing research has proposed residual-driven adaptive sampling methods, such as Residual Adaptive Refinement (RAR) and Residual Adaptive Distribution (RAD). These methods select new sampling points by calculating the residuals of the current PINNs solution, thereby refining regions with large errors. Although residual-driven methods improve sampling efficiency to some extent, they still have two shortcomings: first, they require evaluating PDE residuals on a large number of candidate points, resulting in huge computational overhead; second, these methods are essentially local greedy strategies, passively refining based only on the residual information of the current model, easily getting trapped in local optima and lacking global exploration capabilities. In addition, in high-dimensional problems or highly nonlinear dynamic equations, the residual construction itself may be affected by noise, causing the sampling strategy to fail or become unstable.

[0005] These shortcomings of traditional PINNs sampling methods make it difficult for existing technologies to achieve efficient and high-precision solutions to complex partial differential equations. In practical engineering contexts, such as superconducting Josephson junction models, dynamic response analysis of large-scale mechanical structures, and simulation of nonlinear material behavior, partial differential equations are often highly nonlinear and exhibit dramatic local variations. Existing sampling strategies struggle to simultaneously guarantee solution efficiency, accuracy, and stability. Therefore, there is an urgent need for a novel sampling method with proactive exploration capabilities, capable of adaptively identifying key regions, and effectively improving the accuracy of PINNs solutions. Summary of the Invention

[0006] This invention aims to address the problems of existing Physical Information Neural Networks (PINNs) in solving complex partial differential equations (PDEs), such as fixed sampling point distribution, difficulty in adaptively identifying high-error regions, insufficient global exploration capability, and excessive residual computation overhead. Especially for dynamic control equations of significant engineering importance, such as nonlinear wave equations, whose solutions often contain obvious spatial non-uniformity, local high-gradient regions, or multi-scale structures, traditional uniform sampling methods cannot effectively capture key regions, leading to difficulties in neural network training and high errors.

[0007] To overcome the above-mentioned shortcomings, in a first aspect, the present invention provides a PINN adaptive sampling method based on multi-agent reinforcement learning, comprising the following steps: A multi-agent sampling environment is constructed, and the sampling process is modeled as a Markov decision process. A fixed time anchor mechanism is introduced when setting the sampling environment, and the agents only perform actions in the spatial dimension. Multiple agents are controlled to perform actions in a set sampling environment. The reward is calculated based on the local change characteristics of the predicted solution of the pre-trained PINN at a fixed time anchor point, and the total reward of the agent is obtained by combining the regional reward and the redundancy penalty. The policy network and value network of the agents are updated using a multi-agent reinforcement learning algorithm to obtain a trained adaptive sampling policy. An adaptive sampling point set is generated using a trained adaptive sampling strategy, and then merged with the initial sampling point set to form a hybrid sampling point set. The hybrid sampling point set is used to refine the training of PINN to obtain the target PINN model for solving the target partial differential equation.

[0008] Furthermore, obtaining the pre-trained PINN includes: Obtain the target partial differential equation, determine the space-time solution domain, time interval, initial conditions and boundary conditions, and construct an equation operator expression form that can be used for automatic differentiation; An initial set of sampling points is generated within the solution domain, a physical information neural network is constructed, and the PINN is pre-trained using the initial set of sampling points to obtain a pre-trained PINN.

[0009] Furthermore, generating the initial set of sampling points within the solution domain includes: The method for generating the initial sampling points is as follows:

[0010] in, The space-time solution domain is defined in step one; This represents the initial set of sampling points, where the sampling method is random sampling, uniform partitioning sampling, or Latin hypercube sampling.

[0011] Furthermore, a multi-agent sampling environment is constructed, and the sampling process is modeled as a Markov decision process; a fixed-time anchor mechanism is introduced when setting the sampling environment, and the agents only perform actions in the spatial dimension, including: At the beginning of each sampling round, a time value is randomly selected within the time interval. And remain unchanged during that sampling round; In the spatial domain Each agent is randomly assigned an initial coordinate. And place multiple agents together in the same environment to perform policy learning tasks; spatial domain According to the number of sections Divided into several sub-regions And assign a region identifier to each agent. ; The state space of an agent is defined as follows: ,in The current spatial location of the agent. This identifies the spatial sub-region to which the agent belongs; the agent's action space only includes movement in the spatial dimension.

[0012] Furthermore, the formula for calculating the total reward of the intelligent agent is as follows:

[0013] in, The reward for local changes is calculated based on the local difference magnitude of the predicted solution; For regional rewards, the calculation is based on the distance between the agent and the center of its region; As a redundancy penalty, calculations are based on the distance between agents; Local change reward The calculation method is as follows: Suppose the intelligent agent is located Move to Calculate PINN at a fixed time anchor point The predicted solution difference magnitude ;when When the value exceeds a preset threshold, a positive reward is given; otherwise, the reward is 0. The regional rewards and redundancy penalty The calculation method is as follows: The regional reward is represented as ,in The distance from the agent's current location to the center of the region. This is the regional reward coefficient. The attenuation factor is denoted as ; the redundancy penalty is expressed as . ,in The distance between the current agent and its nearest neighbor agents. The penalty coefficient is... The interaction radius.

[0014] Furthermore, the policy network and value network of the agents are updated using a multi-agent reinforcement learning algorithm to obtain the trained adaptive sampling policies, including: Multiple agents share a unified policy network Its update objective is to maximize expected return:

[0015] Policy updates adopt the pruning objective of Proximal Policy Optimization (PPO), in the following form:

[0016] For time step Expectations It is the probability ratio. For the dominant function, For the parameters of the policy network, This is for pruning hyperparameters.

[0017] Building a value network The advantage function is used to predict the expected return of a state. It is calculated as follows:

[0018] in, As a discount factor, For time step The value network is trained by minimizing the following loss to provide immediate rewards:

[0019] in, For actual returns; In each policy iteration, both the policy network and the value network are updated simultaneously, and the following overall objective is optimized using the gradient descent algorithm:

[0020] in For the policy entropy term, These are weighting coefficients, all greater than 0; after multiple rounds of updates, the agent can learn a stable adaptive sampling strategy. Multiple agents share a policy network and evaluate the global state through a centralized value network; the policy network's update objective function uses a shear ratio.

[0021] in It is the probability ratio. This is the dominant function.

[0022] Furthermore, an adaptive sampling point set is generated using the trained adaptive sampling strategy, and this set is then fused with the initial sampling point set to form a hybrid sampling point set, including: Record the locations visited by the agent during the exploration process, and select points whose individual rewards meet preset conditions as adaptive sampling points:

[0023] Will With a uniformly distributed initial set of sampling points The samples are then merged to obtain a mixed set of sampling points.

[0024] Furthermore, the mixed sampling training loss function for PINN is:

[0025] in The physical equation residual loss on the mixed sampling point set, Loss due to initial conditions This is the boundary condition loss.

[0026] Secondly, the present invention also provides a computer device including a processor and a memory, the memory being used to store a computer executable program, the processor reading part or all of the computer executable program from the memory and executing it, and the processor executing part or all of the computer executable program can realize the above-mentioned PINN adaptive sampling method based on multi-agent reinforcement learning.

[0027] Simultaneously, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, it can implement the above-described PINN adaptive sampling method based on multi-agent reinforcement learning.

[0028] Compared with existing technologies, this invention provides a PINN adaptive sampling method based on multi-agent reinforcement learning. It constructs a multi-agent sampling environment, models the configuration point selection process as a Markov decision process, and allows multiple agents to collaboratively explore in the spatial domain. By analyzing the local solution change characteristics of the neural network at a fixed time anchor point, it autonomously identifies regions where partial differential equation solutions exhibit rapid changes or complex structures, thereby generating a high-value configuration point set for subsequent refined training of PINNs. This is a PINNs adaptive sampling method with active exploration capabilities, effectively discovering and encrypting important sampling regions, significantly improving the solution accuracy and training stability of complex PDEs such as nonlinear wave equations, reducing sampling point waste, and lowering the high computational cost of traditional residual-driven methods that require residual scanning across the entire domain. This invention not only improves the numerical stability and accuracy of solving partial differential equations mathematically, but also has significant value in practical engineering applications. For fields such as superconducting device simulation, structural dynamics analysis, nonlinear material response prediction, and wave propagation control, this invention can provide high-quality numerical solutions with limited computational cost, offering a new and effective approach for the application of physical information neural networks in scientific computing and engineering simulation. Attached Figure Description

[0029] Figure 1 This is an overall flowchart of the PINN adaptive sampling method based on multi-agent reinforcement learning provided in this embodiment of the invention; Figure 2 This is a structural framework diagram of the PINN adaptive sampling method based on multi-agent reinforcement learning provided in the embodiments of the present invention; Figure 3 This is a comparison diagram of the predicted three-dimensional surface of the model trained by different sampling methods in the embodiments of the present invention and the true solution; Figure 4 This is a comparison diagram of the distribution of L2 error in the spatiotemporal domain of models trained by different sampling methods in embodiments of the present invention; Figure 5 This is a comparison chart of the convergence curves of the loss function during the training process for different sampling methods in this embodiment of the invention. Detailed Implementation

[0030] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Other implementation methods that can be obtained by those skilled in the art based on the disclosure of the present invention without creative effort are all within the protection scope of the present invention.

[0031] Please see Figure 1 and Figure 2The first embodiment of the present invention provides a PINN adaptive sampling method based on multi-agent reinforcement learning for efficiently solving partial differential equations. The method includes the following steps: Step 1: Determine the solution domain, time interval, initial conditions, and boundary conditions based on the objective partial differential equation, and construct an equation operator expression form that can be used for automatic differentiation.

[0032] Step 2: Generate an initial set of sampling points within the solution domain to construct the training input for the physical information neural network, and pre-train the PINN to enable it to have a preliminary ability to fit the overall structure of the equation.

[0033] Step 3: Construct a multi-agent sampling environment, set a fixed time anchor point mechanism, the initial position of the agents, and the region division method, and establish a sampling framework based on Markov decision process.

[0034] Step 4: Execute actions in a multi-agent environment and calculate rewards based on the changes in the solution of the neural network under fixed time slices. At the same time, combine regional rewards and redundancy penalties to obtain the total reward for the agent.

[0035] Step 5: Use the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm to update the agent policy network and value network, so that the agent can gradually learn an adaptive sampling strategy that can identify important regions of the equation.

[0036] Step 6: Generate an adaptive sampling point set based on the rewards obtained during the agent's exploration process, and merge it with the initial sampling point set to perform hybrid sampling training on the physical information neural network, thereby obtaining the optimized target PINN model, and solve the target partial differential equation based on the model.

[0037] In step one, the spatial-temporal solution domain, time interval, initial conditions, and boundary conditions are determined based on the objective partial differential equation, including: The general form of the objective partial differential equation can be abstracted as follows:

[0038] in, The solution to the equation; To act on Differential operators, integral operators, nonlinear operators, or combinations thereof; This represents the control operator used to describe the dynamic laws of the system; This represents the domain of the equation.

[0039] The general form of initial and boundary conditions is: , , in, , and These are known functions or known boundary values, used to ensure the well-posedness of partial differential equations.

[0040] Constructing operator representations of equations that can be automatically differentiated, including: Using parameterized neural networks Approximate the solution of the partial differential equation and obtain the operator through automatic differentiation technique. The required derivative terms or related operator values ​​are then substituted into the operator expression of the equation to obtain the physical residuals of the partial differential equation.

[0041] The residual is used for subsequent training and error constraints of the PINN. This step aims to construct a differentiable PDE computation graph, enabling the network to learn the equation structure through backpropagation.

[0042] It should be noted that the adaptive sampling strategy of the present invention is based on the local solution change characteristics of the neural network on a fixed time slice, and does not rely on the above-mentioned PDE residuals for reward calculation. The purpose is to reduce computation and increase sampling efficiency. The residual term is only used for the training of the PINN network itself, and not for the evaluation of multi-agent policies.

[0043] Step two involves generating an initial set of sampling points within the space-time solution domain to construct the training input for the physical information neural network (PINN), and pre-training PINN to enable it to initially fit the overall structure of the equations, including: The method for generating the initial sampling points is represented as follows:

[0044] in, The space-time solution domain is defined in step one; This represents the initial set of sampling points. In this invention, the sampling method can be random sampling, uniform partitioning sampling, or Latin hypercube sampling to ensure basic coverage of the solution domain.

[0045] The initial sampling points are used to construct the PINN input, and their corresponding physical constraints include points inside the physical equations, initial condition points, and boundary condition points. These three types of points correspond to the physical equation constraints, initial condition constraints, and boundary condition constraints described in step one, respectively.

[0046] The training loss of a physical information neural network is expressed as:

[0047] in, This represents the residual loss in the physical equations; Indicates the initial condition loss; This represents the loss due to boundary conditions.

[0048] Specifically, the automatic differentiation mechanism of the deep learning framework is first used to calculate the first or higher-order partial derivatives of the neural network output with respect to the spatiotemporal coordinate input, thereby constructing a loss function term that includes physical constraints; then, the gradient information of the total loss function with respect to the network weights and biases is calculated through the backpropagation algorithm.

[0049] Subsequently, based on the gradients calculated above, deep learning optimization algorithms (such as Adam or L-BFGS) are used to iteratively update the parameters of the physical information neural network, enabling the network to capture the overall evolution law of the partial differential equations in the initial stage, thereby obtaining a PINN model that satisfies the basic physical constraints.

[0050] In this invention, the basic PINN model will serve as the environment function in the subsequent multi-agent sampling stage, and will be fixed to provide the predicted values ​​and their changes at different spatial locations, thus providing an evaluation basis for agent policy learning.

[0051] Step 3: Construct a multi-agent sampling environment, setting a fixed-time anchor point mechanism, initial agent positions, and region partitioning method, and establishing a sampling framework based on Markov decision processes, including: The method for setting fixed time anchor points is generally as follows: within the time interval determined in step one. Randomly select a time value The sampling process remains constant throughout, allowing the multi-agent to move only in the spatial dimension, thus enabling the evaluation of the spatial variation characteristics of the neural network's predicted solutions with fixed-time slices. This fixed-time anchoring mechanism aims to improve the stability of the sampling process and avoid policy learning difficulties caused by actions in the temporal dimension.

[0052] The initial position of the agent is set in the spatial domain as follows: Each agent is randomly assigned an initial coordinate. Furthermore, multiple agents are placed together in the same environment to perform policy learning tasks. Randomization of initial positions helps improve global coverage and prevents agents from clustering in the early stages.

[0053] The spatial division of the sampling environment is generally as follows: dividing the spatial domain According to the number of sections Divided into several sub-regions And assign a region identifier to each agent. Regional division is used to enhance the spatial division of labor capabilities of agents and introduces regional coverage factors into the reward mechanism, enabling agents to maintain stable exploration within their respective regions.

[0054] After constructing the aforementioned multi-agent sampling environment, the agents are state-based. Perform left and right movements within the discrete motion space, using a step size. This invention achieves spatial location updates. A unified state definition and action update method is used, and the calculation process will not be elaborated further. This environment serves as an interactive platform for agent policy learning, enabling the agent to conduct spatial exploration based on the prediction results of the physical information neural network at fixed time slices, thereby generating high-value sampling points for subsequent adaptive sampling.

[0055] As an example, the sampling process is modeled as a Markov decision process. To improve learning stability, this invention uses a fixed time anchor point so that the time dimension is not controlled by the action but is only used to evaluate the spatial changes of the solution at that moment. State space:

[0056] in The current spatial location of the agent. For each Randomly sampled and fixed time anchor points, This indicates the identifier of the spatial sub-region to which the intelligent agent belongs.

[0057] The motion space only includes left and right movement in the spatial dimension: ,in: action Move to the left ;action Move to the right .

[0058] The position update formula is:

[0059] Under the condition that the time anchor point remains unchanged, the present invention constructs a reward function based on the movement result of the agent in the spatial coordinates to measure the local complexity of the partial differential equation solution at that position and achieve cooperative coverage of multiple agents.

[0060] Step four involves executing actions in a multi-agent sampling environment and calculating rewards based on the changes in the neural network's predicted solutions over fixed time slices. Simultaneously, the total reward for the agent is obtained by combining regional rewards and redundancy penalties, including: To measure the change in the predicted solution between the agent's current position and the updated position, this invention uses the local rate of change as the reward criterion. Let the agent be at the time anchor point... Below by position Move to The local difference component is defined as:

[0061] To enhance direction independence, this invention takes its absolute value and uses a threshold to determine whether to award a reward, specifically defined as:

[0062] in, This is the change threshold. The reward is used to highlight regions where the solution changes rapidly, guiding the agent to prioritize exploring locations with higher local complexity. Not less than the preset threshold Individuals are rewarded at appropriate times:

[0063] This guides the agent to preferentially navigate to spatial regions where solutions change significantly, and where nonlinear behavior or structural mutations may occur.

[0064] To enhance spatial domain coverage, this invention introduces a regional reward. An agent receives an additional reward when active near the center of its assigned region. This regional reward can be defined as:

[0065] in This represents the distance between the agent's current position and the center of the region; All adjusted parameters are greater than 0. This encourages agents to maintain reasonable coverage within their respective regions, avoiding insufficient region exploration.

[0066] To reduce excessive aggregation of agents in space, this invention constructs a redundancy penalty based on the minimum distance between agents. If the distance between agents is less than a set radius... , The penalty is then defined as:

[0067] in, Let be the distance between any two agents. This is the penalty coefficient. The penalty increases when agents get too close, thus increasing the dispersion of space exploration. No penalty is incurred, thus encouraging agents to distribute their work across the entire domain. Simultaneously, to enhance coverage stability within their respective areas of responsibility, this invention sets regional rewards based on spatial partitioning: the distance from the region center to the current location is denoted as... The regional reward can then be expressed as:

[0068] This invention will include the above-mentioned rewards Regional rewards and punishment The signal combination serves as a single-step reward to evaluate the agent's overall performance under the current action. The reward format is as follows:

[0069] Without introducing time-dimensional action decisions, local complexity can be effectively measured solely by spatial solution difference over fixed time slices. Furthermore, collaborative coverage and deduplication among multiple agents are achieved through region and redundancy terms, automatically generating high-value adaptive sampling points without global residual scanning. This reward system, through the combined effects of local change detection, region coverage, and agent decentralization, enables agents to more effectively identify high-value sampling points, thereby improving the learning efficiency of the adaptive sampling strategy.

[0070] Step 5: Update the agent policy network and value network using the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm, enabling the agents to gradually learn adaptive sampling strategies capable of identifying important regions of partial differential equations, including: After completing the reward construction in step four, this invention records the state-action-reward sequence obtained by the agent's interactions in the sampled environment, forming an experience dataset. For any agent, its trajectory can be represented as:

[0071] Where T is the total number of time steps for each trajectory; this trajectory is used to learn the parameters of the policy network and the value network.

[0072] This invention employs the MAPPO framework, which features centralized training and distributed execution. Multiple agents share a unified policy network. Its update objective is to maximize expected return:

[0073] Policy updates adopt the pruning objective of Proximal Policy Optimization (PPO), in the following form:

[0074] For time step Expectations It is the probability ratio. For the dominant function, For the parameters of the policy network, For pruning hyperparameters; To estimate the advantage function, this invention constructs a value network. This is used to predict the expected reward of a state. The advantage calculation method can be expressed as:

[0075] in As a discount factor, For time step The immediate reward. The value network is trained by minimizing the following loss:

[0076] in For actual returns.

[0077] In each policy iteration, this invention simultaneously updates the policy network and the value network, optimizing the following overall objective through gradient descent:

[0078] in For the policy entropy term, These are weighting coefficients, all greater than 0, used to improve exploration capabilities and training stability.

[0079] After multiple updates, the agent can learn a stable adaptive sampling strategy, maintaining sensitivity to complex regions, highly variable regions, or regions with structural abrupt changes, thereby providing higher quality sampling points for PINN and significantly improving training efficiency and PDE solution accuracy.

[0080] Step six: Generate an adaptive sampling point set based on the rewards obtained during the agent's exploration process, and merge it with the initial sampling point set. Perform hybrid sampling training on the physical information neural network to obtain the optimized target PINN model. Solve the target partial differential equation based on this model, including: After multiple rounds of policy training, each agent executes the learned policy in the spatial domain, obtaining several sampling position sequences. For the agent at the time anchor point... The execution trajectory below The points with higher local change rewards are selected as adaptive sampling points, defined as:

[0081] The initial sampling point set generated in step two Compared with the adaptive sampling point set obtained in step six The samples are then fused to construct a hybrid set of sampling points.

[0082] This hybrid sampling point set retains basic coverage of the global region while achieving higher sampling density in complex local regions. It can significantly reduce the cost of residual calculation while greatly improving the model's solution accuracy and convergence speed on complex partial differential equations, thereby enhancing the training effect of PINN.

[0083] Based on the mixed sampling point set The training input for the physical information neural network is constructed, and the derivative terms required for the equation are obtained through automatic differentiation, forming a hybrid loss function:

[0084] Subsequently, an optimization algorithm is used to update the network parameters, enabling the neural network to be retrained under the new sampling distribution and obtain the converged target model, thereby realizing the numerical solution of the target partial differential equation.

[0085] Implementation example: This embodiment selects a class of nonlinear wave equations with important physical background to verify the effectiveness of the method of the present invention. These equations can be used to describe typical engineering scenarios such as the phase difference evolution of Josephson superconducting junctions under external excitation, forced wave propagation in nonlinear media, and large-amplitude vibrations of flexible materials. Their important characteristics include the presence of significant nonlinear restoring force terms and non-uniform driving terms, leading to steep structures or complex oscillation modes in local solutions, thus placing higher demands on the sampling strategy. A comparison of the predicted three-dimensional surface of the model trained by different sampling methods with the actual solution in this embodiment of the invention is provided in reference [reference missing]. Figure 3 .

[0086] Josephson superconducting junctions are a class of superconducting devices exhibiting quantum tunneling effects, whose dynamic behavior can be determined by phase difference. The evolution description. Under weak coupling conditions, this phase satisfies the following typical nonlinear dynamic equation:

[0087] in The term corresponds to the nonlinear response of the Josephson current; This indicates an external electromagnetic drive; while the second-order spatial term describes coupling and propagation on the junction surface. In this embodiment, parameter normalization is used to obtain... This form is consistent with the generalized Sine-Gordon model and can truly reflect the wave behavior of the Josephson junction under forced driving.

[0088] The spatial and temporal intervals were selected as follows The initial conditions are set as follows:

[0089] The initial state of the standing wave corresponds to the static state; while the boundary conditions are:

[0090] It can simulate the situation where the two ends of a Josephson junction are subjected to phase modulation or external periodic electromagnetic excitation.

[0091] The driver item is set to:

[0092] External driving forces that represent spatial non-uniformity and temporal periodicity.

[0093] The specific solution process follows the steps described in this invention: First, initial sampling points are constructed and the physical information neural network is pre-trained to establish a preliminary fit of the model to the overall structure of the equation. Then, a multi-agent environment with a fixed-time anchor mechanism is constructed. The agents, through collaborative exploration, autonomously locate sensitive regions where the solution changes significantly, based on a composite reward mechanism that includes the predicted solution difference magnitude, region coverage, and redundancy penalty. After optimizing the strategy using the MAPPO algorithm, the adaptive sampling points generated by the agents are merged with the initial point set for refined training of the PINN.

[0094] To verify the effectiveness of the above method, this embodiment compares the method of the present invention (MAPPO-2) with the traditional uniform sampling method (UNIFORM) and the existing residual adaptive refinement method (RAR). Reference Figure 4 and Figure 5 Experimental results show that the UNIFORM method performs well with 2000 sampling points. The error was 2.119e-02, and the time taken was 89.5 seconds. Although the RAR method reduced the error to 1.846e-02, a 12.9% improvement relative to the benchmark, the improvement was limited. In contrast, the MAPPO-2 method proposed in this invention, through the active exploration of the agent, can accurately capture nonlinear fluctuation characteristics and, with a moderate increase in the total number of sampling points, achieve a higher accuracy. The error was significantly reduced to 9.805e-03. Compared with the UNIFORM benchmark, the solution accuracy of this method was improved by 53.7%, and the total training time was shortened to 52.8 seconds, which fully demonstrates that the present invention can achieve both high-precision solution and excellent computational efficiency when dealing with complex nonlinear dynamic equations.

[0095] On the other hand, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the PINN adaptive sampling method based on multi-agent reinforcement learning described in the present invention.

[0096] The present invention can also provide a computer device, including a processor and a memory, wherein the memory is used to store a computer executable program, the processor reads the computer executable program from the memory and executes it, and the processor can implement the PINN adaptive sampling method based on multi-agent reinforcement learning described in the present invention when executing the computer executable program.

[0097] The computer device may be a laptop, a desktop computer, or a workstation.

[0098] The processor can be a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or an off-the-shelf programmable gate array (FPGA).

[0099] The memory described in this invention can be an internal storage unit of a laptop, desktop computer, or workstation, such as memory or hard disk; or it can be an external storage unit, such as a portable hard disk or flash memory card.

[0100] Computer-readable storage media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media can include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. Random access memory can include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).

[0101] The above content is only for illustrating the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solution based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.

Claims

1. A PINN adaptive sampling method based on multi-agent reinforcement learning, characterized in that, Includes the following steps: Construct a multi-agent sampling environment and model the sampling process as a Markov decision process; When setting up the sampling environment, a fixed time anchor point mechanism is introduced, and the agent only performs actions in the spatial dimension. Multiple agents are controlled to perform actions in a set sampling environment. The reward is calculated based on the local change characteristics of the predicted solution of the pre-trained PINN at a fixed time anchor point, and the total reward of the agent is obtained by combining the regional reward and the redundancy penalty. The policy network and value network of the agents are updated using a multi-agent reinforcement learning algorithm to obtain a trained adaptive sampling policy. An adaptive sampling point set is generated using a trained adaptive sampling strategy, and then merged with the initial sampling point set to form a hybrid sampling point set. The hybrid sampling point set is used to refine the training of PINN to obtain the target PINN model for solving the target partial differential equation.

2. The PINN adaptive sampling method based on multi-agent reinforcement learning according to claim 1, characterized in that, Obtaining pre-trained PINN includes: Obtain the target partial differential equation, determine the space-time solution domain, time interval, initial conditions and boundary conditions, and construct an equation operator expression form that can be used for automatic differentiation; An initial set of sampling points is generated within the solution domain, a physical information neural network is constructed, and the PINN is pre-trained using the initial set of sampling points to obtain a pre-trained PINN.

3. The PINN adaptive sampling method based on multi-agent reinforcement learning according to claim 2, characterized in that, Generating the initial set of sampling points within the solution domain includes: The method for generating the initial sampling points is as follows: in, The space-time solution domain is defined in step one; This represents the initial set of sampling points, where the sampling method is random sampling, uniform partitioning sampling, or Latin hypercube sampling.

4. The PINN adaptive sampling method based on multi-agent reinforcement learning according to claim 1, characterized in that, Construct a multi-agent sampling environment and model the sampling process as a Markov decision process; When setting the sampling environment, a fixed-time anchor point mechanism is introduced, and the agent only performs actions in the spatial dimension, including: At the beginning of each sampling round, a time value is randomly selected within the time interval. And remain unchanged during that sampling round; In the spatial domain Each agent is randomly assigned an initial coordinate. And place multiple agents together in the same environment to perform policy learning tasks; spatial domain According to the number of sections Divided into several sub-regions And assign a region identifier to each agent. ; The state space of an agent is defined as follows: ,in The current spatial location of the agent. This identifies the spatial sub-region to which the agent belongs; the agent's action space only includes movement in the spatial dimension.

5. The PINN adaptive sampling method based on multi-agent reinforcement learning according to claim 1, characterized in that, The formula for calculating the total reward of the intelligent agent is: in, The reward for local changes is calculated based on the local difference magnitude of the predicted solution; For regional rewards, the calculation is based on the distance between the agent and the center of its region; As a redundancy penalty, calculations are based on the distance between agents; Local change reward The calculation method is as follows: Suppose the intelligent agent is located Move to Calculate PINN at a fixed time anchor point The predicted solution difference magnitude ;when When the value exceeds a preset threshold, a positive reward is given; otherwise, the reward is 0. The regional rewards and redundancy penalty The calculation method is as follows: The regional reward is represented as ,in The distance from the agent's current location to the center of the region. This is the regional reward coefficient. The attenuation factor is denoted as ; the redundancy penalty is expressed as . ,in The distance between the current agent and its nearest neighbor agents. The penalty coefficient is... The interaction radius.

6. The PINN adaptive sampling method based on multi-agent reinforcement learning according to claim 1, characterized in that, By using multi-agent reinforcement learning algorithms to update the policy and value networks of agents, well-trained adaptive sampling policies are obtained, including: Multiple agents share a unified policy network Its update objective is to maximize expected return: Policy updates adopt the pruning objective of Proximal Policy Optimization (PPO), in the following form: For time step Expectations It is the probability ratio. For the dominant function, For the parameters of the policy network, For pruning hyperparameters; Building a value network The advantage function is used to predict the expected return of a state. It is calculated as follows: in, As a discount factor, For time step The value network is trained by minimizing the following loss to provide immediate rewards: in, For actual returns; In each policy iteration, both the policy network and the value network are updated simultaneously, and the following overall objective is optimized using the gradient descent algorithm: in For the policy entropy term, These are weighting coefficients, all greater than 0; after multiple rounds of updates, the agent can learn a stable adaptive sampling strategy. Multiple agents share a policy network and evaluate the global state through a centralized value network; the policy network's update objective function uses a shear ratio. in It is the probability ratio. This is the dominant function.

7. The PINN adaptive sampling method based on multi-agent reinforcement learning according to claim 1, characterized in that, An adaptive sampling point set is generated using a trained adaptive sampling strategy, and then fused with the initial sampling point set to form a hybrid sampling point set, including: Record the locations visited by the agent during the exploration process, and select points whose individual rewards meet preset conditions as adaptive sampling points: Will With a uniformly distributed initial set of sampling points The combined sampling point set is obtained. : 。 8. The PINN adaptive sampling method based on multi-agent reinforcement learning according to claim 1, characterized in that, The mixed sampling training loss function for PINN is: in The physical equation residual loss on the mixed sampling point set, Loss due to initial conditions This is the boundary condition loss.

9. A computer device, characterized in that, It includes a processor and a memory, the memory being used to store a computer-executable program, the processor reading part or all of the computer-executable program from the memory and executing it, and the processor executing part or all of the computer-executable program being able to implement the PINN adaptive sampling method based on multi-agent reinforcement learning as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program that, when executed by a processor, implements the PINN adaptive sampling method based on multi-agent reinforcement learning as described in any one of claims 1-7.