A model training method and related device based on a physics-informed neural network
By introducing multiple residual network channels and partial differential equations into the physical information neural network model, the problem of excessive time and low accuracy of electromagnetic simulation calculation in the existing technology is solved, and efficient and accurate calculation of electromagnetic simulation performance indicators is realized, supporting the optimized design of antennas.
Patent Information
- Application Number
- CN202111069844.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-09-13
AI Technical Summary
The prior art takes too long to calculate the electromagnetic simulation performance indicators, and the accuracy of calculation through the physical information neural network model is not high, making it difficult to effectively optimize the antenna.
Using a model training method based on physical information neural network, the target physical information neural network model is trained by introducing multiple residual network channels into the first neural network and using partial differential equations as part of the loss function.
It improves the accuracy and efficiency of model training, and can calculate high-precision electromagnetic simulation performance indicators more quickly, thus helping to optimize the antenna design.
Smart Images

Figure CN115809695B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a model training method and related devices based on a physics-informed neural network. Background Art
[0002] Electromagnetic simulation is the main technology for the design, optimization, and analysis of various antennas and antenna arrays. Through electromagnetic simulation, some performance indicators of the simulated antenna can be calculated, such as return loss, antenna energy efficiency, etc., so as to guide the design or optimization of the antenna.
[0003] The traditional calculation method of the performance indicators of electromagnetic simulation can be to first divide the simulation domain of the antenna into grids, and then solve the Maxwell equations on the discrete grids to calculate the full amount of electromagnetic fields for the next optimization analysis. Statistical results show that the discrete grid division usually takes from dozens of minutes to several hours. For a computational grid of about ten million, the solution of the governing equations takes 4 to 8 hours, and this calculation method takes too much time.
[0004] Currently, there is also a solution to calculate the performance indicators of electromagnetic simulation through a Physics Informed Neural Networks (PINNs) model. However, the accuracy of the performance indicators of electromagnetic simulation calculated by the currently trained PINNs model is not high, which is not conducive to the optimization of the antenna. Summary of the Invention
[0005] This application provides a model training method based on a Physics Informed Neural Networks (PINNs) to improve the accuracy of model training. This application also provides corresponding devices, computer equipment, computer-readable storage media, computer program products, etc.
[0006] The first aspect of the present application provides a model training method based on the physics-informed neural network (PINNs). The physics-informed neural network includes a first neural network and a partial differential equation. The first neural network includes at least two residual network channels. The method includes: obtaining a plurality of sampling point data from the simulation domain of the antenna. The plurality of sampling point data includes sampling point data of the active region, sampling point data of the passive region, data of the boundary of the simulation domain, and initial data of the simulation domain. The simulation domain includes the active region and the passive region; inputting the product of each training sample in the plurality of training samples and the coefficient corresponding to each residual network channel into each residual network channel of the first neural network. Each training sample includes a sampling point data and a latent vector corresponding to the simulation domain, and the coefficients corresponding to each residual network channel are different; processing the data input into each residual network channel through the first neural network to obtain an output data set. The output data set includes active output data, passive output data, boundary output data, and initial output data; processing the output data set through the partial differential equation to obtain a total loss function. The total loss function is related to the active loss function, the passive loss function, the boundary loss function, and the initial loss function; updating the parameters in the first neural network according to the total loss function to obtain a second neural network; using the second neural network as the first neural network, and iteratively executing the above training process until the second neural network reaches the convergence condition to obtain a target physics-informed neural network model for electromagnetic simulation of the antenna.
[0007] In the present application, PINNs is to add physical equations as constraints to the neural network so that the training results satisfy physical laws. This constraint is actually achieved by adding the residuals before and after the iteration of the physical equation to the loss function of the neural network, allowing the physical equation to also "participate" in the training process. In this way, when the neural network is training and iterating, it optimizes not only the loss function of the network itself, but also the residuals of each iteration of the physical equation, so that the finally trained results satisfy physical laws.
[0008] In the present application, the first neural network is used to represent the neural network before one iteration, and the second neural network is used to represent the neural network after one iteration. The first neural network includes a plurality of residual network channels. The plurality in the present application includes two or more. Each residual network channel can convert the input data into output data in electromagnetic form.
[0009] In the present application, the partial differential equation can be the point-source Maxwell equation.
[0010] In the present application, the simulation domain of the antenna refers to the coverage area of the simulated antenna electromagnetic wave. The antenna can be understood as the antenna of the terminal or the antenna of the network device. The antennas of different terminals or network devices are usually different, so the simulation domains of different antennas are also different.
[0011] In this application, the simulation domain includes an active region, a passive region, and a boundary. The active region refers to the near-source region that simulates adding an excitation source in the antenna array and is affected by the excitation source, including the excitation source. The boundary refers to the edge of the simulation domain, and the passive region refers to the region in the simulation domain other than the active region and the boundary. The boundary of the simulation domain usually has a bounce boundary or an absorbing boundary, and different types of boundaries have a great impact on the results of electromagnetic simulation.
[0012] In this application, the simulation domain of the antenna can include the simulation domains of multiple different antennas, and the corresponding hidden vectors of each simulation domain can be different.
[0013] In this application, the sampling point data refers to the data corresponding to the sampling points. The sampling point data has four types: the sampling point data in the active region, the sampling point data in the passive region, the data on the boundary of the simulation domain, and the initial data of the simulation domain. The initial data of the simulation domain usually refers to the electric field data and magnetic field data in the initial state of the simulation domain (usually at t = 0 in the time dimension), and the electric field data and magnetic field data of the simulation domain in the initial state are usually zero. The sampling point data is usually four-dimensional, including the three-dimensional spatial coordinates of the sampling point and the one-dimensional time information of the sampling point. The form of the sampling point data can be expressed as U=(x, y, z, t).
[0014] In this application, the training sample refers to the sample data used to train the model. The training sample not only includes the sampling point data but also includes the hidden vector Z corresponding to the simulation domain. The training sample can be expressed in the form of X=(Z, U).
[0015] In this application, the hidden vector Z is used to characterize the parameter settings of different electromagnetic simulation scenarios. In this application, the hidden vector Z adopts a low-dimensional vector, and the commonly used dimension selections can be 16, 32, 64, 128, etc.
[0016] In this application, because there are four types of sampling point data, there are also four types of training samples, namely the training samples containing the sampling point data in the active region, the training samples containing the sampling point data in the passive region, the training samples containing the data on the boundary of the simulation domain, and the training samples containing the initial data of the simulation domain.
[0017] In this application, each type of training sample is input into each residual network channel one by one. Each residual network channel will obtain the output data of this type, and then the output data of each residual network channel is summarized to obtain an output data corresponding to the input. Therefore, there are also four types of output data, namely active output data, passive output data, boundary output data, and initial output data. In addition, the coefficients of each residual network channel are different, so that the same training sample can be differentially changed, thereby improving the training accuracy of the model.
[0018] In this application, since there are four types of training samples, there are also four types of output data and four types of loss functions. The total loss function is obtained through the four types of loss functions and then the parameters in the first neural network are updated to obtain the second neural network.
[0019] In this application, the method of gradient descent can be used to update the parameters in the first neural network.
[0020] In this application, the target PINNs model is relative to the initial PINNs model before starting model training. The parameters in the first neural network of the initial PINNs are usually large. During the model training process, through the training samples, the parameters in the first neural network are continuously updated until the convergence condition is reached to obtain the second neural network. At this time, the parameters in the second neural network can be understood as fixed, and the entire model at this time is called the target PINNs model.
[0021] As can be seen from the above description of the first aspect, since the first neural network of PINNs includes multiple residual network channels and the coefficients corresponding to each residual network channel are different, in this way, during the model training stage, the same training samples can be multiplied by different coefficients, and one data can be expanded into multiple data. Moreover, different frequency signals can be captured through multiple residual network channels, thereby improving the accuracy of model training.
[0022] In a possible implementation manner of the first aspect, the active region is a region in the simulation domain centered on the point source corresponding to the excitation source with a first length as the radius. The first length is related to the first parameter in the continuous probability density function, and the continuous probability density function approaches the Dirac function. The function of the point source is the product of the continuous probability density function and the signal of the excitation source; the passive region is the region in the simulation domain except the active region and the boundary.
[0023] In this application, the excitation source is regarded as a point source, and the function of the point source can be expressed as J(x,t) = η α (x)g(t). Compared with the existing function of the point source J(x,t) = δ(x - x 0 ), the Dirac function δ(x - x 0 ) is replaced by the continuous probability density function η α (x). Among them, J(x,t) represents the function of the point source, δ(x - x 0 ) represents the Dirac function, g(t) represents the signal of the excitation source, and x 0 represents the position of the excitation source. The function of this point source represents applying an excitation source signal in the form of g(t) at x 0 in the simulation domain.
[0024] In this application, the continuous probability density function η α(x) replaces δ(x - x 0 ), and this continuous probability density function approaches the Dirac function, which can be expressed as δ(x - x 0 ) ~ η α (x). This η α (x) represents an abstracted typical distribution, and the specific form can be the form of a Gaussian distribution, the form of a Cauchy distribution, or the form of an exponential distribution.
[0025] In this possible implementation, the continuous probability density function η α (x) that approaches the Dirac function is used to replace the Dirac function, overcoming the bottleneck that PINNs cannot handle point source problems.
[0026] In a possible implementation of the first aspect, when the active output data is the sum of the output data of each residual network channel when a training sample in multiple training samples contains sampling point data of the active region, the passive output data is the sum of the output data of each residual network channel when a training sample in multiple training samples contains sampling point data of the passive region, the boundary output data is the sum of the output data of each residual network channel when a training sample in multiple training samples contains boundary data, and the initial output data is the sum of the output data of each residual network channel when a training sample in multiple training samples contains initial data.
[0027] In a possible implementation, it can be multiplying the data output by each residual network channel by some coefficients and then summing them up.
[0028] In this possible implementation, the output data of each residual network channel can be summed, or the output data of each residual network channel can be multiplied by some coefficients and then summed up. This method of taking the sum of the output data of multiple residual network channels and then taking the partial derivative in this application can improve the accuracy of model training.
[0029] In a possible implementation of the first aspect, each residual network channel includes a sine periodic activation function; the sine periodic activation function is used to convert the data in each residual network channel into electric field parameters and magnetic field parameters as the output data of each residual network channel.
[0030] In this possible implementation, each residual network channel can include a residual network and a sine periodic activation function. The residual network can optimize the first neural network model and improve the performance of the first network model. The sine periodic activation function can obtain electric field data and magnetic field data. This combination of the residual network and the sine periodic activation function can effectively improve the accuracy of the model.
[0031] In a possible implementation of the first aspect, the coefficients corresponding to each residual network channel increase exponentially.
[0032] In this possible implementation, the coefficients corresponding to each of the multiple residual network channels increase exponentially. For example, if there are four residual network channels, the coefficients of the four residual network channels can be 1, 2, 4, and 8 respectively. This exponentially increasing method is conducive to quickly widening the gap of the same data, thereby improving the accuracy of model training.
[0033] In a possible implementation of the first aspect, the above step: processing the output data set through a partial differential equation to obtain a total loss function, includes: each time taking an output data in the output data set as a known quantity of the partial differential equation, performing operations on the partial differential equation to obtain a loss function corresponding to an output data; accumulating the loss functions corresponding to each output data in the output data set according to a preset relationship to obtain a total loss function.
[0034] In a possible implementation of the first aspect, the preset relationship includes learnable parameters and hyperparameters. The learnable parameters corresponding to different loss functions related to the total loss function are different, and the learnable parameters will be updated as the parameters in the first neural network are updated. The hyperparameters are used to assist in weighting the loss functions corresponding to the learnable parameters.
[0035] In a possible implementation of the first aspect, when updating the parameters in the first neural network according to the total loss function, the method further includes: updating the hidden vector of the simulation domain and the learnable parameters in the preset relationship.
[0036] The second aspect of the present application provides an incremental learning method, which includes: obtaining a plurality of sampling point data from the simulation domain of the antenna to be optimized, the plurality of sampling point data including sampling point data of the active region, sampling point data of the passive region, data of the boundary of the simulation domain, and initial data of the simulation domain, and the simulation domain includes an active region and a passive region; inputting a plurality of sample data into the target physical information neural network, where each sample data includes a sampling point data and a first hidden vector of the simulation domain, and the target physical information neural network is the target physical information neural network model trained by the first aspect or any possible implementation of the first aspect; obtaining output data corresponding to each sample data through the target physical information neural network; keeping the parameters in the control standard physical information neural network unchanged, adjusting the first hidden vector of the simulation domain according to the output data to obtain a second hidden vector; taking the second hidden vector as the first hidden vector, and iteratively performing the above adjustment of the first hidden vector through different sample data until the output data meets the preset requirements of the antenna to be optimized to obtain a second hidden vector matching the simulation domain.
[0037] In this second aspect, during the incremental learning process, the parameters in the target physics-informed neural network are frozen, and the latent vector of the simulation domain of the antenna to be optimized is repeatedly adjusted through the output data of the target physics-informed neural network until a latent vector matching the simulation domain is obtained. This method can quickly learn the latent vector and improve the acquisition speed of the latent vector for the new electromagnetic simulation scenario.
[0038] In the third aspect of the present application, a method for electromagnetic simulation is provided. The method includes simulating an antenna using the target physics-informed neural network model trained by using the first aspect or any possible implementation manner of the first aspect above to obtain the electromagnetic field distribution of the antenna.
[0039] In the fourth aspect of the present application, a model training device based on a physics-informed neural network is provided. The device has the function of implementing the method of the first aspect or any possible implementation manner of the first aspect above. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, such as: an acquisition unit and one or more processing units.
[0040] In the fifth aspect of the present application, an incremental learning device is provided. The device has the function of implementing the method of the second aspect above. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, such as: an acquisition unit and one or more processing units.
[0041] In the sixth aspect of the present application, an electromagnetic simulation device is provided. The device has the function of implementing the method of the third aspect above. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, such as: one or more processing units.
[0042] In the seventh aspect of the present application, a computer device is provided. The computer device includes at least one processor, a memory, an input / output (I / O) interface, and computer execution instructions stored in the memory and executable on the processor. When the computer execution instructions are executed by the processor, the processor executes the method of the first aspect or any possible implementation manner of the first aspect above.
[0043] In the eighth aspect of the present application, a computer device is provided. The computer device includes at least one processor, a memory, an input / output (I / O) interface, and computer execution instructions stored in the memory and executable on the processor. When the computer execution instructions are executed by the processor, the processor executes the method of the second aspect above.
[0044] The ninth aspect of this application provides a computer device, which includes at least one processor, a memory, an input / output (I / O) interface, and computer-executable instructions stored in the memory and executable on the processor. When the computer-executable instructions are executed by the processor, the processor executes the method of the third aspect as described above.
[0045] The tenth aspect of this application provides a computer-readable storage medium storing one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the processor executes the method of the first aspect or any possible implementation manner of the first aspect as described above.
[0046] The eleventh aspect of this application provides a computer-readable storage medium storing one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the processor executes the method of the second aspect as described above.
[0047] The twelfth aspect of this application provides a computer-readable storage medium storing one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the processor executes the method of the third aspect as described above.
[0048] The thirteenth aspect of this application provides a computer program product storing one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the processor executes the method of the first aspect or any possible implementation manner of the first aspect as described above.
[0049] The fourteenth aspect of this application provides a computer program product storing one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the processor executes the method of the second aspect as described above.
[0050] The fifteenth aspect of this application provides a computer program product storing one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the processor executes the method of the third aspect as described above.
[0051] The sixteenth aspect of this application provides a chip system, which includes at least one processor. The at least one processor is used to implement the functions involved in the first aspect or any possible implementation manner of the first aspect. In a possible design, the chip system may further include a memory for storing the necessary program instructions and data for the device to process the artificial intelligence model. The chip system may be composed of chips or may include chips and other discrete devices.
[0052] The seventeenth aspect of the present application provides a chip system, which includes at least one processor for implementing the functions involved in the second aspect above. In a possible design, the chip system may further include a memory for storing the necessary program instructions and data for the device of data processing based on the artificial intelligence model. The chip system may be composed of chips or may include chips and other discrete devices.
[0053] The eighteenth aspect of the present application provides a chip system, which includes at least one processor for implementing the functions involved in the second aspect above. In a possible design, the chip system may further include a memory for storing the necessary program instructions and data for the device of data processing based on the artificial intelligence model. The chip system may be composed of chips or may include chips and other discrete devices. Description of the Drawings
[0054] Figure 1 is a schematic structural diagram of the physical information neural network model provided by an embodiment of the present application;
[0055] Figure 2 is a schematic diagram of model training provided by an embodiment of the present application;
[0056] Figure 3 is a schematic diagram of the simulation domain of an antenna provided by an embodiment of the present application;
[0057] Figure 4 is a schematic diagram of an embodiment of the model training method provided by an embodiment of the present application;
[0058] Figure 5 is a schematic example diagram of the model training method provided by an embodiment of the present application;
[0059] Figure 6 is a schematic example diagram of the power Maxwell equation provided by an embodiment of the present application;
[0060] Figure 7 is a schematic diagram of an embodiment of the incremental learning method provided by an embodiment of the present application;
[0061] Figure 8 is another schematic diagram of an embodiment of the incremental learning method provided by an embodiment of the present application;
[0062] Figure 9 is a comparison diagram of experimental effects provided by an embodiment of the present application;
[0063] Figure 10 is a schematic diagram of an embodiment of the electromagnetic simulation provided by an embodiment of the present application;
[0064] Figure 11It is a schematic diagram of an embodiment of a model training device provided by an embodiment of the present application;
[0065] Figure 12 It is a schematic diagram of an embodiment of an incremental learning device provided by an embodiment of the present application;
[0066] Figure 13 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0067] Next, the embodiments of the present application will be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Those of ordinary skill in the art can know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0068] Terms such as "first" and "second" in the specification, claims and the above drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0069] The embodiments of the present application provide a model training method based on Physical Informed Neural Networks (PINNs) to improve the accuracy of model training, thereby improving the accuracy of electromagnetic simulation. The present application also provides corresponding devices, computer devices, computer-readable storage media, computer program products, etc. The following will be described in detail respectively.
[0070] Antennas can be optimized through electromagnetic simulation. Currently, neural network models can be pre-trained through artificial intelligence (AI) technology, and the neural network models are used to complete the process of electromagnetic simulation, determine simulation results such as the electromagnetic field distribution and performance indicators of the antenna to be optimized, and then optimize the antenna according to the simulation results.
[0071] Because the electromagnetic field distribution has strong physical properties, most of the neural network models for electromagnetic simulation are PINNs models. PINNs is to add physical equations as constraints into the neural network so that the training results satisfy physical laws. And this constraint is actually achieved by adding the residuals before and after the iteration of the physical equations to the loss function of the neural network, allowing the physical equations to also "participate" in the training process. In this way, when the neural network is training and iterating, it optimizes not only the loss function of the network itself, but also the residuals of each iteration of the physical equations, making the finally trained results satisfy physical laws.
[0072] To better use the PINNs model for electromagnetic simulation, the embodiments of this application provide the following aspects: First, provide a PINNs model with a new structure; Second, train the PINNs model with the new structure based on the simulation domain of the antenna to obtain the target PINNs model; Third, use the target PINNs model for incremental learning to obtain the latent vector of the new electromagnetic simulation scenario; Fourth, use the target PINNs model for electromagnetic simulation to obtain the electromagnetic field data at each point in the antenna simulation domain. The processes of model training, incremental learning, and electromagnetic simulation can all be carried out on a computer device, which can be a server, a terminal device, or a virtual machine (VM).
[0073] A terminal device (which can also be called a user equipment (UE)) is a device with wireless transceiver functions. It can be deployed on land, including indoor or outdoor, handheld or vehicle-mounted; it can also be deployed on the water surface (such as a ship, etc.); it can also be deployed in the air (such as an airplane, a balloon, a satellite, etc.). The terminal can be a mobile phone, a tablet (pad), a computer with wireless transceiver functions, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc.
[0074] A VM can be a virtualized device divided in a virtualized manner on the hardware resources of a physical machine.
[0075] The following will introduce the content involved in the embodiments of the present application in sequence.
[0076] 1. PINNs model with a new structure.
[0077] The PINNs model with a new structure provided by the embodiments of the present application can be understood by referring to Figure 1 As shown in Figure 1 , the PINNs model with a new structure provided by the embodiments of the present application can include a first neural network and a partial differential equation (PDE). The first neural network includes at least two residual network channels. As shown in Figure 1 , the first neural network includes n residual network channels, such as: residual network channel 1, residual network channel 2,..., residual network channel n. The partial differential equation can be the point source Maxwell equation.
[0078] Each residual network channel has a corresponding coefficient, and the coefficients corresponding to the n residual network channels can increase exponentially. For example, when n = 4, there are four residual network channels, and the coefficients corresponding to these four residual network channels can be 1, 2, 4, and 8 respectively. When n = 5, there are five residual network channels, and the coefficients corresponding to these five residual network channels can be 1, 2, 4, 8, and 16 respectively.
[0079] Each residual network channel can include a residual network and a sine periodic activation function. Among them, the residual network and the sine periodic activation function can be expressed as x→φ i (x)=x + sin(W i x + b i ), where x represents the residual, and sin(W i x + b i ) represents the sine periodic activation function.
[0080] In the embodiments of the present application, the residual network can optimize the first neural network and improve the performance of the first neural network. The sine periodic activation function is used to convert the data in each residual network channel into electric field parameters and magnetic field parameters as the output data of each residual network channel. This combination of the residual network and the sine periodic activation function can effectively improve the accuracy of the model.
[0081] 2. Training the PINNs model with the new structure based on the simulation domain of the antenna to obtain the target PINNs model.
[0082] The process of model training provided by the embodiments of the present application can be understood by referring to Figure 2 As shown in Figure 2As shown, training samples are input into the PINNs model, and the training samples are processed by the first neural network to obtain output data. The output data is processed by partial differential equations to obtain a loss function, and then the parameters in the first neural network are updated by the loss function. The computer device iteratively executes this training process until the convergence condition is reached, and a target PINNs model is obtained.
[0083] In the embodiment of the present application, the training samples for training the PINNs model are from the simulation domain of the antenna. The simulation domain of the antenna refers to the coverage area of the simulated antenna electromagnetic wave. The antenna can be understood as the antenna of the terminal or the antenna of the network device. Antennas of different terminals or network devices are usually different, so the simulation domains of different antennas are also different.
[0084] The antenna in the embodiment of the present application can be an antenna with a pulse excitation source. In this way, the simulation domain of the antenna includes an active region, a passive region, and a boundary. The structure of the antenna can be, for example, Figure 3 the butterfly structure 100 shown in the figure. The antenna of this butterfly structure includes two opposite triangular structures. Simulating the area covered by the electromagnetic wave of this antenna can be understood as the simulation domain 101 of this butterfly antenna. For example, Figure 3 in the middle position between the two triangles, a source can be added. This excitation source can be understood as a point source 102. The near-source region containing the point source 102 is the active region 103, and the region in the simulation domain 101 except the active region 103 and the boundary of this simulation domain 101 is the passive region 104.
[0085] It can also be understood that: the active region is the region in the simulation domain centered on the point source corresponding to the excitation source with a first length as the radius. The first length is related to the first parameter in the continuous probability density function. The continuous probability density function approaches the Dirac function, and the function of the point source is the product of the continuous probability density function and the signal of the excitation source; the passive region is the region in the simulation domain except the active region and the boundary, or the region inside the simulation domain except the active region after removing the boundary.
[0086] In the embodiment of the present application, the simulation domain after removing the boundary can be represented by Ω, the active region can be represented by Ω 0 and the passive region can be represented by Ω 1 In this way, Ω 0 ={(x 0 +x)∈Ω,||x||≤3α}, Ω 1 =Ω - Ω 0 . Among them, x 0represents the center of the point source corresponding to the excitation source, x represents the radius of the first length, and α represents the first parameter in the continuous probability density function. In the embodiments of the present application, the value of α can be set according to requirements, usually 1 / 100 to 1 / 200 of the length of the simulation domain, and the time range and space range of the simulation domain can both be determined according to the antenna.
[0087] In the embodiments of the present application, the excitation source is regarded as a point source, and the function of the point source can be expressed as J(x,t) = η α (x)g(t). Compared with the existing function of the point source J(x,t) = δ(x - x 0 ), the Dirac function δ(x - x 0 ) is replaced by the continuous probability density function η α (x). Among them, J(x,t) represents the function of the point source, δ(x - x 0 ) represents the Dirac function, g(t) represents the signal of the excitation source, and x 0 represents the position of the excitation source. The function of this point source represents applying an excitation source signal in the form of g(t) at x 0 in the simulation domain.
[0088] In the embodiments of the present application, the continuous probability density function η α (x) is used to replace δ(x - x 0 ). This continuous probability density function approaches the Dirac function and can be expressed as δ(x - x 0 ) ∼ η α (x). This η α (x) represents an abstracted typical distribution, and the specific form can be in the form of a Gaussian distribution, a Cauchy distribution, or an exponential distribution. The forms of several distributions can be understood by referring to Table 1 below.
[0089] Table 1:
[0090]
[0091] In the embodiments of the present application, by using the continuous probability density function η α (x) that approaches the Dirac function to replace the Dirac function, the bottleneck that PINNs cannot handle point source problems is overcome.
[0092] An embodiment of the model training method based on PINNs provided in the embodiments of the present application can be understood by referring to Figure 4 . As Figure 4 shown, an embodiment of the model training method based on PINNs provided in the embodiments of the present application may include:
[0093] 201. The computer device obtains a plurality of sampling point data from the simulation domain of the antenna.
[0094] Among them, the data of multiple sampling points include the sampling point data of the active region, the sampling point data of the passive region, the data of the boundary of the simulation domain, and the initial data of the simulation domain. The simulation domain includes the active region and the passive region.
[0095] In the embodiments of the present application, there are four types of sampling point data, namely, the sampling point data of the active region, the sampling point data of the passive region, the data of the boundary of the simulation domain, and the initial data of the simulation domain. The boundary of the simulation domain usually has a reflection boundary or an absorption boundary. Different types of boundaries have a great impact on the results of electromagnetic simulation. The initial data of the simulation domain usually refers to the electric field data and magnetic field data in the initial state of the simulation domain (usually at t = 0 in the time dimension). The electric field data and magnetic field data of the simulation domain in the initial state are usually zero. The sampling point data is usually four-dimensional, including the three-dimensional spatial coordinates of the sampling point and the one-dimensional time information of the sampling point. The form of the sampling point data can be expressed as U = (x, y, z, t). According to the type of the sampling point data, the sampling point data of the active region can be expressed as U SRC , the sampling point data U NO_SRC of the passive region, the boundary data U BC of the simulation domain, and the initial data U IC of the simulation domain.
[0096] The computer device inputs the product of each training sample and the corresponding coefficient of each residual network channel in each residual network channel of the first neural network.
[0097] Among them, each training sample includes a sampling point data and a latent vector corresponding to the simulation domain.
[0098] The training sample refers to the sample data used to train the PINNs model. The training sample not only includes the sampling point data but also includes the latent vector Z corresponding to the simulation domain. The training sample can be expressed in the form of X = (Z, U). According to the type of the training sample, the training sample containing U SRC can be expressed as X SRC = (Z, U SRC ), the training sample containing U NO_SRC can be expressed as X NO_SRC = (Z, U NO_SRC ), the training sample containing U BC can be expressed as X BC = (Z, U BC ), and the training sample containing U IC can be expressed as X IC = (Z, U IC ).
[0099] The latent vector Z is used to represent the parameter settings of different electromagnetic simulation scenarios. In the embodiments of the present application, the latent vector Z adopts a low-dimensional vector, and the commonly selected dimensions can be 16, 32, 64, 128, etc.
[0100] The coefficients corresponding to each residual network channel are different. For example, Figure 5 as shown, there are n residual network channels in the first neural network, from residual network channel 1 to residual network channel n. Among them, the coefficient corresponding to residual network channel 1 is a 1 , the coefficient corresponding to residual network channel 2 is a 2 , …, the coefficient corresponding to residual network channel n is a n . The coefficients of these n residual network channels can also be represented in the form of a set as {a 1 , a 2 , …, a n}. In this way, when the training sample is X, the input of each residual network channel can be expressed as {a 1 X, a 2 X, …, a n X}. This X can be any one of the above X SRC , X NO_SRC , X BC and X IC .
[0101] If the training samples come from multiple electromagnetic simulation scenarios, that is, from the simulation domains of multiple different antennas, then, for each different simulation domain, there will be a corresponding latent vector. For example, if there are N different simulation domains, then the N latent vectors can be expressed as {Z 1 , … Z N}. When there are N simulation domains, the training samples from the i-th simulation domain can be expressed as {X i,SRC = (Z i , U i,SRC ), X i,NO_SRC = (Z i , U i,NO_SRC ), X i,IC = (Z i , U i,IC ), X i,BC = (Z i , U i,BC )}.
[0102] 203. The computer device processes the data input into each residual network channel through the first neural network to obtain an output data set.
[0103] Among them, the output data set includes active output data, passive output data, boundary output data, and initial output data.
[0104] Optionally, in the embodiments of the present application, when the active output data is the sampled point data of the active region included in one training sample among multiple training samples, the output data of each residual network channel is the sum; when the passive output data is the sampled point data of the passive region included in one training sample among multiple training samples, the output data of each residual network channel is the sum; when the boundary output data is the boundary data included in one training sample among multiple training samples, the output data of each residual network channel is the sum; when the initial output data is the initial data included in one training sample among multiple training samples, the output data of each residual network channel is the sum.
[0105] In the embodiments of the present application, it is not limited to directly adding and summing the output data of each residual network channel. It can also be multiplying the data output by each residual network channel by some coefficients and then adding and summing.
[0106] The output data set can be expressed as {Y SRC , Y NO_SRC , Y BC , Y IC}. Among them, each Y can be obtained by multiplying the coefficient of each residual network channel by the corresponding type of X and then adding and summing the output of each residual network channel, and can be expressed as Y = Y 1 + Y 2 … + Y n , where Y 1 represents the output data of the first residual network channel, and Y n represents the output data of the nth residual network channel.
[0107] The computer device processes the output data set through partial differential equations to obtain the total loss function.
[0108] Among them, the total loss function is obtained based on the active loss function, the passive loss function, the boundary loss function, and the initial loss function.
[0109] The active loss function refers to the loss function obtained through the active output data, the passive loss function refers to the loss function obtained through the passive output data, the boundary loss function refers to the loss function obtained through the boundary output data, and the initial loss function refers to the loss function obtained through the initial output data. The active loss function can be represented by L SRC , the passive loss function can be represented by L NO_SRC , the boundary loss function can be represented by L BC , and the initial loss function can be represented by L IC .
[0110] Optionally, the process of obtaining the total loss function may be as follows: each time an output data in the output dataset is used as a known quantity of the partial differential equation, the partial differential equation is operated on to obtain a loss function corresponding to an output data; the loss functions corresponding to each output data in the output dataset are accumulated according to a preset relationship to obtain the total loss function. Among them, the preset relationship includes learnable parameters, and the learnable parameters corresponding to different loss functions are different.
[0111] The partial differential equation may be the point-source Maxwell equation. The output data Y is usually six-dimensional and includes three-dimensional electric field data and three-dimensional magnetic field data. As Figure 6 shown, substituting the electric field data and magnetic field data in the output data Y as known quantities into the power Maxwell equation as Figure 6 shown, and then performing calculations, the corresponding loss function can be calculated. Figure 6 In it, E represents the electric field, H represents the magnetic field, and the subscripts x, y, and z respectively represent three-dimensional space.
[0112] The total loss function may be obtained by accumulating according to a preset relationship. The preset relationship includes learnable parameters and hyperparameters. The learnable parameters corresponding to different loss functions related to the total loss function are different. The learnable parameters will be updated as the parameters in the first neural network are updated. The hyperparameters are used to assist the learnable parameters in weighting the corresponding loss functions.
[0113] The preset relationship can be expressed as:
[0114]
[0115] where, L total represents the total loss function, L i represents four types of loss functions, ε is a hyperparameter, and the value of this hyperparameter can be 0.01. Of course, this is only an example of the value of the hyperparameter. In this application, the specific value of the hyperparameter is not limited. λ i is a learnable parameter, and i = 1, 2, 3, 4.
[0116] In the embodiments of this application, the dynamic weighting of the loss function is achieved through hyperparameters and learnable parameters, and the weights of each loss function are balanced, which can accelerate the convergence speed in the neural network training process.
[0117] 205. The computer device updates the parameters in the first neural network according to the total loss function to obtain a second neural network.
[0118] In this application, the first neural network is used to represent the neural network before one iteration, and the second neural network is used to represent the neural network after one iteration.
[0119] In addition, when updating the parameter θ in the first neural network, the latent vector Z in the simulation domain and the learnable parameter λ in the above preset relationship can also be updated. That is, according to L total Update θ, Z, and λ.
[0120] In the embodiments of the present application, to update θ, Z, and λ, the gradient descent method can be used to adjust downward based on the θ, Z, and λ in the current iteration to obtain new θ, Z, and λ, and then start the next iteration process.
[0121] Take the second neural network as the first neural network and iteratively execute the above training process until the second neural network reaches the convergence condition to obtain the target physics-informed neural network model.
[0122] In the embodiments of the present application, the target PINNs model is relative to the initial PINNs model before starting model training. Usually, the parameters in the first neural network of the initial PINNs are relatively large. During the model training process, through training samples, the parameters in the first neural network are continuously updated until the convergence condition is reached to obtain the second neural network. At this time, the parameters in the second neural network can be understood as being fixed, and the entire model at this time is called the target PINNs model.
[0123] In the embodiments of the present application, since the first neural network of PINNs includes multiple residual network channels and the coefficients corresponding to each residual network channel are different, in the model training stage, the same training samples can be multiplied by different coefficients, so that one data can be expanded into multiple data, and different frequencies of signals can be captured through multiple residual network channels, thereby improving the accuracy of model training.
[0124] III. Perform incremental learning using the target PINNs model.
[0125] As Figure 7 shown, an embodiment of the incremental learning provided by the embodiments of the present application includes:
[0126] 301. The computer device obtains multiple sampling point data from the simulation domain of the antenna to be optimized.
[0127] The multiple sampling point data includes sampling point data in the active region, sampling point data in the passive region, data on the boundary of the simulation domain, and initial data of the simulation domain. The simulation domain includes the active region and the passive region.
[0128] The sampling point data in the embodiments of the present application can be understood by referring to the sampling point data in step 201 above. However, the sampling point data in the embodiments of the present application comes from the simulation domain of the antenna to be optimized, or from the simulation domain of a new electromagnetic simulation scenario.
[0129] 302. The computer device inputs multiple sample data into the target physical information neural network. Each sample data includes a sampling point data and a first hidden vector of the simulation domain.
[0130] The target physical information neural network is the target physical information neural network obtained by the model training method based on PINNs.
[0131] 303. The computer device obtains output data corresponding to each sample data through the target physical information neural network.
[0132] 304. The computer device controls the parameters in the target physical information neural network to remain unchanged, and adjusts the first hidden vector of the simulation domain according to the output data to obtain a second hidden vector.
[0133] In the embodiment of the present application, the adjustment of the hidden vector can be performed by the gradient descent method.
[0134] Take the second hidden vector as the first hidden vector, and iteratively execute the above adjustment of the first hidden vector through different sample data until the output data meets the preset requirements of the antenna to be optimized, so as to obtain a second hidden vector that matches the simulation domain.
[0135] In the embodiment of the present application, the first hidden vector can be understood as the hidden vector before iteration, and the second hidden vector can be understood as the hidden vector after iteration.
[0136] In the embodiment of the present application, during the incremental learning process, the parameters in the target physical information neural network are frozen, and the hidden vector of the simulation domain of the antenna to be optimized is repeatedly adjusted through the output data of the target physical information neural network until a hidden vector that matches the simulation domain is obtained. This method can quickly learn the hidden vector and improve the acquisition speed of the hidden vector in the new electromagnetic simulation scenario.
[0137] The above process of incremental learning can refer to Figure 8 for understanding. For example, Figure 8 as shown, for a new electromagnetic simulation scenario, the trained target PINNs model can be used, and keep θ in the target PINNs model unchanged in each round of iteration. Input {Xnew, SRC , =(Znew, Unew - SRC ), Xnew, NO_SRC =(Znew, Unew - NO_SRC ), Xnew, IC =(Znew, U new - IC ), Xnew, BC =(Znew, Unew - BC)}, adjust Z in the input data X through the output data Y of the target PINNs model until the output data Y meets the preset requirements, and obtain the latent vector Z that matches the new electromagnetic simulation scenario.
[0138] Regarding the incremental learning solution, the developers conducted relevant experiments, such as Figure 9 shown in the time comparison diagram of the latent vector Z of the new electromagnetic simulation scenario obtained by using the incremental learning solution provided in this application compared with the latent vector Z of the new electromagnetic simulation scenario obtained by the original method. From Figure 9 it can be seen that in the case of a 5% error, the solution of this application only takes 200 seconds to obtain the latent vector Z of the new electromagnetic simulation scenario, while it takes 3337 seconds to obtain the latent vector Z of the new electromagnetic simulation scenario by using the original method. The solution of this application has greatly improved in speed.
[0139] IV. Use the target PINNs model for electromagnetic simulation to obtain the electromagnetic field data of each point in the antenna simulation domain.
[0140] After training the target PINNs model through the above model training process, the target PINNs model can be stored in the form of a model file. When a computer device (such as a terminal device, a server, or a VM, etc.) used for electromagnetic simulation needs to use the target PINNs model, it can be that the computer device used for electromagnetic simulation actively loads the model file of the target PINNs model. It can also be that the model file storing the target PINNs model actively sends the model file of the target PINNs model to the computer device used for electromagnetic simulation to install the model file of the target PINNs model.
[0141] Such as Figure 10 shown, after installing the target PINNs model on the computer device, the target PINNs model can be used for electromagnetic simulation. The simulation result can be Figure 10 the schematic diagram of the electromagnetic field distribution shown, or some performance indicators of the simulated antenna, such as: the electromagnetic field data of each point in the antenna simulation domain. In the embodiments of this application, the electromagnetic field data includes electric field data and magnetic field data, such as: electric field intensity and magnetic field intensity, etc. In this way, the antenna can be optimized and designed through the results of electromagnetic simulation.
[0142] The electromagnetic simulation solution provided by the embodiments of this application uses the target PINNs model with multiple residual network channels to execute the electromagnetic simulation process, which greatly improves the accuracy of electromagnetic simulation.
[0143] The above describes the model training method based on the physics-informed neural network and the incremental learning method. Next, in combination with the attached Figure 11 introduce the model training device 40 based on the physics-informed neural network provided by the embodiments of this application. The model training device 40 based on the physics-informed neural network includes:
[0144] An acquisition unit 401, configured to acquire multiple sampling point data from the simulation domain of the antenna. The multiple sampling point data includes sampling point data of the active region, sampling point data of the passive region, data of the boundary of the simulation domain, and initial data of the simulation domain. The simulation domain includes an active region and a passive region. The function of the acquisition unit 401 can be understood by referring to step 201 in the above method embodiment.
[0145] A first processing unit 402, configured to input the product of each training sample in multiple training samples and the coefficient corresponding to each residual network channel into each residual network channel of the first neural network. Each training sample includes one sampling point data acquired by the acquisition unit 401 and a latent vector corresponding to the simulation domain. The coefficients corresponding to each residual network channel are different. The function of the first processing unit 402 can be understood by referring to step 202 in the above method embodiment.
[0146] A second processing unit 403, configured to process the data input into each residual network channel by the first neural network to obtain an output data set. The output data set includes active output data, passive output data, boundary output data, and initial output data. The function of the second processing unit 403 can be understood by referring to step 203 in the above method embodiment.
[0147] A third processing unit 404, configured to process the output data set through a partial differential equation to obtain a total loss function. The total loss function is obtained based on an active loss function, a passive loss function, a boundary loss function, and an initial loss function. The function of the third processing unit 404 can be understood by referring to step 204 in the above method embodiment.
[0148] A fourth processing unit 405, configured to update the parameters in the first neural network according to the total loss function to obtain a second neural network. The function of the fourth processing unit 405 can be understood by referring to step 205 in the above method embodiment.
[0149] Use the second neural network as the first neural network, and iteratively execute the above training process until the second neural network reaches the convergence condition to obtain a target physical information neural network model.
[0150] In the embodiment of the present application, since the first neural network of the PINNs includes multiple residual network channels and the coefficients corresponding to each residual network channel are different, in this way, during the model training stage, the same training sample can be multiplied by different coefficients, so that one data can be expanded into multiple data, and different frequency signals can be captured through multiple residual network channels, thereby improving the accuracy of model training.
[0151] Optionally, the active region is a region in the simulation domain centered on the point source corresponding to the excitation source and with a first length as the radius. The first length is related to the first parameter in the continuous probability density function, and the continuous probability density function approaches the Dirac function. The function of the point source is the product of the continuous probability density function and the signal of the excitation source. The passive region is the region in the simulation domain except for the active region and the boundary.
[0152] Optionally, the active output data is the sum of the output data of each residual network channel when a training sample in the multiple training samples contains the sampling point data of the active region. The passive output data is the sum of the output data of each residual network channel when a training sample in the multiple training samples contains the sampling point data of the passive region. The boundary output data is the sum of the output data of each residual network channel when a training sample in the multiple training samples contains the boundary data. The initial output data is the sum of the output data of each residual network channel when a training sample in the multiple training samples contains the initial data.
[0153] Optionally, each residual network channel includes a sine periodic activation function. The sine periodic activation function is used to convert the data in each residual network channel into electric field parameters and magnetic field parameters as the output data of each residual network channel.
[0154] Optionally, the coefficients corresponding to each residual network channel increase exponentially.
[0155] Optionally, the third processing unit 404 is configured to use one output data in the output data set as the known quantity of the partial differential equation each time, operate on the partial differential equation to obtain the loss function corresponding to one output data, and accumulate the loss functions corresponding to each output data in the output data set according to a preset relationship to obtain the total loss function.
[0156] The preset relationship includes learnable parameters and hyperparameters. The learnable parameters corresponding to different loss functions related to the total loss function are different. The learnable parameters will be updated as the parameters in the first neural network are updated. The hyperparameters are used to assist in weighting the loss functions corresponding to the learnable parameters.
[0157] Optionally, the fourth processing unit 405 is further configured to update the hidden vector of the simulation domain and the learnable parameters in the preset relationship.
[0158] Optionally, the preset relationship includes learnable parameters, and the learnable parameters corresponding to different loss functions are different.
[0159] The model training device 40 based on the physics-informed neural network described above can be understood by referring to the corresponding description in the foregoing method embodiment part, and will not be repeated here.
[0160] Such as Figure 12As shown in the figure, an embodiment of the incremental learning device 50 provided by the embodiments of the present application includes:
[0161] An acquisition unit 501, configured to acquire a plurality of sampling point data from the simulation domain of the antenna to be optimized. The plurality of sampling point data includes sampling point data of the active region, sampling point data of the passive region, data of the boundary of the simulation domain, and initial data of the simulation domain. The simulation domain includes an active region and a passive region. The acquisition unit 501 may execute step 301 in the above method embodiment.
[0162] A first processing unit 502, configured to input a plurality of sample data into the target physics-informed neural network. Each sample data includes a sampling point data and a first hidden vector of the simulation domain. The target physics-informed neural network is a target physics-informed neural network obtained by a model training method based on the physics-informed neural network. The first processing unit 502 may execute step 302 in the above method embodiment.
[0163] A second processing unit 503, configured to obtain output data corresponding to each sample data through the target physics-informed neural network. The second processing unit 503 may execute step 303 in the above method embodiment.
[0164] A third processing unit 504, configured to keep the parameters in the target physics-informed neural network unchanged, and adjust the first hidden vector of the simulation domain according to the output data to obtain a second hidden vector. The third processing unit 504 may execute step 304 in the above method embodiment.
[0165] Take the second hidden vector as the first hidden vector, and iteratively execute the above adjustment of the first hidden vector through different sample data until the output data meets the preset requirements of the antenna to be optimized, so as to obtain a second hidden vector that matches the simulation domain.
[0166] In the embodiments of the present application, during the incremental learning process, the parameters in the target physics-informed neural network are frozen, and the hidden vector of the simulation domain of the antenna to be optimized is repeatedly adjusted through the output data of the target physics-informed neural network until a hidden vector that matches the simulation domain is obtained. This method can quickly learn the hidden vector and improve the acquisition speed of the hidden vector in the new electromagnetic simulation scenario.
[0167] The embodiments of the present application provide an electromagnetic simulation device. The electromagnetic simulation device is installed with the above target physics-informed neural network model. The electromagnetic simulation device can simulate the antenna through the target physics-informed neural network model to obtain the electromagnetic field distribution of the simulation domain of the antenna.
[0168] Figure 13As shown in the figure, it is a schematic diagram of a possible logical structure of a computer device 60 provided by an embodiment of the present application. The computer device 60 can be a model training device based on a physical information neural network, an incremental learning device, or an electromagnetic simulation device. The computer device 60 includes: a processor 601, a communication interface 602, a memory 603, and a bus 604. The processor 601, the communication interface 602, and the memory 603 are interconnected through the bus 604. In the embodiment of the present application, the processor 601 is used to control and manage the actions of the computer device 60. For example, the processor 601 is used to execute Figures 1 to 9 the process in the method embodiment, and the communication interface 602 is used to support the computer device 60 to communicate. The memory 603 is used to store the program code and data of the computer device 60.
[0169] Among them, the processor 601 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure of the present application. The processor 601 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 604 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 13 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0170] In another embodiment of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores computer-executable instructions. When the processor of the device executes the computer-executable instructions, the device executes the above-mentioned model training method based on a physical information neural network, the incremental learning method, or executes the above-mentioned electromagnetic simulation method.
[0171] In another embodiment of the present application, a computer program product is further provided. The computer program product includes computer-executable instructions, and the computer-executable instructions are stored in a computer-readable storage medium; when the processor of the device executes the computer-executable instructions, the device executes the above-mentioned model training method based on a physical information neural network, the incremental learning method, or executes the above-mentioned electromagnetic simulation method.
[0172] In another embodiment of the present application, a chip system is further provided. The chip system includes a processor, which is used to implement the above-mentioned model training method based on physical information neural network, the incremental learning method, or execute the above-mentioned electromagnetic simulation method. In a possible design, the chip system may further include a memory, which is used to store the necessary program instructions and data for the inter-process communication device. The chip system may be composed of chips or may include chips and other discrete devices.
[0173] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0174] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0175] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces. The indirect coupling or communication connection of devices or units may be in an electrical, mechanical, or other form.
[0176] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0177] In addition, in each embodiment of the embodiments of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0178] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0179] The above is only the specific implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the embodiments of the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be subject to the protection scope of the claims.
Claims
1. A model training method based on the Physics-Informed Neural Networks (PINNs), characterized in that, the Physics-Informed Neural Networks include a first neural network and a partial differential equation, the first neural network includes at least two residual network channels, and the method includes: obtaining a plurality of sampling point data from the simulation domain of the antenna, the plurality of sampling point data including sampling point data of the active region, sampling point data of the passive region, data of the boundary of the simulation domain, and initial data of the simulation domain, and the simulation domain includes the active region and the passive region; inputting the product of each training sample in a plurality of training samples and the coefficient corresponding to each residual network channel into each residual network channel of the first neural network, each training sample includes a sampling point data and a latent vector corresponding to the simulation domain, and the coefficients corresponding to each residual network channel are different; processing the data input into each residual network channel through the first neural network to obtain an output data set, wherein the output data set includes active output data, passive output data, boundary output data, and initial output data; processing the output data set through the partial differential equation to obtain a total loss function, and the total loss function is related to an active loss function, a passive loss function, a boundary loss function, and an initial loss function; updating the parameters in the first neural network according to the total loss function to obtain a second neural network; using the second neural network as the first neural network, and iteratively executing the above training process until the second neural network reaches a convergence condition to obtain a target Physics-Informed Neural Networks model for electromagnetic simulation of the antenna.
2. The model training method according to claim 1, characterized in that, the active region is a region in the simulation domain centered on the point source corresponding to the excitation source and with a first length as the radius, and the first length is related to a first parameter in the continuous probability density function, and the continuous probability density function approaches the Dirac function; the passive region is a region in the simulation domain except the active region and the boundary of the simulation domain.
3. The model training method according to claim 1 or 2, characterized in that, the active output data is the sum of the output data of each residual network channel when a training sample in the plurality of training samples includes the sampling point data of the active region; the passive output data is the sum of the output data of each residual network channel when a training sample in the plurality of training samples includes the sampling point data of the passive region; the boundary output data is the sum of the output data of each residual network channel when a training sample in the plurality of training samples includes the data of the boundary of the simulation domain; the initial output data is the sum of the output data of each residual network channel when a training sample in the plurality of training samples includes the initial data.
4. The model training method according to any one of claims 1-3, characterized in that, each residual network channel includes a sine periodic activation function; The sine periodic activation function is used to convert the data in each residual network channel into electric field parameters and magnetic field parameters as the output data of each residual network channel.
5. The model training method according to any one of claims 1-4, wherein, the coefficients corresponding to each residual network channel increase exponentially.
6. The model training method according to any one of claims 1-5, wherein, processing the output data set through the partial differential equation to obtain a total loss function, including: each time taking an output data in the output data set as a known quantity of the partial differential equation, and operating on the partial differential equation to obtain a loss function corresponding to the one output data; accumulating the loss functions corresponding to each output data in the output data set according to a preset relationship to obtain the total loss function.
7. The model training method according to claim 6, wherein, the preset relationship includes learnable parameters and hyperparameters. The learnable parameters corresponding to different loss functions related to the total loss function are different. The learnable parameters will be updated as the parameters in the first neural network are updated. The hyperparameters are used to assist the learnable parameters in weighting the corresponding loss functions.
8. The model training method according to claim 7, wherein, when updating the parameters in the first neural network according to the total loss function, the method further includes: updating the hidden vector of the simulation domain and the learnable parameters in the preset relationship.
9. An incremental learning method, wherein, including: obtaining a plurality of sampling point data from the simulation domain of the antenna to be optimized. The plurality of sampling point data includes sampling point data of the active region, sampling point data of the passive region, data of the boundary of the simulation domain, and initial data of the simulation domain. The simulation domain includes the active region and the passive region; inputting a plurality of sample data into the target physics-informed neural network. Each sample data includes a sampling point data and a first hidden vector of the simulation domain. The target physics-informed neural network is the target physics-informed neural network obtained by the model training method according to any one of claims 1-8 above; obtaining output data corresponding to each sample data through the target physics-informed neural network; keeping the parameters in the target physics-informed neural network unchanged, and adjusting the first hidden vector of the simulation domain according to the output data to obtain a second hidden vector; taking the second hidden vector as the first hidden vector, and iteratively performing the above adjustment of the first hidden vector through different sample data until the output data meets the preset requirements of the antenna to be optimized, so as to obtain a second hidden vector matching the simulation domain.
10. A model training device based on a physics-informed neural network, wherein, the physics-informed neural network includes a first neural network and a partial differential equation. The first neural network includes at least two residual network channels. The model training device includes: An acquisition unit, configured to acquire multiple sampling point data from the simulation domain of the antenna, where the multiple sampling point data includes sampling point data of the active region, sampling point data of the passive region, data of the boundary of the simulation domain, and initial data of the simulation domain, and the simulation domain includes the active region and the passive region; A first processing unit, configured to input the product of each training sample in the multiple training samples and the coefficient corresponding to each residual network channel into each residual network channel of the first neural network, where each training sample includes a sampling point data and a hidden vector corresponding to the simulation domain, and the coefficients corresponding to each residual network channel are different; A second processing unit, configured to process the data input into each residual network channel through the first neural network to obtain an output data set, where the output data set includes active output data, passive output data, boundary output data, and initial output data; A third processing unit, configured to process the output data set through the partial differential equation to obtain a total loss function, where the total loss function is related to an active loss function, a passive loss function, a boundary loss function, and an initial loss function; A fourth processing unit, configured to update the parameters in the first neural network according to the total loss function to obtain a second neural network; Use the second neural network as the first neural network, and iteratively execute the above training process until the second neural network reaches a convergence condition to obtain a target physical information neural network model for electromagnetic simulation of the antenna.
11. The model training apparatus according to claim 10, wherein, the third processing unit is configured to use one output data in the output data set as a known quantity of the partial differential equation each time, and operate on the partial differential equation to obtain a loss function corresponding to the one output data; Accumulate the loss functions corresponding to each output data in the output data set according to a preset relationship to obtain the total loss function.
12. The model training apparatus according to claim 10, wherein, the fourth processing unit is further configured to update the hidden vector of the simulation domain and the learnable parameters in the preset relationship.
13. An incremental learning apparatus, wherein, comprising: An acquisition unit, configured to acquire multiple sampling point data from the simulation domain of the antenna to be optimized, where the multiple sampling point data includes sampling point data of the active region, sampling point data of the passive region, data of the boundary of the simulation domain, and initial data of the simulation domain, and the simulation domain includes the active region and the passive region; A first processing unit, configured to input multiple sample data into a target physical information neural network, where each sample data includes a sampling point data and a first hidden vector of the simulation domain, and the target physical information neural network is the target physical information neural network obtained by the model training method according to any one of claims 1-8; A second processing unit, configured to obtain output data corresponding to each sample data through the target physical information neural network; A third processing unit, configured to keep the parameters in the physical information neural network unchanged, and adjust a first hidden vector in the simulation domain according to the output data to obtain a second hidden vector; Use the second hidden vector as the first hidden vector, and iteratively execute the above adjustment of the first hidden vector through different sample data until the output data meets the preset requirements of the antenna to be optimized, so as to obtain a second hidden vector matching the simulation domain.
14. A computing device, characterized in that, it includes one or more processors and a computer-readable storage medium storing a computer program; When the computer program is executed by the one or more processors, the method according to any one of claims 1-8 or the method according to claim 9 is implemented.
15. A chip system, characterized in that, it includes one or more processors, and the one or more processors are called to execute the method according to any one of claims 1-8 or the method according to claim 9.
16. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by one or more processors, the method according to any one of claims 1-8 or the method according to claim 9 is implemented.
17. A computer program product, characterized in that, it includes a computer program, and when the computer program is executed by one or more processors, it is used to implement the method according to any one of claims 1-8 or the method according to claim 9.
Citation Information
Patent Citations
Motion parameter prediction method and device of fluid mechanics and storage medium
CN112784496A
Plasma equation numerical calculation method based on double neural network framework
CN113297534A