Data processing method and device, equipment and storage medium

By using the moment transfer analysis algorithm to calculate expectations and variance layer by layer in physical perception neural networks, replacing Monte Carlo sampling, the problem of high computing costs in the existing technology is solved, real-time and reliable uncertainty evaluation is achieved, and data processing efficiency is improved.

CN120354903AActive Publication Date: 2025-07-22ZHUHAI SHUZHOU TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510847204.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

When solving complex problems containing physical constraints in the prior art, the Monte Carlo sampling method leads to high calculation costs and long calculation time, making it difficult to meet the real-time and efficiency requirements.

Method used

The moment transfer analytical algorithm is used to calculate the expectations and variance of neuron output layer by layer in physical perception neural network, replacing Monte Carlo simulation, and realize the uncertainty quantification of single forward propagation.

Benefits of technology

Without sacrificing quantitative accuracy, real-time and reliable uncertainty evaluation of model prediction is achieved, data processing efficiency is improved, and real-time requirements are applicable to high-dimensional and nonlinear physical systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354903A_ABST
    Figure CN120354903A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, a data processing device, data processing equipment and a storage medium, and aims to solve the efficiency bottleneck of uncertainty evaluation when a complex problem containing physical constraints is solved in the prior art by providing an uncertainty quantification method for a physical perception neural network. According to the method, in a pre-trained PINNs model, expectation and variance output by neurons are calculated layer by layer in single forward propagation through a moment transfer analysis algorithm, Monte Carlo simulation depending on mass sampling is replaced by the method, and the calculation efficiency is improved. The pain point that the method is difficult to apply due to too high calculation cost in high-dimensional and nonlinear physical system simulation is fundamentally solved. Finally, on the premise of not sacrificing quantization precision, real-time and reliable uncertainty evaluation of model prediction is achieved, and key technical support is provided for application of PINNs in the fields of digital twinning, medical diagnosis, aerospace and the like which have strict requirements for safety and real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing of physical perception neural networks, and in particular, to a data processing method, apparatus, device, and storage medium. Background Art

[0002] In recent years, deep learning technology has achieved numerous unprecedented results with its powerful feature extraction and fitting capabilities. For example, in the field of image processing, deep learning is developing in the direction from low-level pixel processing to high-level semantic understanding, becoming the core driving force in the field of computer vision. Against this background, Scientific Machine Learning, as an emerging interdisciplinary discipline, has emerged. This method uses the advantages of deep neural networks in high-dimensional function approximation to provide new solutions for partial differential equation problems that are difficult to handle by traditional numerical methods. In scientific computing, the core computational cost often focuses on the solution process of the forward problem (forward model). Therefore, using a neural network to construct a surrogate of the forward model naturally becomes a choice to reduce the computational cost. In the existing technology of scientific machine learning, such as MC-dropout (a method for neural network uncertainty quantification), it relies on the Monte Carlo sampling strategy and needs to perform dozens to hundreds of forward propagations to statistically output the distribution. In high-dimensional scientific computing scenarios, the time-consuming of a single forward propagation is significant, which makes the total computational time of the Monte Carlo method increase sharply and is difficult to meet the real-time requirements. In addition, N samplings are required for one forward propagation, and a large number of repeated samplings will also cause the data processing volume to increase linearly, further reducing the data processing efficiency, making the feasibility of the algorithm in practical applications severely limited. Summary of the Invention

[0003] The purpose of the embodiments of the present invention is to provide an efficient physical neural network uncertainty quantification algorithm, including a data processing method, apparatus, device, and storage medium, which can effectively improve the real-time performance and efficiency of data processing.

[0004] To achieve the above purpose, the embodiments of the present invention provide a data processing method, including: Obtain data to be predicted; Input the data to be predicted into a pre-trained physical perception neural network model, where the training process of the physical perception neural network model incorporates physical laws as constraint conditions, so that the physical perception neural network model performs moment transfer calculations and outputs a predicted value and a target variance; Construct a confidence interval according to the predicted value and the target variance; Determine the uncertainty quantification result of the predicted value according to the confidence interval; where the uncertainty quantification result is used to characterize the reliability of the predicted value.

[0005] As an improvement to the above solution, the data to be predicted is data characterizing the state of a physical system, and the data is spatio-temporal coordinates for solving differential and partial differential equations, sparse monitoring signals, noisy monitoring signals, snapshot data of boundaries and variables, industrial simulation cloud maps, medical image data, meteorological observation data, or industrial inspection images.

[0006] As an improvement to the above solution, the physics-informed neural network model includes at least one moment transfer architecture; wherein, each layer in the moment transfer architecture transfers the expected value and variance value backward, and each moment transfer architecture includes a linear layer, a normalization layer, an activation function layer, and a dropout layer connected in sequence.

[0007] As an improvement to the above solution, the linear layer uses a weight matrix and a bias to perform a linear transformation on the input first expected value and first variance, and outputs a second expected value and a second variance; wherein, the first expected value is the distribution mean of the data to be predicted, and the first variance is the distribution variance of the data to be predicted; The normalization layer uses normalization parameters to perform a normalization calculation on the input second expected value and second variance, and outputs a third expected value and a third variance; The activation function layer uses an activation function to perform a functional operation on the input third expected value and third variance, and outputs a fourth expected value and a fourth variance; The dropout layer randomly discards according to the fourth expected value and the fourth variance through a preset dropout probability, and outputs a fifth expected value and a fifth variance.

[0008] As an improvement to the above solution, the physics-informed neural network model is provided with a target linear layer at the end of the last moment transfer architecture, and the physics-informed neural network model uses the second expected value output by the target linear layer as the predicted value, and uses the second variance output by the target linear layer as the target variance.

[0009] As an improvement to the above solution, the activation function is a smooth function that can perform high-order derivatives to meet the calculation requirements of physical law constraints, and the smooth function is a Sigmoid function, a Tanh function, or a Swish function.

[0010] As an improvement to the above solution, constructing a confidence interval according to the predicted value and the target variance includes: Calculating a target standard deviation according to the target variance; Calculating the product of the target standard deviation and a preset confidence coefficient to obtain a confidence deviation value; Calculating the sum of the predicted value and the confidence deviation value to obtain an upper limit value, and calculating the difference between the predicted value and the confidence deviation value to obtain a lower limit value; A confidence interval is constructed with the upper limit value and the lower limit value.

[0011] To achieve the above object, an embodiment of the present invention further provides a data processing device, including: A data acquisition module, used to acquire data to be predicted; A model processing module, used for inputting the data to be predicted into a pre-trained physical perception neural network model, wherein the training process of the physical perception neural network model incorporates physical laws as constraints so that the physical perception neural network model performs moment transfer calculations and outputs predicted values and target variances; A confidence interval construction module, used to construct a confidence interval according to the predicted value and the target variance; A quantification result output module is used to determine the uncertainty quantification result of the predicted value according to the confidence interval; wherein the uncertainty quantification result is used to characterize the reliability of the predicted value.

[0012] To achieve the above objectives, an embodiment of the present invention further provides a data processing device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the data processing method as described in any of the above embodiments is implemented.

[0013] To achieve the above objectives, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the data processing method as described in any of the above embodiments.

[0014] Compared with the prior art, the data processing method, device, equipment and storage medium disclosed in the embodiments of the present invention provide an uncertainty quantification method for physics-informed neural networks (PINNs), aiming to solve the efficiency bottleneck of uncertainty assessment in solving complex problems with physical constraints in the prior art. In the pre-trained PINNs model, the present invention uses a moment transfer analytical algorithm to calculate the expectation and variance of neuron output layer by layer in a single forward propagation. This method replaces the Monte Carlo simulation that relies on massive sampling, and fundamentally solves the pain point that it is difficult to apply in high-dimensional, nonlinear physical system simulation due to high computational costs. Finally, the present invention realizes real-time and reliable uncertainty assessment of model predictions without sacrificing quantization accuracy, providing key technical support for the application of PINNs in fields such as digital twins, medical diagnosis, aerospace, etc. that have strict requirements on safety and real-time performance.

[0015] It should be noted that traditional MC-dropout needs to perform forward propagation statistical distribution dozens to hundreds of times, while the present invention analyzes the expectation and variance layer by layer in a single forward propagation through a moment transfer calculation module. The data processing process only needs to be completed once, avoiding the problem of linear growth of the sampling times and data volume in the traditional method, and is especially suitable for real-time data flow scenarios. The present invention replaces the statistical method of Monte Carlo sampling with an analytical method of moment transfer, fundamentally solving the problems of high computational cost and large data processing volume caused by multiple forward propagations in the prior art, and is especially suitable for the real-time uncertainty quantification requirements in high-dimensional scientific computing scenarios. In addition, the present invention uses moment transfer calculation to replace repeated sampling, improving the data processing efficiency on the premise of ensuring the uncertainty quantification effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a flowchart of a data processing method provided by an embodiment of the present invention; Figure 2 is a schematic diagram of the framework of the PINNs model provided by an embodiment of the present invention; Figure 3 is a schematic diagram of the uncertainty quantification result of the physics-informed neural network by the MC method provided by the prior art; Figure 4 is a schematic diagram of the uncertainty quantification result of the physics-informed neural network by the MP method provided by an embodiment of the present invention; Figure 5 is a block diagram of the structure of a data processing device provided by an embodiment of the present invention; Figure 6 is a block diagram of the structure of a data processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] It should be noted that the "uncertainty quantification result" is a core metric in machine learning and data analysis for measuring the reliability of model predictions, and is particularly crucial in scenarios where prediction risks need to be evaluated (such as medical diagnosis, autonomous driving, scientific simulations). The essence of the uncertainty quantification result is a quantitative metric or distribution used to describe the "degree of uncertainty" of the model's prediction results for input data. In plain terms, it represents how confident the model is in this prediction and reflects the potential deviation range between the predicted value and the true value. In the embodiments of the present invention, the uncertainty quantification result is used to avoid blindly trusting unreliable predictions. For example, in the task of medical image (a type of image data) segmentation, for pixel regions where the confidence interval exceeds a preset threshold (such as the blurred boundary of a tumor), the system automatically triggers an artificial review process to avoid missed or misdiagnosed cases caused by model misjudgment. By deeply integrating the uncertainty quantification result into the model design and application process, the present invention promotes the transformation of machine learning from black-box prediction to transparent decision-making, thereby improving the data processing efficiency while ensuring the effect of uncertainty quantification.

[0019] See Figure 1 , Figure 1 is a flowchart of a data processing method provided by an embodiment of the present invention. The method includes: S11. Obtain the data to be predicted; S12. Input the data to be predicted into a pre-trained physics-informed neural network model, where the training process of the physics-informed neural network model incorporates physical laws as constraint conditions, enabling the physics-informed neural network model to perform moment transfer calculations and output a predicted value and a target variance; S13. Construct a confidence interval based on the predicted value and the target variance; S14. Determine the uncertainty quantification result of the predicted value according to the confidence interval; where the uncertainty quantification result is used to characterize the reliability degree of the predicted value.

[0020] Exemplarily, the data to be predicted is data characterizing the state of a physical system, and the data includes but is not limited to spatio-temporal coordinates for solving differential and partial differential equations, sparse monitoring signals, noisy monitoring signals, snapshot data of boundaries and variables, industrial simulation cloud maps, medical image data, meteorological observation data, or industrial inspection images. First, obtain the data to be predicted, such as medical image data, as the input basis for subsequent processing (this process can preprocess the medical image data, such as cropping, denoising, etc., and then convert the medical image data into a two-dimensional matrix); second, input the data into the physics-informed neural network model, and the physics-informed neural network model calculates the expectations (predicted values) and variances of each layer through moment propagation, and outputs the final predicted value and the target variance; then, use the predicted value and the target variance to construct a confidence interval representing the prediction uncertainty; finally, generate an uncertainty quantification result according to the confidence interval to intuitively characterize the reliability of the predicted value and assist in decision-making.

[0021] In the embodiment of the present invention, since traditional MC-dropout needs to perform forward propagation statistical distributions dozens to hundreds of times, while the present invention parses expectations and variances layer by layer in a single forward propagation through a moment propagation calculation module, and the data processing process only needs to be completed once, avoiding the problem of linear growth of the sampling times and the data volume in the traditional method, and is particularly suitable for real-time data stream scenarios. The present invention replaces the statistical method of Monte Carlo sampling with an analytical method of moment propagation, fundamentally solving the problems of high computational cost and large data processing volume caused by multiple forward propagations in the prior art, and is particularly suitable for the real-time uncertainty quantification requirements in high-dimensional scientific computing scenarios. In addition, the present invention uses moment propagation calculation to replace repeated sampling, improving the data processing efficiency on the premise of ensuring the uncertainty quantification effect.

[0022] Specifically, in step S14, constructing a confidence interval according to the predicted value and the target variance includes: calculating a target standard deviation according to the target variance; calculating the product of the target standard deviation and a preset confidence coefficient to obtain a confidence deviation value; calculating the sum of the predicted value and the confidence deviation value to obtain an upper limit value, and calculating the difference between the predicted value and the confidence deviation value to obtain a lower limit value; constructing a confidence interval with the upper limit value and the lower limit value.

[0023] Exemplarily, the width of the confidence interval directly reflects the uncertainty of the predicted value. The wider the interval, the higher the uncertainty and the less reliable the predicted value. Conversely, the more reliable it is. Assume that the data to be predicted is medical image data in the format of a two-dimensional matrix with dimensions satisfying: 512 * 512 pixels, and the single-channel grayscale value range is 0 to 255. After data processing by the physics-aware neural network model, assume that the predicted value output by the model is 0.6 (i.e., it means that there is a 60% probability that this pixel point is a lesion area), and the output target variance is 0.04, and the confidence coefficient is 2. Then, take the square root of the target variance to obtain the target standard deviation of 0.2. Further calculate to get the confidence deviation value of 0.4. Then, the corresponding upper limit value is 1.0, the lower limit value is 0.2, and the confidence interval is [0.2, 1.0]. This confidence interval indicates that: the predicted true value is most likely 0.6, but the confidence interval is relatively wide, indicating that the model has a high uncertainty in classifying this pixel.

[0024] In the embodiment of the present invention, the confidence interval provides a quantitative basis for decision-making. Through a strict statistical calculation process, the target variance is converted into a confidence interval, and the fluctuation range of the predicted value is intuitively presented with clear upper and lower limit values, accurately quantifying the prediction uncertainty and avoiding ambiguous expressions. In addition, based on the calculation steps of a fixed formula, there is no need for complex iteration or a large number of samplings, and the calculation complexity is low. In scenarios with high real-time requirements, the confidence interval can be quickly output, saving time for system decision-making and ensuring the efficient operation of the system.

[0025] Specifically, the physics-aware neural network model includes at least one moment transfer architecture; wherein, each layer in the moment transfer architecture transfers the expected value and variance value backward, and each moment transfer architecture includes a linear layer, a normalization layer, an activation function layer, and a dropout layer connected in sequence.

[0026] Exemplarily, see Figure 2 , Figure 2It is a schematic diagram of the framework of the Physics-Informed Neural Networks (PINNs) model provided by an embodiment of the present invention. The Physics-Informed Neural Networks model consists of an input layer, n moment transfer architectures, and a target linear layer. By providing a target linear layer at the end of the last moment transfer architecture and using this target linear layer as the output layer, its structure and purpose are the same as those of the linear layer in the moment transfer architecture. Since the training process of the Physics-Informed Neural Networks model integrates physical laws as constraint conditions, the target linear layer outputs the Physics-Informed Neural Networks model required by the present invention. Further, it may also include initial and boundary conditions. The initial and boundary conditions are core concepts in the problem of determining the solution of partial differential equations, used to define the unique solution of the equation, and can be specifically decomposed into initial conditions and boundary conditions. By forcing the network input / output to conform to the initial and boundary conditions, the purpose is to enable the neural network to integrate physical laws, avoid the deviation of pure data-driven methods, and improve the physical rationality of predictions. It should be noted that the initial and boundary conditions are preset for the data used to solve differential and partial differential equations during model training. The setting process can refer to the prior art, and the present invention does not make specific limitations on this.

[0027] The Physics-Informed Neural Networks model uses the second expectation output by the target linear layer as the predicted value and the second variance output by the target linear layer as the target variance, thereby outputting the predicted value and the target variance. For each layer in the moment transfer architecture, the following detailed description is made respectively: 1) The linear layer uses a weight matrix and a bias to perform a linear transformation on the input first expectation and first variance, and outputs a second expectation and a second variance; where the first expectation is the distribution mean of the data to be predicted, and the first variance is the distribution variance of the data to be predicted.

[0028] Exemplarily, the linear layer, as a basic transformation layer, performs a linear combination on the input data and outputs a new feature representation, providing a processing basis for subsequent layers. After the data to be predicted undergoes data processing, its corresponding first expectation and first variance are obtained. For example, when the data to be predicted is medical image data, after the medical image data undergoes gray normalization and region of interest extraction, the average gray value is calculated to obtain the first expectation, and the dispersion degree of the gray values within the neighborhood is calculated to obtain the first variance. The first variance is used to characterize the dispersion degree of the local texture of the image.

[0029] Exemplarily, assume that the first expectation of the data to be predicted is and the first variance is , then the calculation process of the second expectation and the second variance satisfies the following formulas: (1); (2); where is the weight matrix, is the weight matrix of the transposed matrix; is the bias.

[0030] Furthermore, the first variance can be a covariance matrix. In the moment transfer calculation of a neural network, expanding the first variance into a covariance matrix can capture the dependencies between variables and make the uncertainty propagation more in line with physical laws. When calculating the statistics of the output of a certain layer of a neural network, the output of the layer can be regarded as multiple random variables (the output of each neuron). The covariance matrix of these variables reflects their correlation. And in Dropout, whether a neuron is discarded is random. Therefore, in actual calculation, if it is assumed that the neurons are independent of each other, at this time the covariance matrix degenerates into a diagonal matrix, and only the diagonal elements of the covariance matrix need to be calculated. These diagonal elements are the variances of the outputs of each neuron, thus simplifying the calculation process. Finally, after the covariance matrix undergoes a linear transformation by the weight matrix it reflects the covariant relationship between the output variables.

[0031] 2) The normalization layer uses normalization parameters to perform normalization calculations on the input second expectation and the second variance, and outputs a third expectation and a third variance.

[0032] Exemplarily, the normalization layer normalizes the output of the linear layer to standardize the data distribution, thereby accelerating data processing and enhancing stability. The third expectation and the third variance The calculation process satisfies the following formula: (3); (4); where, are the first network parameter and the second network parameter respectively, which play a role in scaling and offsetting; are the expectation coefficient and the variance coefficient respectively, which play a role in normalization; together constitute the normalization parameters; is a fixed constant used to prevent the denominator in the formula from being 0.

[0033] 3) The activation function layer uses an activation function to perform a functional operation on the input third expectation and the third variance, and outputs a fourth expectation and a fourth variance.

[0034] Exemplarily, the activation function is a smooth function that can perform high-order derivatives to meet the calculation requirements of the physical law constraints. The smooth function includes but is not limited to the Sigmoid function, Tanh function, or Swish function. The activation function layer introduces non-linearity to the output of the normalization layer, enabling the network to learn complex patterns through a series of expectation and variance calculations. The ReLU function is the most common activation function in the field of deep learning currently due to its characteristics such as effectively avoiding the vanishing gradient. However, when applied to physics-informed neural networks (PINNs), there are significant limitations. The reason is that the physics-informed neural network embeds the partial differential equation model into the loss function, and the partial differential equation model often contains high-order derivative information. At this time, using the ReLU as the activation function will cause the high-order derivative of the network with respect to the input variable to degenerate to 0. Therefore, the ReLU function is not suitable for network architectures trained based on physical information such as PINNs. Taking the Sigmoid function as an example, due to its smoothness and high-order derivative characteristics, the Sigmoid function is more suitable for physics-informed neural networks. Therefore, the present invention generalizes the moment transfer of the Sigmoid activation function to physics-informed neural networks.

[0035] Exemplarily, when the output data of the normalization layer passes through the Sigmoid function, the following formula is satisfied: (5); Where, assuming the input , represents the input follows a normal distribution with a mean of and a variance of . Here , = , that is, the input of the activation function layer is data that conforms to the third expectation and the fourth expectation of the output of the normalization layer. Formula (5) indicates that the Sigmoid function maps the input to the interval (0, 1).

[0036] Moments are statistical quantities that describe the distribution of random variables. For example, the first moment is the expectation, the second central moment is the variance, and higher-order moments describe the shape of the distribution (skewness, kurtosis, etc.). In the moment propagation of neural networks, each layer processes the moments of the input distribution. Layers such as linear layers, normalization layers, and activation function layers will update these moments. Therefore, in the moment propagation architecture, the input of each layer is the moments (such as expectation and variance) of the output distribution of the previous layer, and the operations of the layer (linear transformation, normalization, activation function) will change these moments. So, it is necessary to derive the transformed moment information. In neural networks, the input of the Sigmoid function is usually a random variable (due to the introduction of uncertainty by weights / Dropout, etc.). When calculating the moments of the activated variable, direct integration is very complex. However, introducing the logistic distribution can transform the problem into the derivation of the moments of probability events. Therefore, by using the known moments (expectation, variance) of the logistic distribution, the complex integration of the Sigmoid function can be avoided, greatly reducing the mathematical difficulty of moment derivation.

[0037] Before deriving the moment information in the present invention, a random variable is introduced , and according to the properties of the logistic ( ) distribution, it can be known that the conditional expectation , the conditional variance , and the conditional cumulative distribution function respectively satisfy the following formulas: (6); (7); (8); Among them, the conditional expectation represents the expectation of under the condition of .

[0038] Since there is a connection in form between the cumulative distribution function of the logistic distribution and the Sigmoid function, when is given, the conditional probability that ≥ 0 is the value of the Sigmoid function, that is, it satisfies: (9); Further, it can be obtained that: (10); Among them, is the expectation; is the density function. In the context of the moment propagation architecture, the density function Refers to the Probability Density Function (PDF) of a random variable, which is used to describe the probability distribution pattern of the output variables of each layer in a neural network. In the moment transfer architecture, the input and output of each layer are regarded as random variables, and their distributions are characterized by density functions characterized.

[0039] Using the moment matching technique, can be used to approximate , so , since the Gaussian distribution is its own conjugate prior, the marginal distribution can be inferred and satisfies the following formula: (11); Applying moment matching again gives approximately follows , so that the expectation of can be transformed into the form of a sigmoid function with respect to the mean and variance for subsequent calculations, that is, it satisfies the following formula: (12); where represents the sigmoid function with input .

[0040] Next, consider the transformation of the variance and use the following variance change formula for subsequent derivation. The variance change formula satisfies: (13); (14); where represents the variance of the random variable , measuring the deviation degree of the value of from its expectation ; represents the derivative of the sigmoid function, that is, the first derivative of , reflecting the change rate of . Formula (13) is based on the definition of variance. By analyzing the characteristics of , its variance is transformed into the form of

[0041] Using the moment matching technique, the density function of a Gaussian distribution with an expectation of 0 and a variance of can be used to approximate the expectation of , satisfying the following formula: (15); where formula (15) uses the moment matching technique and uses a Gaussian distribution with an expectation of 0 and a variance of Gaussian distribution density function Approximate and calculate Expectation of , this step is to transform the complex integral operation into a tractable form, and finally substitute it into the variance formula (13) to complete the accurate solution of the Variance.

[0042] Let , it is easy to obtain the extreme point of the function as being , is obtained by taking the derivative of and setting the derivative to 0, which is the center of the Taylor expansion. Expanding using Taylor expansion gives the following formula: (16); wherein, represents the derivative after taking the second derivative of , that is, the quadratic term coefficient in the Taylor expansion; formula (16) means first expanding using Taylor expansion, and then performing approximate integral calculation, approximating the complex distribution with a simple distribution (Gaussian), and combining Taylor expansion to reduce the difficulty of dimensionality reduction integration.

[0043] Applying the moment matching technique again can transform into the form of the derivative of the Sigmoid function for subsequent variance calculation, that is: (17); wherein, formula (17) is to transform into a form directly related to the derivative of the Sigmoid function for substitution into the variance formula.

[0044] Substituting it into the above variance change formula (13), we can get: (18).

[0045] Furthermore, is the fourth expectation, is the fourth variance.

[0046] 4) The dropout layer randomly drops according to the fourth expectation and the fourth variance with a preset dropout probability, and outputs the fifth expectation and the fifth variance.

[0047] Exemplarily, the dropout layer randomly drops some neuron data during the data processing process, and adjusts the expectation and variance to prevent overfitting and enhance the generalization ability. Let the input neuron be a random variable (with an expectation of and a variance of ), with a discard probability of , the output is , where . In the Bernoulli distribution , is a discrete random variable that can only take two values, 0 or 1. has a probability of , representing a successful event (in the context of a dropout layer, it can be understood as "retaining the input neuron information"); has a probability of , representing a failed event (in the context of a dropout layer, it can be understood as "discarding the input neuron information"). In the dropout layer formula , is used to randomly control the retention or discard of the input neuron : when , is retained and scaled by ; when , is discarded, and the output is . Through this random discard mechanism, overfitting of the neural network can be effectively prevented.

[0048] Since is independent of , we can obtain , that is, the fifth expectation and the fifth variance satisfy the following formulas: (19); (20); where is the fourth expectation , is the fourth variance , and formula (19) means that although the input random variable has experienced random discard, after being scaled by , the expectation of the output is equal to the expectation of the input ; formula (20) is derived based on the characteristics of the Bernoulli distribution and the variance operation rules, combined with the influence of the discard operation on , reflecting the adjustment of the output variance by the discard probability .

[0049] It should be noted that based on the above analysis, the analytical transfer form of neuron moment information in common network architectures is obtained. On this basis, a moment transfer dropout network (characterized by MP-dropout) is constructed to replace the Monte Carlo simulation process of the MC-dropout method in the prior art during the prediction stage, thereby achieving prediction acceleration.

[0050] In the embodiment of the present invention, the output statistics (expectation, variance / covariance) of each layer are used as the input statistical characteristics of the next layer. The expectation and variance of the output of the linear layer are input into the normalization layer, and new expectation and variance are obtained through calculation and then input into the Sigmoid activation function layer for the derivation of expectation and variance under nonlinear transformation. Finally, the dropout layer performs a random dropout operation according to the input statistical characteristics and adjusts the output statistics. This layer-by-layer transfer of statistical characteristics ensures that within the moment transfer framework, the network can quantify and accumulate the uncertainty of each layer, and finally output the prediction mean and confidence interval, realizing the accurate characterization of the uncertainty of the entire network output.

[0051] Furthermore, during the training stage of the physics-informed neural network model, it is necessary to retain the randomness of the dropout layer and the network weights, that is, retain the dropout probability. The moment transfer algorithm can analytically calculate the expectation and variance of the output in a single forward propagation without relying on multiple Monte Carlo samplings, then calculate the neuron expectation and variance layer by layer, and finally output the predicted value and the target variance to construct a confidence interval. That is, the expectation of the final output of the model is used as the predicted value, and the standard deviation is obtained by taking the square root of the variance of the final output to construct the confidence interval.

[0052] Furthermore, the present invention provides an embodiment for predicting the temperature field of tumor hyperthermia in medical images (using a one-dimensional nonlinear Poisson equation): In tumor hyperthermia treatment planning, it is necessary to accurately predict the temperature distribution generated by microwave radiation in biological tissues. Traditional methods are based on the bioheat transfer equation (Pennes equation), but there are the following pain points: ① There are individual differences in tissue thermophysical parameters, and the measurement data are sparse and noisy; ② The dynamic change of blood perfusion rate during the treatment process leads to model uncertainty; ③ It is necessary to calculate the temperature field in real time and quantify the prediction reliability to ensure treatment safety. The clinical value of the MP-dropout provided by the present invention is to display the temperature distribution in the tumor core area (near x = 0) in real time, and observe whether it is necessary to initiate adaptive temperature measurement verification according to the uncertain area of thermophysical parameters (confidence bandwidth). Consider the following heat transfer model: (21); Wherein, represents the tissue temperature distribution (i.e., the data to be predicted, which is a physical quantity); is the tissue thermal conductivity, which can be obtained by preoperative puncture measurement; is actually expressed as is the second partial derivative of temperature with respect to the spatial coordinates, which is used in mathematics to characterize the rate of heat diffusion (the larger the second derivative, the faster the heat diffusion), and physically corresponds to the spatial rate of change of heat conduction; is the blood perfusion rate, which is a dynamically changing parameter; is the hyperbolic tangent function, which describes the non-linear feedback of temperature on blood perfusion. When the temperature increases, the response of blood perfusion is non-linear; is the microwave radiation heat source. The input data is the spatial coordinates , ∈ [0.7, 0.7] (corresponding to the tissue layer at a depth of 7 cm); the boundary condition indicates that the epidermal temperature (corresponding to the boundary at a depth of 7 cm) satisfies: ( 0.7), (0.7).

[0053] In the experiment of the present invention, 32 data of heat source observation points obtained in real time (measured by an implanted fiber optic temperature sensor), and for the sake of simplifying the calculation, only numerical simulation is carried out here, and it is not required to conform to the actual situation. By combining "physical equation constraints" and "observation data" through a physics-informed neural network (PINNs), and combining the Dropout method to quantify the prediction uncertainty, the steps are as follows: 1.1 Assume that the true temperature distribution is , and at this time, the true solution can be set as . In tumor thermotherapy, the temperature distribution is affected by microwave heat generation, heat conduction, and blood perfusion together, and will show the characteristics of multi-peak and non-linear fluctuations. can shorten the time period to π / 3, can simulate the rapid fluctuation of temperature in a local area (such as the temperature gradient between the tumor edge and normal tissue), and the cubic non-linearity can enhance the "asymmetry" and "steep change" of the waveform, simulating the non-linear response of heat diffusion in biological tissues (such as the formation of a high-temperature area in the tumor core). And set the remaining parameters: is unknown. In ∈ [0.7, 0.7], 32 data points are obtained through an intraoperative infrared temperature measurement probe, and Gaussian noise is added to simulate the noise characteristics of clinical measurement.

[0054] 1.2 Using these observation data, solve the tissue temperature distribution and quantify the uncertainty. Adopting the idea of PINNs, approximate with a neural network, regard as a function with spatial coordinates as the input and temperature value as the output, approximate this unknown function with a neural network, and construct according to formula (21) The proxy function, and construct a loss function in combination with the observed data to train the network. Since derivative calculations are involved, a relatively small dropout probability is set, such as = 0.01 to ensure training stability.

[0055] 1.3 Use 300 Monte Carlo simulations to quantify the uncertainty of the prediction results using MC-dropout, and use MP-dropout provided by the present invention to quantify the uncertainty of the prediction results, respectively obtaining Figures 3 - 4 Results, Figure 3 is a schematic diagram of the uncertainty quantification result of the physics-informed neural network by the MC method provided by the prior art, Figure 4 is a schematic diagram of the uncertainty quantification result of the physics-informed neural network by the MP method provided by the embodiment of the present invention. The black solid line in the figure is the true temperature distribution, representing the ideal prediction target; the black dashed line is the prediction mean (that is, the average result of 300 Dropout samplings, reflecting the "best prediction" of the model); the gray area represents the ±2σ (standard deviation) range, representing the uncertainty interval (that is, the confidence interval). The wider the width, the lower the credibility of the prediction.

[0056] Exemplarily, through Figure 3 and Figure 4 , it can be concluded that due to sparse and noisy data, there are various uncertainties in the model prediction, manifested as significant error differences in different intervals, and the intervals with large errors correspond to wider confidence bands. This shows that the present invention has good error calibration ability and can reasonably reflect the uncertainty of the model prediction. In addition, MP-dropout has similar results to MC-dropout in the network using the Sigmoid activation function. This is because both MC-dropout (Monte Carlo Dropout) and MP-dropout simulate model uncertainty through the random perturbation of Dropout, and their prediction results converge in the network. Both are based on the theory that "Dropout is equivalent to integrating multiple sub-models" (multiple samplings of Dropout during inference are equivalent to training multiple sub-models with similar structures and integrating them), so the statistical laws of uncertainty (trends of mean and variance) are the same. In addition, in the case of few samples, the uncertainty of the model mainly comes from the ambiguity of parameter estimation, rather than model structure differences. Therefore, the difference in the uncertainty quantification results of the two Dropout methods is reduced.

[0057] Furthermore, the time overhead of MC-dropout stems from the repeated calculations of multiple random samplings, while MP-dropout avoids this problem through an optimization strategy. Therefore, MC-dropout has a greater time cost because it obtains results through multiple simulations of random sampling. This is because in the few-shot setting, MC-dropout requires more samplings to make the uncertainty distribution "smooth", but the number of samplings is linearly related to time, resulting in a sharp increase in time cost. Especially when the number of sample points is relatively small, the edge of the result has a non-smooth property, that is, the uncertainty band in the boundary region of MC-dropout appears "sawtooth-shaped", indicating large sampling fluctuations and poor stability. On the other hand, even when the samples are few and the network is deep (e.g., using the Sigmoid activation function, the gradient is prone to vanishing), the confidence band of MP-dropout remains smooth and has stronger robustness, indicating that the present invention is still effective in more complex network structures and has good generality. Therefore, the MP-dropout proposed by the present invention has more advantages in complex networks and few-shot scenarios. It not only retains the uncertainty quantification ability of Dropout but also solves the data processing efficiency and stability problems of MC-dropout.

[0058] See Figure 5 , Figure 5 FIG. is a structural block diagram of a data processing device 100 provided by an embodiment of the present invention. The data processing device 100 includes: A data acquisition module 11, configured to acquire data to be predicted; A model processing module 12, configured to input the data to be predicted into a pre-trained physics-informed neural network model, wherein the training process of the physics-informed neural network model incorporates physical laws as constraint conditions, so that the physics-informed neural network model performs moment transfer calculations and outputs a predicted value and a target variance; A confidence interval construction module 13, configured to construct a confidence interval according to the predicted value and the target variance; A quantization result output module 14, configured to determine an uncertainty quantization result of the predicted value according to the confidence interval; wherein the uncertainty quantization result is used to characterize the reliability degree of the predicted value.

[0059] It should be noted that the specific working processes of the various modules in the data processing device 100 according to the embodiments of the present invention may refer to the working process of the data processing method described in the above embodiments, and will not be elaborated herein.

[0060] See Figure 6 , Figure 6It is a structural block diagram of a data processing device 200 provided by an embodiment of the present invention. The data processing device 200 includes a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the processor 21 executes the computer program, the steps in the above-mentioned various embodiments of the data processing method are implemented, such as steps S11 to S14.

[0061] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory 22 and executed by the processor 21 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the data processing device 200.

[0062] The data processing device 200 may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art can understand that the schematic diagram is only an example of the data processing device 200, and does not constitute a limitation on the data processing device 200. It may include more or fewer components than shown, or combine certain components, or different components. For example, the data processing device 200 may further include input / output devices, network access devices, buses, etc.

[0063] The processor 21 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor 21 is the control center of the data processing device 200, and connects all parts of the entire data processing device 200 through various interfaces and lines.

[0064] The memory 22 can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory 22 and invoking the data stored in the memory 22, the processor 21 realizes various functions of the data processing device 200. The memory 22 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 22 can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0065] Among them, if the modules / units integrated in the data processing device 200 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 21, the steps of the above-mentioned various method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0066] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A data processing method, characterized in that, Including: Obtain the data to be predicted; Input the data to be predicted into a pre-trained physics-informed neural network model, wherein the training process of the physics-informed neural network model incorporates physical laws as constraint conditions, so that the physics-informed neural network model performs moment transfer calculation and outputs a predicted value and a target variance; Construct a confidence interval based on the predicted value and the target variance; Determine the uncertainty quantification result of the predicted value according to the confidence interval; wherein, the uncertainty quantification result is used to characterize the reliability degree of the predicted value.

2. The data processing method according to claim 1, wherein The data to be predicted is data characterizing the state of a physical system, and the data is spatio-temporal coordinates for solving differential and partial differential equations, sparse monitoring signals, noisy monitoring signals, snapshot data of boundaries and variables, industrial simulation cloud maps, medical image data, meteorological observation data or industrial inspection images.

3. The data processing method according to claim 1, wherein The physics-informed neural network model includes at least one moment transfer architecture; wherein, each layer in the moment transfer architecture transfers the expected value and the variance value backward, and each moment transfer architecture includes a linear layer, a normalization layer, an activation function layer and a dropout layer connected in sequence.

4. The data processing method according to claim 3, characterized in that The linear layer uses a weight matrix and a bias to perform a linear transformation on the input first expected value and first variance, and outputs a second expected value and a second variance; wherein, the first expected value is the distribution mean of the data to be predicted, and the first variance is the distribution variance of the data to be predicted; The normalization layer uses normalization parameters to perform a normalization calculation on the input second expected value and second variance, and outputs a third expected value and a third variance; The activation function layer uses an activation function to perform a functional operation on the input third expected value and third variance, and outputs a fourth expected value and a fourth variance; The dropout layer randomly drops according to the fourth expected value and the fourth variance through a preset dropout probability, and outputs a fifth expected value and a fifth variance.

5. The data processing method according to claim 4, wherein, The physics-informed neural network model is provided with a target linear layer at the end of the last moment transfer architecture, and the physics-informed neural network model uses the second expected value output by the target linear layer as the predicted value, and uses the second variance output by the target linear layer as the target variance.

6. The data processing method according to claim 4, wherein The activation function is a smooth function capable of performing high-order derivatives to meet the calculation requirements of the physical law constraints, and the smooth function is a Sigmoid function, a Tanh function or a Swish function.

7. The data processing method according to claim 1, characterized in that The constructing the confidence interval according to the predicted value and the target variance includes: Calculate the target standard deviation according to the target variance; Calculate the product of the target standard deviation and a preset confidence coefficient to obtain a confidence deviation value; Calculate the sum of the predicted value and the confidence deviation value to obtain an upper limit value, and calculate the difference between the predicted value and the confidence deviation value to obtain a lower limit value; Construct a confidence interval with the upper limit value and the lower limit value.

8. A data processing device, characterized in that, Including: A data acquisition module for obtaining the data to be predicted; A model processing module, configured to input the data to be predicted into a pre-trained physics-informed neural network model, wherein the training process of the physics-informed neural network model incorporates physical laws as constraint conditions, so that the physics-informed neural network model performs moment transfer calculations and outputs a predicted value and a target variance; A confidence interval construction module, configured to construct a confidence interval based on the predicted value and the target variance; A quantization result output module, configured to determine an uncertainty quantization result of the predicted value according to the confidence interval; wherein the uncertainty quantization result is used to characterize the reliability of the predicted value.

9. A data processing device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the data processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the data processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Underground water level probability forecasting method based on interpretable Bayesian convolutional network

    CN117933316A

  • Life degradation calculation method based on Monte Carlo BiGRU gravity wave instrument

    CN118261033A

  • Automatically quantifying uncertainty of predictions provided by trained regression models

    CN118765399A

  • Uncertainty quantification method based on embedded physical information neural network

    CN120106142A

  • Information estimation apparatus and information estimation method

    US20180181865A1