Flow field data interpolation method based on optimal transportation theory
By transforming the flow field data interpolation problem into a discrete data distribution transportation problem, and utilizing optimal transportation theory and cost functions, the problem of insufficient interpolation accuracy and reliability in traditional methods is solved, achieving more accurate flow field data interpolation and extrapolation.
Patent Information
- Application Number
- CN202511081128.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-14
AI Technical Summary
Traditional flow field data interpolation methods ignore the inherent relationships in data distribution, resulting in decreased interpolation accuracy and reliability in complex flow fields or sparse data regions, making extrapolation impossible.
The interpolation problem is transformed into a transportation problem of discrete data distribution. The cost function and optimal transportation matrix are established using optimal transportation theory. Interpolation is achieved by minimizing the transformation cost, taking into account the overall characteristics and directional information of the data distribution.
It generates interpolation results that better match the actual physical characteristics of the flow field, improving the accuracy and reliability of interpolation. It can perform extrapolation in complex flow fields and sparse data regions, expanding the scope of applications.
Smart Images

Figure CN120950820A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of flow field data processing technology, and in particular to a flow field data interpolation method based on optimal transport theory. Background Technology
[0002] In many fields such as fluid mechanics, meteorology, and oceanography, flow field data interpolation is a crucial step. Traditional interpolation methods, such as linear interpolation, polynomial interpolation, Kriging interpolation, and triangulation interpolation, are typically based on numerical operations and do not consider the characteristics of flow field data as a "field," which is essentially a distribution. These methods ignore the inherent relationships between data distributions, leading to interpolation results that may not accurately reflect the true physical properties of the flow field, especially when dealing with complex flow fields or sparse data regions, where interpolation accuracy and reliability are significantly compromised. Furthermore, these methods, which only consider numerical values, cannot perform numerical extrapolation. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention provides a flow field data interpolation method based on optimal transport theory, which solves the technical problem that traditional interpolation methods often suffer from decreased interpolation accuracy and reliability due to their inability to accurately capture data distribution characteristics.
[0004] Traditional interpolation methods often focus on numerical relationships between data points, neglecting the overall characteristics of flow field data as a distribution. This invention, however, transforms the interpolation problem into a transportation problem of discrete data distributions, considering the distribution level and fully taking into account the inherent relationships between data distributions. Based on this, this invention provides a technical solution for flow field data interpolation based on optimal transport theory, which includes the following steps: S1. Obtain the flow field data of any flow field and divide it into N equal sub-regions, and normalize the data of each sub-region. S2. Randomly extract two batches of sizes from each sub-region. The data is treated as two discrete data distributions, and a discrete distribution is constructed based on the probability mass of any data point in the assigned discrete data distribution. and ,in and These are data points , The probability mass; S3. Establish a cost function to measure the data direction and value information in the flow field data. ; S4. Establish an optimal transportation plan to transform the flow field data interpolation problem into a discrete distributed transportation problem, and use a cost function. Minimize the optimal transport matrix used for interpolating flow field data. Obtain the interpolation result; S5. Normalize the interpolation result back to the original numerical range and perform smoothing.
[0005] This invention introduces optimal transportation theory into the field of flow field data interpolation, utilizing it to solve for the transportation scheme with the minimum conversion cost between two distributions, thereby achieving interpolation. By defining a reasonable cost function and constraints to solve for the optimal transportation matrix, a novel quantitative method based on distribution characteristics is provided for interpolation. This is the key technical means that distinguishes this invention from traditional interpolation methods.
[0006] Furthermore, in step S2, the specific process includes the following steps: Density is calculated based on the probability mass being proportional to the density of data points. The calculation formula is: in, Data points The density; Data points and The Euclidean distance between them; According to density Calculate and allocate data points probability mass The calculation formula is: in, For data points in a discrete data distribution , It represents the number of data points in a discrete data distribution.
[0007] Furthermore, in step S3, the data value information includes the uniformity of data spatial distribution, the consistency of unit dimensions, and the differences in values across different dimensions.
[0008] Furthermore, the cost function The expression is: in, Indicates from data points To data point The cost; and Data points , The numerical value; and It is the direction of speed. and It is a weight parameter, and .
[0009] Furthermore, the flow field data includes flat regions with simple field data variation trends and vortex regions with complex field data variation trends. Weighting parameters are used in the flat regions. Weighting parameters in the vortex region .
[0010] Furthermore, in step S4, the entropy-regularized optimal transportation plan is expressed as: in, It's the transportation cost; It is the entropy regularization term; It is a regularization parameter; Representing the cost matrix The Middle OK, The cost function value of the column; Represents the optimal transportation matrix The first in OK, The optimal transportation plan for the train.
[0011] Furthermore, the optimal transportation matrix The probability quality that satisfies the condition that the sum of the rows and columns equals the source distribution. and ,Right now: in, Representing the cost matrix The Middle OK, The cost function value of the column; Represents the optimal transportation matrix The first in OK, The optimal transportation plan for the train.
[0012] Furthermore, in step S4, the optimal transport matrix used to achieve flow field data interpolation is solved. The interpolation results are obtained, including: Initialization: Define the logarithmic cost matrix Initialize the logarithmic vector and It is a zero vector; in, It is a length of The vector, corresponding to the logarithmic scaling of the source distribution, is used to adjust the transport matrix. The rows such that the sum of each row equals the probability mass of the source distribution. ; It is a length of The vector, corresponding to the logarithmic scaling of the target distribution, is used to adjust the transportation matrix. The columns are such that the sum of each column equals the probability mass of the target distribution. ; Iterative update Update the logarithmic vector and ,Right now: in, and This indicates element-wise logarithmic and exponential operations; and These are the probability mass vectors of the source distribution and the target distribution, respectively; Indicates the first The next iteration; Represents the logarithmic cost matrix Transpose of; Constructing the optimal transport matrix: using the updated logarithmic vector and Calculate the optimal transportation matrix ,Right now: in, logarithmic vector and The and the One element; Convergence condition determination: Set a convergence threshold. When the logarithmic vector of two consecutive iterations and The difference is less than the convergence threshold When the algorithm converges, it is considered to have converged. in, ; Interpolation calculation: for the target point According to the optimal transportation matrix Calculate its interpolation The calculation formula is: in, Data points The corresponding flow field value; Data points in transportation To the target point The transportation volume is the corresponding transportation weight in the transportation matrix.
[0013] By employing the above technical solution, the present invention provides a flow field data interpolation method based on optimal transport theory, which has at least the following beneficial effects: 1. This invention transforms the interpolation problem into a transportation problem of discrete data distributions, considering it from a distributional perspective and fully taking into account the inherent relationships between data distributions. It utilizes this to solve for the optimal transportation matrix that minimizes the transformation cost between two discrete distributions, thereby achieving interpolation. By defining reasonable cost functions and constraints to solve for the optimal transportation matrix, it provides a novel, distribution-based quantitative method for interpolation.
[0014] 2. Because this invention considers the distribution characteristics of flow field data, it can generate interpolation results that better reflect the actual physical properties of the flow field. In complex flow fields and sparse data regions, traditional methods often suffer from decreased interpolation accuracy and reliability due to their inability to accurately capture data distribution characteristics. However, this invention, through optimization of optimal transport theory, can better reflect the true physical behavior of the flow field, improving the accuracy and reliability of the interpolation results in practical applications.
[0015] 3. Based on the overall characteristics of data distribution and the optimization of optimal transport theory, this invention possesses a certain extrapolation capability. This means that even in areas outside the known data range, relatively reasonable flow field predictions can be obtained, breaking through the limitations of traditional interpolation methods, which are usually limited to internal interpolation. This provides a broader application space for the prediction and analysis of flow field data and expands the application scope of interpolation methods. Attached Figure Description
[0016] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of the flow field data interpolation method in this invention. Detailed Implementation
[0017] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. This will allow for a full understanding of how the present application uses technical means to solve technical problems and achieve technical effects, and to facilitate its implementation.
[0018] This embodiment proposes a flow field data interpolation method based on optimal transport theory. For the first time, it transforms the interpolation problem into a transportation problem involving discrete data distributions, considering the distribution level and fully taking into account the inherent relationships between data distributions. It utilizes this method to solve for the optimal transport matrix that minimizes the transformation cost between two discrete distributions, thereby achieving interpolation. By defining reasonable cost functions and constraints to solve for the optimal transport matrix, it provides a novel, distribution-based quantitative method for interpolation. Figure 1 As shown, the method includes the following steps: S1. Obtain flow field data for any flow field and divide it into N equal sub-regions. Normalize the data in each sub-region, normalizing the velocity and pressure values to the [0,1] interval. In this embodiment, the flow field data is divided into N sub-regions, each containing M data points. Then, the data in each sub-region is normalized, normalizing the velocity and pressure values to the [0,1] interval. The normalization formula is: in, It is the raw data; , These are the minimum and maximum values of the data within that sub-region, respectively. It is the normalized data.
[0019] S2. Randomly extract two batches of sizes from each sub-region. The data is treated as two discrete data distributions, and a discrete distribution is constructed based on the probability mass of any data point in the assigned discrete data distribution. and ,in and These are data points , The probability quality.
[0020] This embodiment randomly extracts two batches of data from each sub-region, with data sizes of respectively. (Generally speaking) Two discrete data distributions are constructed. Probability masses are assigned to the data points based on their density in the flow field; for example, data points in high-density regions are assigned higher probability masses. It is assumed that the probability mass is proportional to the density of the data points, and the density is calculated using the following formula: in, Data points The density; Data points and The Euclidean distance between them.
[0021] Then based on density Calculate and allocate data points probability mass The calculation formula is: in, For data points in a discrete data distribution , It represents the number of data points in a discrete data distribution.
[0022] S3. Establish a cost function to measure the data direction and value information in the flow field data. The data value information includes the uniformity of data spatial distribution, the consistency of unit dimensions, and the differences in values across different dimensions.
[0023] For flow field data, choosing a suitable metric requires comprehensive consideration of the data's characteristics, physical meaning, and computational efficiency. Taking into account the non-uniform spatial distribution of the data, inconsistencies in unit dimensions, significant differences in values across different dimensions, and the directional information of the data, the cost function for flow field data is defined as follows: in, and Data points , The numerical value; and It is the direction of speed. and It is a weight parameter, and In the cost function, and Used to balance the contributions of velocity magnitude and direction. If If the difference in speed is larger, then the proportion of the total cost will be even greater; if If the velocity direction is smaller, then the difference in velocity direction accounts for a larger proportion of the total cost. This scheme has two configurations: one for regions with simple and gradual changes in field data, and another for regions with relatively flat field data trends. For regions with complex field data variations, such as vortex regions, .
[0024] This embodiment introduces a flow field-specific measurement method to calculate the cost function, taking into account the differences in velocity magnitude and direction. This method emphasizes the information of data direction and data value, and can more accurately reflect the physical characteristics of the flow field data.
[0025] S4. Establish an optimal transportation plan to transform the flow field data interpolation problem into a discrete distributed transportation problem, and use a cost function. Minimize the optimal transport matrix used for interpolating flow field data. Obtain the interpolation result. Based on the core idea of this embodiment, the interpolation problem is regarded as the optimal transportation problem of discrete data distribution. Therefore, the interpolation problem can be described using the following mathematical model: Given two discrete distributions , ,in and These are data points , The probability quality is determined by the cost function corresponding to the data points. The constructed cost matrix This indicates that the data points Transport to The cost. The goal is to find an optimal transportation matrix. This minimizes the total cost while satisfying the probability mass conservation and non-negativity constraints.
[0026] This embodiment uses an alternating normalized transport matrix. The rows and columns make the transport matrix This method satisfies specific row and column constraints and gradually approximates the optimal transportation weight matrix through iterative minimization of residuals. A regularization term is also introduced to ensure that the problem is a convex optimization problem. Compared to linear programming and dual methods, this method is faster and particularly suitable for solving transportation problems on large-scale datasets.
[0027] The optimal transportation plan with entropy regularization can be expressed as: in, It's the transportation cost; It is the entropy regularization term; This is the regularization parameter, which is usually set to 0.1.
[0028] This embodiment uses an iterative algorithm based on entropy regularization to efficiently solve the optimal transportation problem. By using entropy regularization, the algorithm transforms the optimal transportation problem into a convex optimization problem, significantly improving the solution efficiency, making it particularly suitable for large-scale datasets. Furthermore, the iterative update steps of the algorithm ensure the convergence of the transportation matrix.
[0029] The optimal transportation matrix The probability quality that satisfies the condition that the sum of the rows and columns equals the source distribution. and This constraint will be addressed during the algorithm iteration process by alternately normalizing the optimal transport matrix. This is implemented using rows and columns. It is denoted as: in, Representing the cost matrix The Middle OK, The cost function value of the column; Represents the optimal transportation matrix The first in OK, The optimal transportation plan for the train.
[0030] This embodiment solves the optimal solution to the constrained convex problem by using an iterative approach. The basic steps are as follows: Initialization: Define the logarithmic cost matrix Initialize the logarithmic vector and It is the zero vector. It is a length of The vector corresponds to the logarithmic scaling of the source distribution. Its function is to adjust the transportation matrix. The rows such that the sum of each row equals the probability mass of the source distribution. .
[0031] It is a length of The vector, corresponding to the logarithmic scaling of the target distribution, is used to adjust the transportation matrix. The columns are such that the sum of each column equals the probability mass of the target distribution. .
[0032] Iterative update Update the logarithmic vector and ,Right now: in, and This represents element-wise logarithmic and exponential operations. Using logarithmic space can avoid underflow and overflow problems in numerical computation and improve the numerical stability of the algorithm. and These are the probability mass vectors of the source distribution and the target distribution, respectively; Indicates the first The next iteration; Represents the logarithmic cost matrix The transpose of .
[0033] Constructing the optimal transport matrix: using the updated logarithmic vector and Calculate the optimal transportation matrix ,Right now: in, logarithmic vector and The and the Each element.
[0034] Convergence condition determination: A convergence threshold is usually set. When the logarithmic vector of two consecutive iterations and The difference is less than the convergence threshold When the algorithm converges, it is considered to have converged. in, ; Interpolation calculation: for the target point According to the optimal transportation matrix Calculate its interpolation The calculation formula is: in, Data points The corresponding flow field value; Data points in transportation To the target point The transportation volume is the corresponding transportation weight in the transportation matrix.
[0035] This embodiment transforms the flow field data interpolation problem into a discrete data distribution transportation problem, and uses optimal transportation theory to minimize the transformation cost to achieve interpolation. This method can better capture the distribution characteristics of flow field data and generate interpolation results that are more consistent with physical properties, especially in complex flow fields and sparse data regions.
[0036] S5. The interpolation result is denormalized to the original numerical range and smoothed. In this embodiment, the denormalization formula is: in, It is the interpolation result, that is, the normalized data; The value after inverse normalization; , These are the minimum and maximum values of the data within the sub-region, respectively.
[0037] Then, the inverse normalized interpolation results are smoothed to eliminate potential numerical fluctuations and ensure the physical validity of the interpolation results. Smoothing can be achieved using moving averages or other methods.
[0038] This invention treats the interpolation problem as a transportation problem of discrete data distribution and uses optimal transport theory to minimize the transformation cost, thus better capturing the distribution characteristics of flow field data. Compared with traditional methods, the interpolation results generated by this method better match the actual physical characteristics of the flow field, especially when dealing with complex flow fields and sparse data regions, where its advantages are more obvious.
[0039] Because this invention performs interpolation based on the overall characteristics of data distribution, it can perform extrapolation to a certain extent. This allows for relatively reasonable flow field predictions even in regions outside the known data range, providing broader applicability for practical applications.
[0040] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0041] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. Since the above embodiments are substantially similar to the method embodiments, they are described relatively simply; relevant parts can be referred to the descriptions of the method embodiments.
[0042] The above embodiments provide a detailed description of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A flow field data interpolation method based on optimal transport theory, characterized in that, The method includes the following steps: S1. Obtain the flow field data of any flow field and divide it into N equal sub-regions, and normalize the data of each sub-region. S2. Randomly extract two batches of data of size m and n from each sub-region as two discrete data distributions, and construct a discrete distribution based on the probability quality of any data point in the assigned discrete data distribution; S3. Establish a cost function to measure the data direction information and data value information in the flow field data; S4. Establish the optimal transportation plan to transform the flow field data interpolation problem into a discrete distribution transportation problem, and solve the optimal transportation matrix to achieve flow field data interpolation with the goal of minimizing the cost function to obtain the interpolation result. S5. Normalize the interpolation result back to the original numerical range and perform smoothing.
2. The flow field data interpolation method according to claim 1, characterized in that, In step S2, the specific process includes the following steps: Density is calculated based on the probability mass being proportional to the density of data points. The calculation formula is: in, Data points The density; Data points and The Euclidean distance between them; According to density Calculate and allocate data points probability mass The calculation formula is: in, For data points in a discrete data distribution , It represents the number of data points in a discrete data distribution.
3. The flow field data interpolation method according to claim 1, characterized in that, In step S3, the data value information includes the uniformity of data spatial distribution, the consistency of unit dimensions, and the differences in values across different dimensions.
4. The flow field data interpolation method according to claim 1 or 3, characterized in that, The cost function The expression is: in, Indicates from data points To data point The cost; and Data points , The numerical value; and It is the direction of speed. and It is a weight parameter, and .
5. The flow field data interpolation method according to claim 4, characterized in that, The flow field data includes flat regions with simple and gentle trends in field data changes, as well as vortex regions with complex field data changes. In the flow field of a flat region, the weighting parameters ; In the flow field of the vortex region, the weighting parameter .
6. The flow field data interpolation method according to claim 1, characterized in that, In step S4, the entropy-regularized optimal transportation plan is expressed as: in, It's the transportation cost; It is the entropy regularization term; It is a regularization parameter; Representing the cost matrix The Middle OK, The cost function value of the column; Represents the optimal transportation matrix The first in OK, The optimal transportation plan for the train.
7. The flow field data interpolation method according to claim 1, characterized in that, The optimal transportation matrix The probability quality that satisfies the condition that the sum of the rows and columns equals the source distribution. and ,Right now: in, Representing the cost matrix The Middle OK, The cost function value of the column; Represents the optimal transportation matrix The first in OK, The optimal transportation plan for the train.
8. The flow field data interpolation method according to claim 1, characterized in that, In step S4, the optimal transport matrix used to achieve flow field data interpolation is solved. The interpolation results are obtained, including: Initialization: Define the logarithmic cost matrix Initialize the logarithmic vector and It is a zero vector; in, It is a length of The vector, corresponding to the logarithmic scaling of the source distribution, is used to adjust the transportation matrix. The rows such that the sum of each row equals the probability mass of the source distribution. ; It is a length of The vector, corresponding to the logarithmic scaling of the target distribution, is used to adjust the transportation matrix. The columns are such that the sum of each column equals the probability mass of the target distribution. ; Iterative update Update the logarithmic vector and ,Right now: in, and This indicates element-wise logarithmic and exponential operations; and These are the probability mass vectors of the source distribution and the target distribution, respectively; Indicates the first The next iteration; Represents the logarithmic cost matrix Transpose of; Constructing the optimal transport matrix: using the updated logarithmic vector and Calculate the optimal transportation matrix ,Right now: in, logarithmic vector and The and the One element; Convergence condition determination: Set a convergence threshold. When the logarithmic vector of two consecutive iterations and The difference is less than the convergence threshold When the algorithm converges, it is considered to have converged. in, ; Interpolation calculation: for the target point According to the optimal transportation matrix Calculate its interpolation The calculation formula is: in, Data points The corresponding flow field value; Data points in transportation To the target point The transportation volume is the corresponding transportation weight in the transportation matrix.