Data compression method and apparatus, and data decompression method and apparatus
By constructing a mechanism model based on partial differential equations and using low-dimensional boundary terms and background fields to compress high-performance computing data, the problem of high data storage cost of high-dimensional data is solved, and efficient data compression and stable reconstruction results are achieved.
Patent Information
- Application Number
- PCT/CN2025/075445
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-07
- Filing Date
- 2025-01-27
- Publication Date
- 2026-02-12
AI Technical Summary
High-dimensional data storage for high-performance computing is costly, general data compression algorithms have poor compression ratios, and existing methods fail to fully utilize the mechanistic characteristics of the data, resulting in low compression rates and high computational complexity.
By constructing a mechanism model based on partial differential equations, low-dimensional boundary terms and background fields are used to compress the data to be compressed. The background field of the partial differential equations is optimized and adjusted to reduce the residuals, thereby achieving effective data compression.
It effectively reduces the compression rate of data compression, reduces storage overhead, and ensures the stability and accuracy of compression effect through the inversion process.
Smart Images

Figure CN2025075445_12022026_PF_FP_ABST
Abstract
Description
Data compression and decompression method and device
[0001] The present application claims priority to the Chinese patent application No. 202411076424.6, filed on August 7, 2024, entitled "Data compression and decompression method and device", the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0002] The present application relates to the technical field of data processing, in particular to a data compression and decompression method and device. BACKGROUND
[0003] High performance computing (HPC) data is an important part of unstructured data storage. In scientific computing and physical field simulation processes, a large amount of time evolution data of various scalar fields (space N dimensions + time dimension) is generated. The storage cost of high-dimensional data is high, and the compression ratio of general data compression algorithms is not good. SUMMARY
[0004] Embodiments of the present application provide a data compression and decompression method and device. The data compression method can effectively reduce the compression ratio of data compression and reduce storage overhead.
[0005] In a first aspect, the present application provides a data compression method, which comprises: obtaining to-be-compressed data, the to-be-compressed data having a mechanism characteristic; determining a mechanism model and a boundary term, the mechanism model describing a mechanism of the to-be-compressed data, the mechanism model comprising a partial differential equation, and the boundary term indicating a distribution boundary of the to-be-compressed data; performing inversion on the mechanism model based on the to-be-compressed data to obtain a parameter of the mechanism model, the parameter comprising a background field of the partial differential equation or a support set of the background field; and compressing the to-be-compressed data based on the background field or the support set of the background field and the boundary term.
[0006] The present application obtains the model parameter of the mechanism model constructed based on the partial differential equation (PDE) through inversion of the to-be-compressed data. The mechanism model can describe the mechanism of the to-be-compressed data. The to-be-compressed data is compressed by using the low-dimensional boundary term and the model parameter of the mechanism model, which effectively reduces the compression ratio of data compression and reduces storage overhead.
[0007] In a possible implementation, the inversion of the mechanism model based on the data to be compressed obtains a specific implementation of the parameters of the mechanism model, which is: initializing the background field based on the evolution trend of the data to be compressed to obtain the initial value of the background field; solving the partial differential equation based on the boundary term and the initial value of the background field to obtain the solution of the partial differential equation; determining a first residual based on the solution of the partial differential equation and the data to be compressed; and optimizing and adjusting the background field of the partial differential equation to obtain the background field by taking minimizing the first residual as an objective.
[0008] The background field of the partial differential equation is optimized and adjusted by minimizing the residual, so that the partial differential equation has a better approximation degree for the mechanism of the data to be compressed, the residual has sparsity and a small magnitude (i.e., a low storage bit number), has a low information entropy, and the compression effect is improved.
[0009] In another possible implementation, the solving of the partial differential equation based on the boundary term and the initial value of the background field obtains a specific implementation of the solution of the partial differential equation, which is: regularizing the background field, and a parameter of the regularization is determined based on noise of the data to be compressed; and solving the partial differential equation based on the boundary term and the initial value of the regularized background field to obtain the solution of the partial differential equation.
[0010] The stability of the inversion process is ensured by regularizing the background field.
[0011] In another possible implementation, the inversion of the mechanism model based on the data to be compressed obtains a specific implementation of the parameters of the mechanism model, which is: initializing a background field set based on the evolution trend of the data to be compressed to obtain the initial value of the background field set; solving the partial differential equation based on the boundary term and the initial value of the background field set to obtain the solution of the partial differential equation; determining a second residual based on the solution of the partial differential equation and the data to be compressed; and adjusting the set of the background field of the partial differential equation to obtain the set of the background field by taking minimizing the second residual as an objective.
[0012] The data compression is further reduced in compression rate by compressing the background field set with a lower dimension.
[0013] In another possible implementation, the background field's support set includes a boundary term of the background field and a boundary value parameter of the background field, the boundary term of the background field indicates a distribution boundary of the background field, and the boundary term of the background field and the boundary value parameter are used to determine the background field; the second residual is calculated based on the objective function, the objective function includes a residual term and a regularization term, a value of the residual term is determined based on the to-be-compressed data and the solution of the partial differential equation, and a value of the regularization term is determined based on a regularization parameter and the boundary value parameter of the background field; the background field's support set of the partial differential equation is adjusted to minimize the second residual, and a specific implementation of the background field's support set is that: a function value of the objective function is minimized to optimize and adjust the boundary term and the boundary value parameter of the background field; and the background field's support set is obtained based on the boundary term and the boundary value parameter of the background field that are adjusted and optimized.
[0014] In another possible implementation, the compression of the to-be-compressed data is also related to a target residual, the target residual indicates a difference between a solution of a target partial differential equation and the to-be-compressed data, and the solution of the target partial differential equation is obtained by solving the partial differential equation based on the boundary term and the background field that is adjusted and optimized.
[0015] In this possible implementation, the lossless compression of the data is implemented by saving the residual in the compression process, and the accuracy of the data compression is ensured.
[0016] In another possible implementation, the to-be-compressed data is compressed based on the background field or the background field's support set and the boundary term, and the foregoing further includes: the target residual is inspected; and if a mean value of the target residual is less than a preset threshold, the inspection passes.
[0017] The mechanism model obtained by the inversion is determined to be able to well depict the mechanism of the to-be-compressed data by inspecting the residual, so that the compression effect is ensured, or other compression methods are used for compression.
[0018] In another possible implementation, the parameters of the mechanism model further include a source term of the partial differential equation; and a specific implementation of obtaining the parameters of the mechanism model based on the to-be-compressed data includes: the background field is initialized based on an evolution trend of the to-be-compressed data to obtain an initial value of the background field; the solution of the partial differential equation is obtained by solving the partial differential equation based on the boundary term and the initial value of the background field; the source term is determined based on the solution of the partial differential equation and the to-be-compressed data; and the background field of the partial differential equation is adjusted and optimized to minimize the source term, so as to obtain the background field.
[0019] The background field of the partial differential equation is adjusted and optimized by the source term, and the calculation cost is smaller and the efficiency is higher.
[0020] In another possible implementation, the compression of the to-be-compressed data is also related to a target source term, the target source term indicating a difference between a solution of a target partial differential equation and the to-be-compressed data, the solution of the target partial differential equation being obtained by solving the partial differential equation based on the background field adjusted according to the boundary term and the optimization.
[0021] In another possible implementation, the accuracy of the parameters of the mechanism model is determined based on the accuracy of the to-be-compressed data.
[0022] In another possible implementation, the obtaining of the to-be-compressed data further includes: performing block processing on the to-be-compressed data to obtain a plurality of to-be-compressed data blocks. Optionally, the to-be-compressed data can be divided into a plurality of data blocks by considering the hardware processing capability and the compression parameters, and then compression is performed on each data block.
[0023] In another possible implementation, the boundary term includes an initial value and / or a boundary value, the initial value indicating a data distribution of an initial time of the to-be-compressed data, and the boundary value indicating a distribution boundary of the to-be-compressed data in space.
[0024] In a second aspect, the present application further provides a data decompression method, including: obtaining compressed data, the compressed data being obtained by compressing to-be-compressed data based on the data compression method described in the first aspect or any possible implementation manner of the first aspect; performing analysis on the compressed data to obtain an analysis result, the analysis result including parameters of a mechanism model and a boundary term, the mechanism model describing a mechanism of the to-be-compressed data, the mechanism model including a partial differential equation, the parameters including a background field or a branch set of the background field of the partial differential equation, and the boundary term indicating a distribution boundary of the to-be-compressed data; solving the partial differential equation based on the background field or the branch set of the background field and the boundary term to obtain a solution of the partial differential equation; and obtaining decompressed data based on the solution of the partial differential equation.
[0025] The decompression method provided in the present application can quickly decompress and reconstruct the original data distribution (i.e., the data distribution before compression).
[0026] In one possible implementation, the analysis result further includes a residual error, the residual error indicating a difference between the solution of the partial differential equation and the to-be-compressed data; and one specific implementation of obtaining the decompressed data based on the solution of the partial differential equation includes: obtaining the decompressed data based on the solution of the partial differential equation and the residual error. In this way, the original data that has not been compressed is reconstructed losslessly through the solution of the partial differential equation and the residual error.
[0027] In a third aspect, the present application provides a data compression apparatus, comprising a first obtaining module, a determining module, an inversion module and a compression module, wherein the first obtaining module is configured to obtain to-be-compressed data, the to-be-compressed data having a mechanism characteristic; the determining module is configured to determine a mechanism model and a boundary term, the mechanism model describing a mechanism of the to-be-compressed data, the mechanism model comprising a partial differential equation, and the boundary term indicating a distribution boundary of the to-be-compressed data; the inversion module is configured to perform inversion on the mechanism model based on the to-be-compressed data to obtain a parameter of the mechanism model, the parameter comprising a background field of the partial differential equation or a support set of the background field; and the compression module is configured to compress the to-be-compressed data based on the background field or the support set of the background field and the boundary term.
[0028] In one possible implementation, the inversion module is specifically configured to: initialize the background field based on an evolution trend of the to-be-compressed data to obtain an initial value of the background field; solve the partial differential equation based on the boundary term and the initial value of the background field to obtain a solution of the partial differential equation; determine a first residual based on the solution of the partial differential equation and the to-be-compressed data; and optimize and adjust the background field of the partial differential equation to obtain the background field, with the objective of minimizing the first residual.
[0029] In another possible implementation, one specific implementation of solving the partial differential equation based on the boundary term and the initial value of the background field to obtain the solution of the partial differential equation is: performing regularization on the background field, with a regularization parameter being determined based on noise of the to-be-compressed data; and solving the partial differential equation based on the boundary term and the initial value of the regularized background field to obtain the solution of the partial differential equation.
[0030] In another possible implementation, the inversion module is specifically configured to: initialize the support set of the background field based on the evolution trend of the to-be-compressed data to obtain an initial value of the support set of the background field; solve the partial differential equation based on the boundary term and the initial value of the support set of the background field to obtain a solution of the partial differential equation; determine a second residual based on the solution of the partial differential equation and the to-be-compressed data; and adjust the support set of the background field of the partial differential equation to obtain the support set of the background field, with the objective of minimizing the second residual.
[0031] In another possible implementation, the support set of the background field comprises a boundary term and an edge value parameter of the background field, the boundary term of the background field indicating a distribution boundary of the background field, and the boundary term and the edge value parameter of the background field being used to determine the background field; the second residual is calculated based on an objective function, the objective function comprising a residual term and a regularization term, a value of the residual term being determined based on the to-be-compressed data and the solution of the partial differential equation, and a value of the regularization term being determined based on a regularization parameter and the edge value parameter of the background field; and one specific implementation of adjusting the support set of the background field of the partial differential equation to obtain the support set of the background field, with the objective of minimizing the second residual, is: optimizing and adjusting the boundary term and the edge value parameter of the background field with the objective of minimizing a function value of the objective function; and obtaining the support set of the background field based on the boundary term and the edge value parameter of the background field that are optimized and adjusted.
[0032] In another possible implementation, the compression of the to-be-compressed data is further related to a target residual, the target residual indicating a difference between a solution of a target partial differential equation and the to-be-compressed data, the solution of the target partial differential equation being obtained by solving the partial differential equation based on the background field that is adjusted by optimization based on the boundary term.
[0033] In another possible implementation, the data compression apparatus provided in this application further includes an inspection module, which is configured to inspect the target residual; and if a mean value of the target residual is less than a preset threshold, the inspection passes.
[0034] In another possible implementation, the parameters of the mechanism model further include a source term of the partial differential equation; and the inversion module is specifically configured to: initialize the background field based on an evolution trend of the to-be-compressed data to obtain an initial value of the background field; solve the partial differential equation based on the boundary term and the initial value of the background field to obtain a solution of the partial differential equation; determine the source term based on the solution of the partial differential equation and the to-be-compressed data; and optimize and adjust the background field of the partial differential equation to obtain the background field, with minimization of the source term as a target.
[0035] In another possible implementation, the compression of the to-be-compressed data is further related to a target source term, the target source term indicating a difference between a solution of a target partial differential equation and the to-be-compressed data, the solution of the target partial differential equation being obtained by solving the partial differential equation based on the background field that is adjusted by optimization based on the boundary term.
[0036] In another possible implementation, the accuracy of the parameters of the mechanism model is determined based on the accuracy of the to-be-compressed data.
[0037] In another possible implementation, the data compression apparatus provided in this application further includes a blocking module, which is configured to perform blocking processing on the to-be-compressed data to obtain a plurality of to-be-compressed data blocks.
[0038] In another possible implementation, the boundary term includes an initial value and / or a boundary value, the initial value indicating a data distribution of an initial time of the to-be-compressed data, and the boundary value indicating a distribution boundary of the to-be-compressed data in space.
[0039] In a fourth aspect, the present application provides a decompression apparatus, comprising a second obtaining module, an analyzing module, a solving module and a decompression module. The second obtaining module is configured to obtain compressed data, which is obtained by compressing to-be-compressed data based on the data compression method described in the first aspect or any possible implementation manner of the first aspect. The analyzing module is configured to analyze the compressed data to obtain an analysis result, the analysis result comprising parameters of a mechanism model and a boundary term, the mechanism model describing a mechanism of the to-be-compressed data, the mechanism model comprising a partial differential equation, the parameters comprising a background field or a support set of the background field of the partial differential equation, and the boundary term indicating a distribution boundary of the to-be-compressed data. The solving module is configured to solve the partial differential equation based on the background field or the support set of the background field and the boundary term to obtain a solution of the partial differential equation. The decompression module is configured to obtain decompressed data based on the solution of the partial differential equation.
[0040] In a possible implementation, the analysis result further comprises a residual error, the residual error indicating a difference between the solution of the partial differential equation and the to-be-compressed data. The decompression module is specifically configured to obtain the decompressed data based on the solution of the partial differential equation and the residual error. In this way, the original data compressed without loss is reconstructed based on the solution of the partial differential equation and the residual error.
[0041] In a fifth aspect, the embodiments of the present application provide a computing device, comprising a memory and a processor, and the memory stores instructions, when the instructions are executed by the processor, the data compression method described in the first aspect or any possible implementation manner of the first aspect, and / or the data decompression method described in the first aspect or any possible implementation manner of the first aspect is implemented.
[0042] In a sixth aspect, the embodiments of the present application provide a computer readable storage medium, which stores a computer program, when the computer program is executed by a processor, the data compression method described in the first aspect or any possible implementation manner of the first aspect, and / or the data decompression method described in the first aspect or any possible implementation manner of the first aspect is implemented.
[0043] In a seventh aspect, the embodiments of the present application further provide a computer program or a computer program product, the computer program or the computer program product comprising instructions, when the instructions are executed, the computer executes the data compression method described in the first aspect or any possible implementation manner of the first aspect, and / or the data decompression method described in the first aspect or any possible implementation manner of the first aspect.
[0044] In an eighth aspect, the embodiments of the present application further provide a chip, comprising at least one processor and a communication interface, wherein the processor is configured to execute the data compression method described in the first aspect or any possible implementation manner of the first aspect, and / or the data decompression method described in the first aspect or any possible implementation manner of the first aspect. BRIEF DESCRIPTION OF DRAWINGS
[0045] FIG. 1 shows a schematic diagram of dimension reduction of mechanism data;
[0046] FIG. 2 shows a schematic diagram of a data compression algorithm provided by the embodiments of the present application;
[0047] FIG. 3 shows a schematic diagram of data compression based on background field inversion;
[0048] FIG. 4 shows a schematic diagram of a system architecture to which the data compression method provided by the embodiments of the present application can be applied;
[0049] FIG. 5 shows a schematic diagram of an application scenario of the data compression method provided by the embodiments of the present application;
[0050] FIG. 6 shows a flowchart of the data compression method provided by the embodiments of the present application;
[0051] FIG. 7 shows the relationship between variables after dimension reduction;
[0052] FIG. 8 shows a schematic diagram of a compression process of a data compression method provided by the embodiments of the present application;
[0053] FIG. 9 shows a flowchart of a data decompression method provided by the embodiments of the present application;
[0054] FIG. 10 shows a schematic diagram of a compression process of a data decompression method provided by the embodiments of the present application;
[0055] FIG. 11 shows a schematic diagram of the distribution of a solution of a passive transport equation and the distribution of a background field streamline;
[0056] FIG. 12 shows a schematic diagram of the distribution of original data and reconstructed data;
[0057] FIG. 13 shows a schematic diagram of a background field reconstruction result;
[0058] FIG. 14 shows a schematic diagram of the distribution of an initial field and a background field in a translation transport scenario;
[0059] FIG. 15 shows a schematic diagram of a reconstruction result in a translation transport scenario;
[0060] FIG. 16 shows a schematic diagram of the distribution of original data at initial and final time;
[0061] FIG. 17 shows a schematic diagram of the distribution of reconstructed data;
[0062] FIG. 18 is a schematic diagram of an effect summary comparison of a data compression method provided by an embodiment of the present application;
[0063] FIG. 19 is a schematic diagram of a structure of a data compression device provided by an embodiment of the present application;
[0064] FIG. 20 is a schematic diagram of a structure of a data decompression device provided by an embodiment of the present application;
[0065] FIG. 21 is a schematic diagram of a structure of a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0066] The term “and / or” mentioned in the present document is a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. The symbol “ / ” in the present document represents an or relationship of the associated objects, for example, A / B represents A or B.
[0067] The terms “first” and “second” and the like in the description and claims of the present document are used to distinguish different objects, and are not used to describe a specific order of the objects. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is only a distinguishing way adopted in the description of the objects with the same properties in the embodiments of the present application. In addition, the terms “include” and “have” and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device containing a series of units does not have to be limited to those units, but can include other units not clearly listed or inherent to these processes, methods, systems, products or devices.
[0068] In the embodiments of the present application, the words “exemplary” or “for example” are used to mean serving as an example, instance, or illustration. Any embodiment or design presented as “exemplary” or “for example” in the embodiments of the present application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of “exemplary” or “for example” is intended to present relevant concepts in a concrete manner.
[0069] In the description of the embodiments of the present application, unless otherwise specified, “a plurality of” means two or more, for example, a plurality of processing units means two or more processing units, and the like; a plurality of elements means two or more elements, and the like.
[0070] The data compression methods in the related art all have various problems. For example, in the related art, a traditional high-dimensional data compression method such as filtering performs basis transformation on global data, removes spatial redundancy, utilizes the sparsity of frequency domain coefficients, divides different bit planes for coding, and realizes storage space saving of existing scientific computing data compression schemes.
[0071] The method utilizes the dimension reduction idea and has good universality, but does not combine the internal mechanism (physical model) of data, so it fails to fully exert the compression potential for mechanism data compression.
[0072] In the related art, a data compression method based on prediction and entropy coding, common predictors include Lorenzo predictor, interpolation predictor, etc. The principle is to realize data dimension reduction through adjacent point prediction. When predicting a sample point, the Lorenzo predictor estimates the scalar value at the point through its processed adjacent points. It is assumed that the data is stored in a conventional grid form. The compressor and the decompressor both traverse the data points in a linear order. For subsequent compression, this method only needs to compress and store the boundary points and the errors of each point.
[0073] For the predictor method, a fixed empirical model is used, and the compression effect of scientific computing data for complex scenes has limited applicability. For data that meet a certain mechanism, the compression effect is good, but for data that meet another mechanism, the model error may be large due to the inability of the model to well depict the physical law, and the compression effect is limited. Thus, there is a large room for improvement. Usually, the characteristics of scientific computing data in complex scenes differ greatly, and specific models need to be designed according to the data characteristics to improve the compression rate.
[0074] In addition, the traditional predictor is often designed for two-dimensional image data. For high-dimensional data, the design of the predictor is a difficulty, and for spatiotemporal high-dimensional mechanism data, the general predictor is difficult to reflect the difference and correlation of the spatiotemporal mechanism, and the well-posedness of the calculation process in the high-dimensional case is not trivial.
[0075] In the related art, a learning algorithm based on neural network. The data-driven modeling based on deep neural networks (DNN) and the corresponding machine learning method represents the data by using neural network approximation, determines and stores the network parameters by optimization, and stores the data residual. The model obtained by this method is in the form of neural network, which is essentially different from the model in the form of differential equation and control equation. In addition, the model learning and data compression based on embedded physical information neural network (PINN). Essentially, it also approximates the data by using neural network, and the model is used as a constraint but finally realizes the data representation in the form of network.
[0076] For high-dimensional data models, the parameter scale and training overhead involved in neural networks also increase, and it is difficult to make good predictions (interpretation and generalization). Due to the complexity and nonlinearity of various neural network models, the interpretability is weak, and it is not convenient for theoretical analysis, including the estimation and analysis of compression effect. Such methods focus on general data approximation and do not involve physical law control equations, so for data scenes with mechanisms, they often cannot well reflect the data feature differences and strip out the internal data manifold. When the approximation degree is high, the model parameter storage cost is large, and when the approximation degree is low, the model bias (residual) storage is large, so the comprehensive compression effect is limited. At the same time, the compression and decompression calculation complexity is high, and the calculation overhead of the general deep learning method is often large, and the application scope of the black box model obtained by training is limited.
[0077] In data storage, a large amount of data has specific physical mechanisms due to its generation and acquisition mechanism. HPC scientific computing data, such as various physical field numerical simulation data of turbulent flow, atmosphere and ocean, corresponds to a clear model mechanism (control equation); various observation data, especially geophysical satellite remote sensing data such as sea surface height, temperature, atmospheric temperature and pressure, evolve to meet various conservation laws and also contain mechanisms; other industrial process data such as concentration and density, which meet the continuity and conservation of internal laws, also have their specific mechanisms.
[0078] For the above mechanism data, if the mechanism can be described by a model, a significant compression rate can be achieved by extracting the main mechanism. Mechanism data can usually be described by a differential equation, and for high-dimensional data, it is a partial differential equation or equation group (PDE), and the solution of the PDE is determined by appropriate initial and boundary value conditions and source terms, so the entire physical field (model data) is determined by low-dimensional well-posed data such as initial and boundary value data (data manifold), as shown in FIG. 1. Therefore, based on the idea of PDE dimension reduction, it is expected to achieve effective compression of data.
[0079] For two-dimensional data, there are traditional methods such as Lorenzo predictor to describe the data mechanism and achieve mechanism compression. However, for high-dimensional spatiotemporal data, the data has multiple dimensions, and the continuous change of the time dimension and the large difference between the spatial dimensions still have spatiotemporal correlation. Therefore, the idea of traditional predictors is difficult to be directly extended to high-dimensional cases.
[0080] In general, compared with the general compression method based on the empirical model, the data compression method based on the mechanism model has advantages for scientific computing data; at the same time, if the model is designed specifically, it is expected to reduce the computing overhead and facilitate theoretical analysis; further, by introducing the model degree of freedom, the specific model is reconstructed based on the data, and it is expected to make the data more consistent with the mechanism to further optimize the compression effect and generalization. From the data storage point of view, we select the appropriate model to better reflect the physical law, reduce the storage overhead of the model error part, and at the same time, we specifically use a kind of model with a wide range of application, that is, while using PDE for dimension reduction, the stability and robustness of the above process are ensured.
[0081] The specific implementation of the data compression method and device provided by the embodiments of the present application will be described in detail below with reference to the drawings.
[0082] The general data compression algorithm and image compression algorithm in the related art have problems of low data pertinence, low compression rate and low performance for data with clear mechanism such as scientific computing data and high-precision remote sensing data. For high-dimensional spatio-temporal data, the above problems are more prominent. Therefore, we need to provide a compression algorithm that is specific to the characteristics of mechanism data, fully extracts the data mechanism, reflects the difference and relevance of spatio-temporal dimensions, and realizes high compression rate and high performance with simple model calculation, low model deviation and small data support set. The technical problems to be solved include:
[0083] 1. Compression rate requirement
[0084] By using the unique characteristics and mechanism characteristics of mechanism floating point data such as scientific data and high-precision remote sensing data compared with general data and image data, a special compression algorithm is designed, which has higher compression rate than the current traditional dimension reduction compression method such as filtering, and is more effective in mining data spatio-temporal mechanism than the low-dimensional predictor method such as Lorenzo predictor.
[0085] 2. Adaptability and reliability requirement
[0086] A mechanism model with data pertinence can be obtained adaptively based on data, and the obtained model can realize reversible and stable reconstruction when used for data compression, avoiding the problems of low reliability and large calculation amount caused by complex model.
[0087] For mechanism data, the data support set is often very small, and even the low-dimensional information or local information such as initial and boundary values can uniquely determine the spatio-temporal global information. For such data, if the mechanism (model) can be reconstructed, the compression potential can be fully tapped.
[0088] Can the constructed model reconstruct the mechanism with only part of the data, predict / extend the information in the local approximate sense, and have the generalization ability that general compression methods based on neural networks and the like cannot achieve.
[0089] Figure 2 illustrates the principle of the data compression algorithm provided in this application embodiment. As shown in Figure 2, the principle of the data compression algorithm provided in this application is to reduce the dimensionality of high-dimensional mechanism data based on the PDE model to achieve compression. For example, the mechanism model constructed by partial differential equations describes the distribution mechanism of high-dimensional mechanism data. By combining data supports (e.g., initial values and boundary values) and other remainder terms (e.g., residuals, source terms, etc.), the distribution of high-dimensional mechanism data can be determined. The amount of data in the partial differential equations, data supports, and other remainder terms is very small, which can achieve compression of high-dimensional mechanism data. Compared with traditional data compression methods, the compression rate is greatly reduced, and storage overhead is saved.
[0090] Therefore, inverse PDEs can be used to effectively compress mechanistic data, and forward PDE calculations can be used to decompress the data.
[0091] For example, linear convection-diffusion partial differential equations are used to characterize the data mechanism for the basic forms of spatiotemporal evolution of physical fields, such as convection and diffusion. The partial differential equations are shown below:
[0092] Where (v x (x,y),v y (x,y) represents the background transport field. Data compression is achieved by inversely solving the aforementioned partial differential equation, and data decompression is achieved by forward computation of the same partial differential equation.
[0093] It is easy to understand that mechanistic data refers to data with specific mechanisms, such as HPC scientific computing data, including numerical simulation data of various physical fields such as turbulence, atmosphere, and ocean, which correspond to clear model mechanisms (governing equations); various observational data, especially geophysical satellite remote sensing data, such as sea surface height, temperature, and atmospheric temperature and pressure, whose evolution satisfies various conservation laws, also contain mechanisms; other data in industrial processes, such as concentration and density, which, if they conform to inherent laws such as continuity and conservation, also have their specific mechanisms. Mechanistic data is often high-dimensional spatiotemporal data, which has a large storage overhead and needs to be compressed during storage. The following uses high-dimensional spatiotemporal mechanistic data as an example to introduce the data compression principle provided in the embodiments of this application.
[0094] Figure 3 illustrates the principle of data compression based on background field inversion. As shown in Figure 3, the spatiotemporal data to be compressed is first matched with the corresponding mechanism equation type. For example, if the spatiotemporal data is data on the spatiotemporal evolution of a physical field, then the corresponding mechanism equation type is a linear convection-diffusion partial differential equation. Then, the background field of the mechanism equation is obtained by inversion. If the residual passes the test, the high-dimensional data is compressed using the data support (i.e., boundary terms) and the background field.
[0095] It should be noted that the data support set is also called well-posed data, and the data support set, such as initial values and boundary values of data, and a mechanism model can determine the distribution of data to be compressed.
[0096] In order to further improve the stability of the background field inversion, the embodiment of the application further provides an algorithm for stabilizing the coefficient (background field) of the inversion equation based on the regularization method, that is, adaptively determining the model based on the current data.
[0097] After the background field is obtained by inversion, the data storage is converted into the storage of the partial differential equation background field, the boundary term (such as the initial value and / or the boundary value), and the low-entropy source term.
[0098] The data compression method provided by the embodiment of the application can be applied to various compression scenarios. For example, the compression scenario of unstructured data storage scientific computing type floating point data storage.
[0099] FIG. 4 shows a system architecture diagram to which the data compression method provided by the embodiment of the application can be applied. As shown in FIG. 4, the embodiment of the application can be applied to a centralized or distributed storage system, as a compression algorithm of deep coding compression, to provide large-scale compression of scientific computing type floating point data.
[0100] FIG. 5 shows a schematic diagram of an application scenario of the data compression method provided by the embodiment of the application. As shown in FIG. 5, the data compression method provided by the embodiment of the application can be applied to effectively compress mechanism data in unstructured data, such as scientific computing floating point data, remote sensing image data, gene sequencing data, energy exploration data, and ocean data.
[0101] In unstructured data, scientific computing floating point data accounts for a large proportion. Compared with traditional data compression algorithms, the data compression method provided by the embodiment of the application can compress data by a large scale, thereby effectively reducing the storage cost.
[0102] The data compression method provided by the embodiment of the application can also be applied to large-scale compression of various high-precision time scalar field data, such as geophysical satellite remote sensing data, such as sea surface height, temperature, atmospheric temperature and pressure, concentration, and spatiotemporal data, whose spatiotemporal evolution is dominated by various conservation laws, typically showing background field transport, diffusion, and other behaviors, having physical mechanisms and corresponding control equations, and being high-precision floating point data. The general compression method for traditional image data is not applicable or not targeted.
[0103] FIG. 6 is a flowchart of a data compression method according to an embodiment of the present application. The method can be executed by any computing device, apparatus, platform or cluster of devices. The specific computing device for executing the method is not limited in the embodiments of the present application, and a suitable computing device can be selected for execution according to needs. For example, the method can be executed on a terminal device, i.e., the data compression method according to the embodiments of the present application can be deployed as a compressor plug-in in the storage system of the terminal device, and when storing mechanism data, the data compression method according to the embodiments of the present application is executed; or the method can be executed on a cloud device, and the cloud service is provided to users in the form of a cloud service to provide compression services for mechanism data; for the convenience of description, the form of the execution subject is not distinguished below, and all are described as data compressors. As shown in FIG. 6, the data compression method according to the embodiments of the present application at least includes steps S601 to S604.
[0104] In step S601, the data to be compressed is obtained.
[0105] The computing device can receive the data to be compressed sent by an external device (for example, a keyboard, a camera, a voice receiver, an external sensor, another computing device, etc.); or the computing device generates the data to be compressed by running an application.
[0106] The data to be compressed is data with mechanism characteristics. For example, the data to be compressed can be HPC scientific computing data, such as various physical field numerical simulation data of turbulent flow, atmosphere and ocean, etc.; various observation data, especially satellite remote sensing data of geophysics, such as sea surface height, temperature, atmospheric temperature and pressure, etc., which evolves to meet various conservation laws and also contains mechanisms; other data such as concentration and density in industrial processes, which also has its specific mechanism, such as continuity and conservation.
[0107] The obtained data is taken as the input of the data compressor, and at this time, the data to be compressed can also be referred to as the input data, and the data compressor obtains the data to be compressed.
[0108] Optionally, after obtaining the input data, the input data is preprocessed, and the preprocessing includes block processing, for example, the input data is processed in blocks to obtain a plurality of data blocks, and each data block is compressed subsequently.
[0109] For example, the data to be compressed is processed in blocks in space and time according to hardware processing capability and compression parameters. The hardware processing capability refers to the amount of data supported by hardware, such as memory, for compression processing at one time, and the compression parameters are based on the amount of data that can be effectively calculated at one time, for example, the effective amount of mechanism data that can be described by a mechanism model, and the accuracy will be affected when the amount exceeds this amount.
[0110] The preprocessing also includes validity inspection and denoising of the data.
[0111] In another example, after the preprocessing, the coefficient and the number of truncated bits of the calculation process are set according to the precision (decimal / binary, number of significant bits / absolute precision) of the input data, that is, the precision truncation is performed on the subsequent calculation process to keep consistent with the precision of the input data.
[0112] In step S602, the mechanism model and the boundary term are determined.
[0113] In this step, the mechanism model is determined according to the data distribution of the input data. For example, when the input data is the data of the spatio-temporal evolution of a physical field, a linear convection-diffusion partial differential equation can be fitted as the mechanism model; for example, when the distribution of the input data conforms to passive transport, a passive transport partial differential equation can be fitted as the mechanism model.
[0114] The boundary term is determined, which includes the initial value of the data to be compressed, that is, the data distribution of the data block at the earliest time, and / or the boundary value, which refers to the boundary value of the spatial distribution of the data to be compressed. In the following, the scheme of the embodiments of the present application is introduced by taking the initial value and the boundary value as examples (also referred to as the initial boundary value for short).
[0115] In step S603, the mechanism model is inverted based on the data to be compressed to obtain the parameters of the mechanism model.
[0116] First, the initial value of the background field of the partial differential equation is determined according to the evolution trend of the input data, and then the partial differential equation is solved based on the initial value of the boundary term and the background field to obtain the solution of the partial differential equation. The solution of the partial differential equation is matched with the data to be compressed, and the residual or source term is minimized to optimize and adjust the background field, thereby obtaining the background field.
[0117] In this way, the background field and the boundary term need to be encoded to realize data compression.
[0118] In another example, in order to further reduce the compression rate of data compression, the initial value of the background field set of the partial differential equation can be determined according to the evolution trend of the input data, and then the partial differential equation is solved based on the initial value of the boundary term and the background field set to obtain the solution of the partial differential equation. The solution of the partial differential equation is matched with the input data, and the residual or source term is minimized to optimize and adjust the background field set, thereby obtaining the background field set.
[0119] It can be understood that the meaning of matching the solution of the partial differential equation with the input data is that the solution of the partial differential equation is spatio-temporal data, and the input data is also spatio-temporal data. The solution of the partial differential equation and the input data at the same space-time are matched. For example, the solution of the partial differential equation u`(x1, y1, t1) and the input data u(x1, y1, t1) are matched.
[0120] Thus, the background field subset, the boundary term with smaller data volume are encoded subsequently, realizing data compression.
[0121] In another example, after the background field is obtained by optimization, the mechanism model also needs to be verified, the residual or source term is calculated, the obtained background field is tested, and the verification residual is calculated. If the residual meets the preset condition, for example, the residual distribution is approximately a Gaussian distribution with mean value of 0, it indicates that the mechanism model better describes the distribution mechanism of the data to be compressed, then the verification is passed, otherwise, the verification is failed, and other compression methods are used for compression.
[0122] In other examples, in order to realize lossless compression, after the background field is obtained by inversion, the mechanism model is solved based on the background field and the boundary term, and then the residual or source term is calculated according to the solution of the mechanism model and the corresponding data to be compressed. Subsequently, the background field / background field subset, the residual or source term and the boundary term are encoded to realize lossless compression of data.
[0123] For example, it is assumed that the spatio-temporal data to be compressed satisfies the passive transport equation:
[0124] where v x (x,y), v y (x,y) are two components of the velocity field (i.e. the background field), and it is assumed that they only depend on space. In practical applications, the evolution of scalar fields such as temperature, density, dye, floating pollutants in the corresponding fluid (such as air, water) is a typical passive transport problem. The data (well-posed condition) of the above model includes data initial value and boundary value. The forward problem is to calculate u(x,y,t) given the initial and boundary value conditions.
[0125] For actual structural data, where (i,j) represents a spatial grid, and n is a time layer. After being discretized by finite difference, the following difference equation is obtained:
[0126] That is, the data of the n+1 time layer can be determined by the data of the previous layer u n+1 =L UV u n , further, the values of all internal points only depend on the values of the external boundary and the initial layer, so in this case, only the boundary, the initial low-dimensional data, and the background field U i,j ,V i,j are stored, and the internal data can be directly calculated.
[0127] If the passive transport and diffusion effect are considered at the same time, the model is:
[0128] Where b.c represents boundary conditions (i.e. boundary values), and i.c represents initial conditions (i.e. initial values).
[0129] In general practical problems, if the data does not exactly satisfy the above ideal transport properties, i.e. the above homogeneous discrete equation, then
[0130] Method 1: Consider the residual, and let the solution of the above homogeneous equation be The original data is u, and the storage When the model is well approximated, It has sparsity and small magnitude (low storage bit number), corresponding to lower low information entropy, which can achieve compression.
[0131] Method 2: Consider the influence of the source term, i.e.
[0132] In this case, the corresponding difference equation is
[0133] At this time, in addition to storing the above low-dimensional well-posedness data, the full-dimensional source term data But when the model is well approximated, often f has sparsity or is a low-value disturbance, corresponding to lower low information entropy, at this time even if the full-dimensional source term data is stored, it is also smaller than the storage of the original data, and the compression effect is achieved.
[0134] Model inversion: According to the above principle, the key to the compression effect is whether the model better approximates the data mechanism. The degree of freedom of the above model is the two-dimensional background field U i,j ,V i,j (When considering diffusion, the diffusion coefficient constant v is also a model to be determined). Therefore, the above background field (model parameters) can be determined by 1, priori preset, 2, according to data optimization (learning), the latter can obtain a better approximation model and thus better compression effect. The "learning" of the above model is transformed into the following inversion problem:
[0135] Given u(x, t) on [0, t] x Ω, find v x ,v y , and predict the information at time [T, T]
[0136] Here, prediction means that since the background field is independent of time, the background field can be determined by using part of the data in a time period, so that longer time data prediction can be performed, and the calculation amount of the compression process is reduced.
[0137] In general, when the assumed model is homogeneous, from , the U i,j , V i,jThe system of linear algebraic equations, when the time level n≥3, forms an overdetermined system and can be solved using methods such as least squares. However, mathematically, this is an ill-posed problem; data perturbations and noise can make the process of solving the background field unstable, especially u. i,j In regions of minimal change, the aforementioned linear system becomes ill-conditioned, leading to significant biases in model reconstruction and greatly impacting prediction and compression performance. Therefore, this scheme further applies low-dimensional assumptions and regularization to the background field to ensure the stability of the inversion process. Let the background field be a potential flow (irrotational and incompressible):
[0138] Then the velocity potential ψ satisfies Δψ=0, according to the properties of harmonic functions. That is, the velocity potential is determined by the single-layer potential density. Decision, among which Ω represents the background field region, and Φ(x,y) represents the fundamental solution of the harmonic function in free space. Furthermore, due to the unique extension of the harmonic function, only the internal local data are needed. This determines the velocity potential ψ, and consequently the background field v. In summary, the background field inversion problem boils down to:
[0139] Given u(x,t)on[0,t]×Ω, find the boundary value μ of the velocity potential.
[0140] Figure 7 shows the relationships between the variables after dimensionality reduction.
[0141] In specific implementation, for Method 1 (optimizing residuals): let the single-layer potential density of the background field be discretized as... {e i (x)} is Discrete basis functions on the upper surface, and denote θ = {c k}, then ψ| can be determined from θ. Ω Then determine U i,j V i,j Finally, the model solution is obtained by solving the difference equation in the forward direction. Let θ be the above. The reflection of Where u0, u bord Given the initial and boundary value data, the background field inversion calculation is transformed into solving the following optimization problem:
[0142] in
[0143] α is the regularization parameter, determined a priori by the noise level of the data itself. θ is then obtained. * Then, calculate the corresponding And calculate the residuals And save θ, u0, u bord, the compression is completed. In the application, the matching data can be local data to save the algorithm overhead because the obtained equation has the extension calculation ability.
[0144] In the decompression, the original data is obtained by solving the difference equation in the forward direction. After the superposition of the residual error, the original data is obtained.
[0145] For the method 2 (optimizing the source term), according to the prior estimation of the equation, the small modulus of the source term can control the modulus of the solution, so the compression can be realized by equivalently optimizing and storing the source term according to the foregoing description. Similarly to the method 1, let be the background field determined by the boundary parameter θ, then the objective function to be optimized here is: 2 F(u; θ) = ‖f(θ)‖ 2
[0146] where the source term f(θ)
[0147] The source term f(θ) and θ, u0, u bord are stored. In the decompression, the non-homogeneous difference equation is solved, and the data is obtained.
[0148] For the above two optimization methods, the method 1 can directly optimize the residual error, so that the compression effect is directly reflected in the modulus of the residual error. The calculation of the objective function involves the explicit time marching equation. The calculation amount of the method 2 is generally equivalent to that of the method 1, and the program implementation of the method 2 is relatively simple.
[0149] In step S604, the data to be compressed is compressed based on the background field or the branch set of the background field and the boundary term.
[0150] For the above step, only the background field is obtained, and the background field and the boundary term are compressed and encoded to realize the lossy compression of the input data, and the compressed data is obtained.
[0151] For the above step, the branch set of the background field is obtained, and the branch set of the background field and the boundary term are compressed and encoded to realize the lossy compression of the input data, and the compressed data is obtained.
[0152] For the above step, the background field and the residual error are obtained, and the background field, the boundary term and the residual error are compressed and encoded to realize the lossless compression of the input data, and the compressed data is obtained.
[0153] For the above step, the branch set of the background field and the residual error are obtained, and the branch set of the background field, the boundary term and the residual error are compressed and encoded to realize the lossless compression of the input data, and the compressed data is obtained.
[0154] After the background field and the source term are obtained in the above steps, the background field, the source term and the boundary term are compressed and encoded to realize lossless compression of the input data and obtain compressed data.
[0155] FIG. 8 shows a compression process diagram of the data compression method provided in the embodiments of the present application. As shown in FIG. 8, the initialization of the background field, the optimization of the velocity field branch set (boundary potential), the calculation of the source term and truncation, the calculation, the inspection of the residual error and the encoding of the velocity field, the source term and the boundary term are sequentially performed on the input data to obtain compressed data. For example, the boundary term (i.e., the initial boundary value condition), the source term and the velocity field (i.e., the background field) are encoded, the header information is saved, and the compressed file is written.
[0156] The data compression method provided in the embodiments of the present application only needs a single current data to perform the background field inversion and does not depend on other training data, databases, large models and the like.
[0157] Based on model learning. The optimization process (low dimension) of the background field data branch set and the explicit solving of the model equation can be processed in parallel. Due to the prediction ability of the model itself, the data required by the optimization process can be only part of the original data
[0158] The data compression method can be suitable for various types of high-dimensional data evolving in space and time and can obtain the background field / mechanism, realize data compression and obtain the physical mechanism. When the data is scientific computing data such as numerical solution of differential equations, the compression effect is outstanding, and the optimal effect is obtained for pure convection and heat equations.
[0159] The data compression method can realize lossless compression based on the source term precision control, and can realize lossy compression with mechanism based on the well-posedness theory of differential equations and the distribution of the data source term
[0160] Therefore, the data compression method provided in the embodiments of the present application is particularly suitable for storing data with physical mechanism, i.e., data satisfying the development equation, such as observation data of atmosphere, ocean or data of industrial processes. Since the compression effect depends on the approximation degree of the model, the equation does not need to be strictly satisfied. It should be noted that the specific implementation of the above algorithm, including the form of the “data branch set” and the uniqueness and stability of the reconstruction of the background field, is based on the related theory of differential equation inversion problems, and is therefore not trivial.
[0161] For the data compression method provided in the embodiments of the present application, the embodiments of the present application further provide a data decompression method.
[0162] FIG. 9 is a flow diagram of a data decompression method according to an embodiment of the present application. The method can be executed by any computing device, apparatus, platform or cluster of devices. The present application does not limit the specific computing device that executes the method, which can be selected as needed. For example, the method can be implemented on a terminal device, i.e., the data decompression method according to an embodiment of the present application can be deployed as a decompressor plug-in in the storage system of a terminal device, and can be executed to reconstruct original data from compressed data compressed using the data compression method according to an embodiment of the present application. The method can also be implemented on a cloud device, i.e., the data decompression method according to an embodiment of the present application can be provided as a cloud service to users, and can be executed to decompress compressed data compressed using the data compression method according to an embodiment of the present application. Hereinafter, the form of the execution subject is not distinguished, and is described as a data decompressor. As shown in FIG. 9, the data compression method according to an embodiment of the present application includes at least steps S901 to S904.
[0163] In step S901, compressed data is obtained.
[0164] When the compressed data needs to be decompressed, the data decompressor is called to decompress the compressed data, and the original data is obtained. The original data refers to data before compression, i.e., the data to be compressed or the input data. The compressed data refers to compressed data compressed using the data compression method according to an embodiment of the present application.
[0165] The compressed data that needs to be decompressed is input into the data decompressor, i.e., the compressed data is obtained.
[0166] In step S902, the compressed data is parsed to obtain a parsing result, which includes the parameters of the mechanism model and the boundary term.
[0167] The compressed data is parsed to obtain the parameters of the mechanism model and the boundary term. The content of the parsing result obtained is related to the specific compression scheme used.
[0168] For example, the compressed data is obtained by compressing and encoding the background field and the boundary term, and the parsing result includes the background field and the boundary term.
[0169] For another example, the compressed data is obtained by compressing and encoding the background field basis set and the boundary term, and the parsing result includes the background field basis set and the boundary term.
[0170] For another example, the compressed data is obtained by compressing and encoding the background field, the boundary term and the residual error, and the parsing result includes the background field basis set, the boundary term and the residual error.
[0171] For another example, the compressed data is obtained by compressively encoding the background field, the boundary term and the source term, and the analysis result includes the background field, the boundary term and the source term.
[0172] For another example, the compressed data is obtained by compressively encoding the background field, the boundary term and the source term, and the analysis result includes the background field, the boundary term and the source term.
[0173] For another example, the compressed data is obtained by compressively encoding the background field, the boundary term and the source term, and the analysis result includes the background field, the boundary term and the source term.
[0174] In step S903, the partial differential equation is solved based on the background field or the background field set and the boundary term, to obtain a solution of the partial differential equation.
[0175] After the analysis result is obtained, the partial differential equation solver is invoked to solve the partial differential equation, to obtain a solution of the partial differential equation.
[0176] For example, the background field is substituted into the partial differential equation, and the partial differential equation is solved based on the boundary term, to obtain a solution of the partial differential equation.
[0177] For another example, the complete background field is calculated based on the background field set, the background field is substituted into the partial differential equation, and the partial differential equation is solved based on the boundary term, to obtain a solution of the partial differential equation.
[0178] In step S904, the decompressed data is obtained based on the solution of the partial differential equation.
[0179] For lossy compression, i.e., the residual or the source term is not saved during compression, the solution of the partial differential equation is directly used as the original data, i.e., the decompressed data.
[0180] For lossless compression, i.e., the residual or the source term is saved during compression, after the solution of the partial differential equation is obtained, the residual or the source term is added to the solution of the partial differential equation, to obtain the original data, i.e., the decompressed data.
[0181] FIG. 10 shows a compression process schematic diagram of a data decompression method provided in an embodiment of the present application. As shown in FIG. 10, the compressed data is analyzed to obtain the background field set, the initial boundary value and the source term, then the background field is calculated based on the background field set, the background field and the initial boundary value are substituted into the model solver to solve, to obtain a solution of the model (i.e., a solution of the partial differential equation), and the residual is added to the solution of the model to obtain the decompressed data.
[0182] The application of the data compression method and the data decompression method provided in the embodiments of the present application in practice will be described below through a specific example.
[0183] The compression process includes the following steps:
[0184] Step 1. Preprocessing of input data. The data is blocked, mean and anomaly calculated according to hardware processing capacity and compression parameters. According to the data precision bits (decimal / binary; significant bits / absolute precision), the coefficient and the truncation bits of the calculation process are set. The four direction uniform velocity field (v x ,v y ) = (±1, 0), (0, ±1) and the corresponding velocity potential function ψ(x, y) = ±x, ±y are calculated to determine the optimized initial value of the velocity potential function by comparing the residual error. The input coefficient initial value can also be specified according to other prior tests.
[0185] Step 2. Model background field inversion process. Based on the blocked data, the above reduced dimension background field (velocity potential) is optimized, and the objective function is
[0186] where i, j traverse the boundary space data points, and n traverses the time layer. The optimization process corresponds to minimizing the above objective function. The calculation of the objective function only involves coefficient multiplication and algebraic summation process, and the explicit promotion only depends on the last layer, so the calculation overhead is small, and it is also convenient to promote to GPU deployment calculation. The default optimization algorithm adopted is Levenberg-Marquardt algorithm, which takes into account the convergence speed and stability. The initial value of the parameter is selected in step 1, and other optional parameters include the number of iteration loops, the number of target function calls, the tolerance, etc. If the user specifies global optimization in the settings, the SQP algorithm or other global optimization algorithms can be used, and the objective function remains unchanged. The discrete components of the optimized velocity potential function, the edge values of the original blocked data, and the initial values are used as data support sets.
[0187] Step 3. Calculate, verify, and save the residual error. 1. Based on the background field velocity potential edge value , the velocity field is calculated , and then the approximate solution is obtained by forward solving the transport equation , the residual error is calculated , and the extraction mechanism effect is judged. If the residual error distribution is approximately a Gaussian distribution with a mean of 0, it means that the mechanism has a good extraction. 2. According to the data precision, the residual term is truncated. Compare the residual magnitude and the original data magnitude. If the residual modulus is significantly lower than the original data, the compression effect is better. In the case of lossless compression, the information entropy of the residual term is also affected by the data quality. If the data itself has a large noise, the compression effect is limited, and if the data is white noise, there is no compression effect.
[0188] Step 4. Encode and store data. The data required for storage is the above background field support set (c kk = 1, 2…m, i.e., m coefficients), residual term (if the original data is an N1×N2×N3 matrix, then the residual term is (N1-2)×(N2-2)×(N3-1), boundary data i = 0, N1 or j = 0, N2; n = 2, ..., N3, initial data (N1×N2 matrix). Then, the data to be stored is encoded and output.
[0189] The decompression process mainly includes the following specific steps:
[0190] Step 1. After parsing the compressed file, obtain the residual term, initial value, boundary value term, and background field support.
[0191] Step 2. Based on the background field support... Calculate the background field U i,j V i,j .
[0192] Step 3. Substitute into the difference equation Explicit advancement to obtain approximate solutions The final decompressed data is obtained by superimposing the residuals. The above equations are solved in an explicit format in the forward direction, which results in fast computation speed, low memory overhead, and parallel processing capability.
[0193] In order to verify the compression effect of the data compression method provided in this application, an experiment was conducted on the data compression method provided in the application.
[0194] This application selects several scientific computing data examples to illustrate that the data compression method provided in this application has the best performance when the data itself is a solution to a transport equation. It should be noted that in most cases, although the data satisfies a differential equation, it is not necessarily a linear equation, and the coefficients are not necessarily time-independent. To demonstrate that this algorithm still has good applicability, we examined the following scenarios:
[0195] 1. Numerical solution of the passive transport equation
[0196] We choose the initial distribution u0(x,y)=exp(-16((x-0.5)) 2 +(y-0.5) 2 (x,y)∈[0,1]×[0,1], discretized into 51×51×75 array blocks, with a time step dt=0.005, and a non-reflective boundary condition. The background field distribution is as follows: r1 = (x - x1) 2 +(y-y1) 2 r² = (x - x²) 2 +(y-y2) 2x1 = 0, x2 = 1, y1 = -0.1, y2 = 1.1
[0197] The original data is obtained by solving the difference equation with the initial boundary conditions, and the data precision is 16 bits. The data distribution is shown in (a) of FIG. 11, and the background field streamline distribution is shown in (b) of FIG. 11. It should be noted that for larger arrays, on the one hand, parallel processing can be performed after blocking, and the calculation processes between blocks are independent of each other; on the other hand, from the perspective of local linearization, linear approximation on a local region is more effective for cases where the background field changes greatly.
[0198] For the original binary file size of 1565K, after zip encoding, it is 527K, the compression rate is 33.7%, and after the data compression method provided by the application is used, it can be further compressed by 60% to 318K, which shows good effect. The specific data summary is shown in Table 1.
[0199] Table 1: Comparison of scientific computing data compression effect
[0200] The background field reconstruction is shown in (c) of FIG. 13. Due to the ill-posedness of the inversion problem itself and the further dimension reduction and discretization of the background field, the reconstruction result is relatively ideal. For the reconstruction of the final data field, the original data and the reconstruction data distribution at the final time are shown in FIG. 12, and the residual error is as small as 10 -3 Since only the residual error is stored, the actual effective number of bits of the stored data is effectively reduced, thereby explaining the compression effect.
[0201] It should be noted that 1. The passive transport process described above is not trivial, and compared with simple translation and rotation, it increases the deformation effect. For the case of simple translation, the method can be covered as a special case, as shown in the results of FIGS. 14-15. 2. The model reconstruction (background field and intermediate quantities or parameters) allows for some deviation, as long as the final data field error has a small modulus to achieve compression. 2. By further modeling the background field, stable reconstruction can be obtained, and the data support set can be further reduced. 3. In the embodiments of the application, the reconstruction of the background vector field adopts the dimension reduction idea and the regularization method of boundedness constraint, to realize the stable reconstruction of the background field. If the inversion algorithm is not suitable, such as direct point-by-point reconstruction, due to the instability of the problem, it will have a great impact on the result. As shown in (right) of FIG. 13, without using the regularization method described above, the background field is directly inverted, and the background field reconstruction has a large deviation, and even cannot reduce the residual error or predict the data.
[0202] In addition, we try to understand the function case, take u(x, y, t) = sin(πt-x)cos(2x+2(y-t))e -0.01tThe solution does not strictly correspond to the passive transport equation, but the method can also achieve good compression effect, as shown in Table 2.
[0203] Table 2: Comparison of compression effect of analytical function data
[0204] Corresponding to the technical problems to be solved by the embodiments of the present application, the above examples solve the following key technical problems: first, for high-precision mechanism data with space-time correlation, in the case that general compression methods are difficult to compress, the method based on the PDE mechanism model dimension reduction idea compression method realizes effective compression, and achieves a compression rate of 60% or even lower on the basis of ordinary coding. Second, the method has good adaptability to various mechanism data, and even if the data itself does not strictly satisfy the control equation, the compression effect can be obtained.
[0205] Compared with related technologies, the improvements of the data compression method provided by the embodiments of the present application include the following aspects. Improvement one: compared with general high-dimensional data compression methods, the method can realize effective compression of high-precision mechanism data. Mechanism data, especially the solution of differential equation, often does not have sparsity, the variable changes continuously, the value distribution is wide, and the corresponding information entropy is high, so the traditional image compression or coding method is difficult to obtain effective compression. Improvement two: compared with fixed equation models, the method introduces model degrees of freedom such as background field coefficients, which can effectively expand the model applicability, can adaptively optimize different data to obtain a model that is more suitable for the data, thereby reducing the model deviation and obtaining better compression effect. In addition, in the case of allowing lossy compression, since the random noise part can be discarded, the method will further highlight the advantages.
[0206] For actual observation data, remote sensing data and the like, there is often a certain level of random noise, so in the reconstruction process, on the one hand, noise will affect the compression effect, and pure white noise data cannot be effectively compressed by a mechanism model, so for these actual data, the compression effect will generally be reduced.
[0207] In addition, for mechanism under the action of diffusion, in order to better make the mechanism model fit the data, the method considers introducing the diffusion term vΔu into the equation, wherein the diffusion coefficient v is a to-be-determined model parameter, and in calculation, the background field is optimized together. It should be noted that after introducing diffusion, in order to avoid the instability of the forward problem solution caused by inverse diffusion, the size of the mode of the data at both ends of the time dimension is used to judge the time evolution direction, so as to correctly select the diffusion initial layer.
[0208] The following takes the ISABEL-Hurricane dataset as an example for the spatio-temporal evolution data of atmospheric composition. A 150x150x10 floating-point array is losslessly compressed, and the data distribution at the initial and final time is shown in FIG. 16. It can be seen that there is a significant attenuation phenomenon, and noise is contained. The reconstructed data through model inversion is shown in FIG. 17. It can be seen that the reconstruction residual is significantly reduced by an order of magnitude (note that here there is no need for accurate reconstruction, as long as the residual order is effectively reduced, the compression effect can be achieved). The specific compression effect is shown in Table 3.
[0209] Table 3: Comparison of data compression effect
[0210] The effect summary comparison of the data compression method provided by the embodiments of the present application is shown in FIG. 18. It can be seen that the data compression method provided by the embodiments of the present application can better extract the data mechanism, and still achieve the compression effect under the disadvantage of large noise.
[0211] The data compression method provided by the embodiments of the present application is based on the idea of model inversion. For data that generally follows a certain spatio-temporal evolution, the model degrees of freedom such as the background field and the diffusion coefficient are introduced, and the equation coefficients are inverted through the current data, so as to obtain the specific model / equation that best matches the data. The effect is outstanding for data with significant inherent mechanism.
[0212] The data compression method provided by the embodiments of the present application is based on the idea of differential equation theory and local linearization. Through a single or even part of the data, the background field can be stably reconstructed, and the mechanism that the data approximately satisfies can be obtained. The cost is small, and it is consistent with the law of the data itself.
[0213] Compared with the two-dimensional section mechanism model compression method, the correlation in the time direction is considered, so that multiple layers of data can be processed at one time. The degrees of freedom of the model itself are limited to the equation coefficients, and the dimensionality of the coefficients is also reduced, achieving a balance between computational efficiency and compression effect.
[0214] The mechanism model of the embodiments of the present application has good prediction ability and generalization characteristics, and has potential applications in high-dimensional data completion and other scenarios.
[0215] The algorithm for reconstructing the background field can be applied to velocity field measurement and other practical engineering applications.
[0216] The data compression method provided in the embodiments of the present application is suitable for compression of high-dimensional data with spatial and temporal continuity. The method uses the idea of model inversion, local linearization and related theories of differential equations, and proposes a model dominated by transport and diffusion mechanisms and a complete set of algorithms for obtaining optimal model coefficients and compressed data based on data. The method can support the more optimal compression needs of massive data era, various scientific calculations and measurement data, and high-precision remote sensing data. For other types of data, it is also expected to try the mechanism compression algorithm provided in the embodiments of the present application after preprocessing or suitable scale decomposition.
[0217] Subsequently, it can be further applied to more general models, such as diffusion coefficient and zero-order variable coefficient cases. These coefficients can be sparsely or low-dimensionally assumed, so there is no essential increase in storage overhead, but better compression effects can be achieved. For discrete formats, it can also be extended to implicit or multi-layer formats in the time direction, and compatible with time second-order derivative cases. The uniqueness and stability of the coefficient inversion of the above equation are guaranteed by the theory of partial differential equation inversion. The method represents a new attempt to cross the differential equation inversion theory and data compression.
[0218] Based on the same idea as the foregoing embodiment of the data compression method, the embodiments of the present application also provide a data compression device 1900. The data compression device can compress mechanism data, reduce compression rate, and reduce storage overhead. The data compression device 1900 includes units or modules for implementing each step of the data compression method shown in FIGS. 2-8.
[0219] FIG. 19 is a structural schematic diagram of a data compression device provided in the embodiments of the present application. As shown in FIG. 19, the data compression device 1900 includes a first acquisition module 1901, a determination module 1902, an inversion module 1903, and a compression module 1904. The first acquisition module 1901 is configured to acquire to-be-compressed data, and the to-be-compressed data has a mechanism characteristic. The determination module 1902 is configured to determine a mechanism model and a boundary term. The mechanism model describes the mechanism of the to-be-compressed data, and the mechanism model includes a partial differential equation. The boundary term indicates the distribution boundary of the to-be-compressed data. The inversion module 1903 is configured to perform inversion on the mechanism model based on the to-be-compressed data, to obtain parameters of the mechanism model. The parameters include a background field of the partial differential equation or a support set of the background field. The compression module 1904 is configured to compress the to-be-compressed data based on the background field or the support set of the background field and the boundary term.
[0220] In a possible implementation, the inversion module 1903 is specifically configured to: initialize the background field based on the evolution trend of the data to be compressed, to obtain an initial value of the background field; solve the partial differential equation based on the boundary term and the initial value of the background field, to obtain a solution of the partial differential equation; determine a first residual based on the solution of the partial differential equation and the data to be compressed; and optimize and adjust the background field of the partial differential equation to obtain the background field, with the objective of minimizing the first residual.
[0221] In another possible implementation, a specific implementation of solving the partial differential equation based on the boundary term and the initial value of the background field to obtain the solution of the partial differential equation includes: regularizing the background field, with a parameter of the regularizing being determined based on noise of the data to be compressed; and solving the partial differential equation based on the boundary term and the initial value of the regularized background field, to obtain the solution of the partial differential equation.
[0222] In another possible implementation, the inversion module 1903 is specifically configured to: initialize a background field set based on the evolution trend of the data to be compressed, to obtain an initial value of the background field set; solve the partial differential equation based on the boundary term and the initial value of the background field set, to obtain a solution of the partial differential equation; determine a second residual based on the solution of the partial differential equation and the data to be compressed; and adjust the set of the background field of the partial differential equation to obtain the set of the background field, with the objective of minimizing the second residual.
[0223] In another possible implementation, the set of the background field includes a boundary term and a boundary value parameter of the background field, the boundary term of the background field indicates a distribution boundary of the background field, and the boundary term and the boundary value parameter of the background field are used to determine the background field; the second residual is calculated based on an objective function, the objective function includes a residual term and a regularization term, a value of the residual term is determined based on the data to be compressed and the solution of the partial differential equation, and a value of the regularization term is determined based on a regularization parameter and the boundary value parameter of the background field; and a specific implementation of adjusting the set of the background field of the partial differential equation to obtain the set of the background field includes: optimizing and adjusting the boundary term and the boundary value parameter of the background field, with the objective of minimizing a function value of the objective function; and obtaining the set of the background field based on the boundary term and the boundary value parameter of the background field that are adjusted and optimized.
[0224] In another possible implementation, the compression of the data to be compressed is also related to a target residual, the target residual indicates a difference between a solution of a target partial differential equation and the data to be compressed, and the solution of the target partial differential equation is obtained by solving the partial differential equation based on the boundary term and the background field that is adjusted and optimized.
[0225] In another possible implementation, the data compression apparatus 1900 provided in the present application further includes an inspection module 1905, configured to inspect the target residual; and if a mean value of the target residual is less than a preset threshold, the inspection passes.
[0226] In another possible implementation, the parameters of the mechanism model further include a source term of the partial differential equation; and the inversion module 1903 is specifically configured to: initialize the background field to obtain an initial value of the background field based on an evolution trend of the data to be compressed; solve the partial differential equation based on the boundary term and the initial value of the background field to obtain a solution of the partial differential equation; determine the source term based on the solution of the partial differential equation and the data to be compressed; and optimize and adjust the background field of the partial differential equation to obtain the background field, with the objective of minimizing the source term.
[0227] In another possible implementation, the compression of the data to be compressed is further related to a target source term, the target source term indicating a difference between a solution of a target partial differential equation and the data to be compressed, the solution of the target partial differential equation being obtained by solving the partial differential equation based on the boundary term and the background field that is adjusted and optimized.
[0228] In another possible implementation, the accuracy of the parameters of the mechanism model is determined based on the accuracy of the data to be compressed.
[0229] In another possible implementation, the data compression apparatus 1900 provided in the present application further includes a blocking module 1906, configured to perform blocking processing on the data to be compressed to obtain a plurality of data blocks to be compressed.
[0230] In another possible implementation, the boundary term includes an initial value and / or a boundary value, the initial value indicating a data distribution at an initial time of the data to be compressed, and the boundary value indicating a distribution boundary of the data to be compressed in space.
[0231] The data compression apparatus 1900 according to the embodiments of the present application can correspond to performing the methods described in the embodiments of the present application, and the above and other operations and / or functions of each module in the data compression apparatus 1900 are respectively for realizing the corresponding flow of each method in FIGS. 2-8, which will not be described herein again for brevity.
[0232] Based on the same idea as the foregoing embodiment of the data decompression method, the present embodiment further provides a data decompression apparatus 2000, which can decompress compressed data compressed by the data compression method provided in the embodiments of the present application, and reconstruct the original data. The data decompression apparatus 2000 includes units or modules for realizing each step of the data decompression method shown in FIGS. 9-10.
[0233] FIG. 20 is a structural schematic diagram of a data decompression apparatus according to an embodiment of the present application. As shown in FIG. 20, the data decompression apparatus 2000 includes a second obtaining module 2001, an analyzing module 2002, a solving module 2003 and a decompressing module 2004. The second obtaining module 2001 is configured to obtain compressed data, the compressed data being obtained by compressing to-be-compressed data based on the data compression method described in the first aspect or any possible implementation manner of the first aspect. The analyzing module 2002 is configured to analyze the compressed data to obtain an analysis result, the analysis result including parameters of a mechanism model and a boundary term, the mechanism model describing a mechanism of the to-be-compressed data, the mechanism model including a partial differential equation, the parameters including a background field or a support set of the background field of the partial differential equation, and the boundary term indicating a distribution boundary of the to-be-compressed data. The solving module 2003 is configured to solve the partial differential equation based on the background field or the support set of the background field and the boundary term to obtain a solution of the partial differential equation. The decompressing module 2004 is configured to obtain decompressed data based on the solution of the partial differential equation.
[0234] In a possible implementation, the analysis result further includes a residual error, the residual error indicating a difference between the solution of the partial differential equation and the to-be-compressed data. The decompressing module 2004 is specifically configured to obtain the decompressed data based on the solution of the partial differential equation and the residual error. In this way, the original data before compression is reconstructed losslessly through the solution of the partial differential equation and the residual error.
[0235] The data decompression apparatus 2000 according to the embodiments of the present application can correspond to performing the data decompression method described in the embodiments of the present application, and the above and other operations and / or functions of each module in the data decompression apparatus 2000 are respectively for realizing the corresponding flow of each method in FIGS. 9-10, which will not be described herein again for brevity.
[0236] The embodiments of the present application further provide a computing device including at least one processor, a memory and a communication interface, the processor being configured to execute the method described in FIGS. 2-10.
[0237] FIG. 21 is a structural schematic diagram of a computing device according to an embodiment of the present application.
[0238] As shown in FIG. 21, the computing device 2100 includes at least one processor 2101, a memory 2102, and a communication interface 2103. Among them, the processor 2101, the memory 2102 and the communication interface 2103 are communicatively connected, which can be realized by wired (for example, bus) communication connection or wireless communication connection. The communication interface 2103 is used to send and / or receive data sent by other devices; the memory 2102 stores computer instructions, and the processor 2101 executes the computer instructions to execute the data compression method in the foregoing method embodiments, to compress data with high quality, reduce compression rate, reduce storage overhead, and execute the data decompression method in the foregoing method embodiments, to decompress the compressed data compressed by the data compression method provided in the embodiments of the present application, and reconstruct the original data.
[0239] It should be understood that, in the embodiments of the present application, the processor 2101 can be a central processing unit CPU, and the processor 1801 can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0240] The memory 2102 can include read-only memory and random access memory, and provide instructions and data for the processor 2101. The memory 2102 can also include non-volatile random access memory. Optionally, the random access memory can be a high bandwidth memory (HBM).
[0241] The memory 2102 can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. Where the nonvolatile memory is a read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), or flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example, and not limitation, many forms of RAM are available, for example, static RAM (SRAM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), double-data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0242] It should be understood that the computing device 2100 according to the embodiments of the present application can execute the method shown in FIGS. 2-10 of the embodiments of the present application, detailed description of which is referred to the above, and for brevity, will not be repeated here.
[0243] The embodiments of the present application provide a computer readable storage medium, which stores a computer program, when the computer program is executed by a processor, the above-mentioned method is implemented.
[0244] The embodiments of the present application provide a chip, which includes at least one processor and an interface, the at least one processor determines program instructions or data through the interface; the at least one processor is used to execute the program instructions to implement the above-mentioned method.
[0245] The embodiments of the present application provide a computer program or computer program product, which includes instructions, when the instructions are executed, the computer executes the above-mentioned method.
[0246] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0247] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented using hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0248] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A data compression method characterized by, The method comprises: obtaining to-be-compressed data, the to-be-compressed data having mechanism characteristics; determining a mechanism model and a boundary term, the mechanism model describing a mechanism of the to-be-compressed data, the mechanism model comprising a partial differential equation, the boundary term indicating a distribution boundary of the to-be-compressed data; inverting the mechanism model based on the to-be-compressed data to obtain parameters of the mechanism model, the parameters comprising a background field of the partial differential equation or a support set of the background field; compressing the to-be-compressed data based on the background field or the support set of the background field and the boundary term.
2. The method of claim 1, wherein, The inverting the mechanism model based on the to-be-compressed data to obtain parameters of the mechanism model comprises: initializing the background field based on an evolution trend of the to-be-compressed data to obtain an initial value of the background field; solving the partial differential equation based on the boundary term and the initial value of the background field to obtain a solution of the partial differential equation; determining a first residual based on the solution of the partial differential equation and the to-be-compressed data; optimizing and adjusting the background field of the partial differential equation to obtain the background field by taking minimizing the first residual as an objective.
3. The method of claim 2, wherein, The solving the partial differential equation based on the boundary term and the initial value of the background field to obtain a solution of the partial differential equation comprises: regularizing the background field, a parameter of the regularizing being determined based on noise of the to-be-compressed data; solving the partial differential equation based on the boundary term and the initial value of the regularized background field to obtain a solution of the partial differential equation.
4. The method of claim 1, wherein, The inverting the mechanism model based on the to-be-compressed data to obtain parameters of the mechanism model comprises: initializing the support set of the background field based on an evolution trend of the to-be-compressed data to obtain an initial value of the support set of the background field; solving the partial differential equation based on the boundary term and the initial value of the support set of the background field to obtain a solution of the partial differential equation; determining a second residual based on the solution of the partial differential equation and the to-be-compressed data; adjusting the support set of the background field of the partial differential equation to obtain the support set of the background field by taking minimizing the second residual as an objective.
5. The method of claim 4, wherein, The support set of the background field comprises a boundary term and a boundary value parameter of the background field, the boundary term of the background field indicating a distribution boundary of the background field, the boundary term and the boundary value parameter of the background field being used to determine the background field; The second residual is calculated based on an objective function, the objective function comprising a residual term and a regularization term, a value of the residual term being determined based on the to-be-compressed data and the solution of the partial differential equation, a value of the regularization term being determined based on a regularization parameter and the boundary value parameter of the background field; The adjusting the support set of the background field of the partial differential equation to obtain the support set of the background field by taking minimizing the second residual as an objective comprises: optimizing and adjusting the boundary term of the background field and the boundary value parameter by taking minimizing a function value of the objective function as an objective; obtaining the support set of the background field based on the boundary term of the background field and the boundary value parameter after the optimization and adjustment.
6. The method according to any one of claims 2-5, characterized in that, The compression of the to-be-compressed data is also related to a target residual, the target residual indicating a difference between a solution of a target partial differential equation and the to-be-compressed data, the solution of the target partial differential equation being obtained by solving the partial differential equation based on the background field or the branch set of the background field and the boundary term.
7. The method of claim 6, wherein, The compression of the to-be-compressed data based on the background field or the branch set of the background field and the boundary term further comprises: checking the target residual; if the mean value of the target residual is less than a preset threshold, the checking passes.
8. The method of claim 1, wherein, The parameters of the mechanism model further comprise a source term of the partial differential equation; The inversion of the mechanism model based on the to-be-compressed data to obtain the parameters of the mechanism model comprises: initializing the background field based on the evolution trend of the to-be-compressed data to obtain an initial value of the background field; solving the partial differential equation based on the initial value of the background field and the boundary term to obtain a solution of the partial differential equation; determining the source term based on the solution of the partial differential equation and the to-be-compressed data; optimizing and adjusting the background field of the partial differential equation to obtain the background field, with the goal of minimizing the source term.
9. The method of claim 8, wherein, The compression of the to-be-compressed data is also related to a target residual, the target residual indicating a difference between a solution of a target partial differential equation and the to-be-compressed data, the solution of the target partial differential equation being obtained by solving the partial differential equation based on the background field or the branch set of the background field and the boundary term.
10. The method according to any one of claims 1 to 9, characterized in that, The accuracy of the parameters of the mechanism model is determined based on the accuracy of the to-be-compressed data.
11. The method according to any one of claims 1 to 10, characterized in that, The obtaining of the to-be-compressed data further comprises: blocking the to-be-compressed data to obtain a plurality of to-be-compressed data blocks.
12. The method according to any one of claims 1 to 11, characterized in that, The boundary term comprises an initial value and / or a boundary value, the initial value indicating a data distribution of an initial time of the to-be-compressed data, and the boundary value indicating a distribution boundary of the to-be-compressed data in space.
13. A method of data decompression, characterized by, comprises: obtaining compressed data, the compressed data being obtained by compressing to-be-compressed data based on the method of any one of claims 1-12; analyzing the compressed data to obtain an analysis result, the analysis result comprising parameters of a mechanism model and a boundary term, the mechanism model describing a mechanism of the to-be-compressed data, the mechanism model comprising a partial differential equation, the parameters comprising a background field or a branch set of the background field of the partial differential equation, and the boundary term indicating a distribution boundary of the to-be-compressed data; solving the partial differential equation based on the background field or the branch set of the background field and the boundary term to obtain a solution of the partial differential equation; obtaining decompressed data based on the solution of the partial differential equation.
14. The method of claim 13, wherein, The analysis result further comprises a residual, the residual indicating a difference between the solution of the partial differential equation and the to-be-compressed data; obtaining decompressed data based on the solution of the partial differential equation, comprising: obtaining the decompressed data based on the solution of the partial differential equation and the residual.
15. A data compression device, characterized by comprises: a first obtaining module, configured to obtain to-be-compressed data, the to-be-compressed data having mechanism characteristics; determining a mechanism model describing a mechanism of the data to be compressed, the mechanism model comprising a partial differential equation, and a boundary term indicating a distribution boundary of the data to be compressed; inverting the mechanism model based on the data to be compressed to obtain a parameter of the mechanism model, the parameter comprising a background field of the partial differential equation or a branch set of the background field; compressing the data to be compressed based on the background field or the branch set of the background field and the boundary term.
16. A data compression device, characterized by comprising: a second obtaining module configured to obtain compressed data, the compressed data being obtained by compressing data to be compressed based on the method in any one of claims 1-12; a parsing module configured to parse the compressed data to obtain a parsed result, the parsed result comprising a parameter of a mechanism model and a boundary term, the mechanism model describing a mechanism of the data to be compressed, the mechanism model comprising a partial differential equation, the parameter comprising a background field of the partial differential equation or a branch set of the background field, and the boundary term indicating a distribution boundary of the data to be compressed; a solving module configured to solve the partial differential equation based on the background field or the branch set of the background field and the boundary term to obtain a solution of the partial differential equation; a decompressing module configured to obtain decompressed data based on the solution of the partial differential equation.
17. A computing device comprising a memory and a processor, wherein: instructions stored in the memory, when executed by the processor, cause the method in any one of claims 1-14 to be implemented.
18. A computer readable storage medium having stored thereon a computer program, characterized in that, the computer program, when executed by the processor, causes the method in any one of claims 1-14 to be implemented.
Citation Information
Patent Citations
Partial differential equation data processing method and system, storage medium, equipment and application
CN112784205A
Flow field mixing calculation method based on physical neural network
CN116796648A
Metasurface spectral response prediction method based on physical-data cooperative driving
CN116955958A
Multi-physics field time domain non-intrusive model reduction method and system based on data driving
CN117473826A
System and Method for Training of neural Network Model for Control of High Dimensional Physical Systems
US20240152748A1