Data compression and decompression method and device

CN121508541APending Publication Date: 2026-02-10HUAWEI TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411076424.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

High-dimensional data storage for high-performance computing is costly, general data compression algorithms have poor compression ratios, and existing methods are unable to effectively tap the compression potential of mechanistic data.

Method used

By constructing a mechanism model based on partial differential equations, low-dimensional boundary terms and background fields are used to compress the data to be compressed. The background field of the partial differential equations is optimized and adjusted to reduce the residuals, thereby achieving effective data compression.

Benefits of technology

It reduces the compression ratio of data compression, reduces storage overhead, and can quickly decompress and reconstruct the original data distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121508541A_ABST
    Figure CN121508541A_ABST
Patent Text Reader

Abstract

The invention provides a data compression method and device and a data decompression method and device, and the data compression method comprises the steps: obtaining to-be-compressed data which has mechanism characteristics; a mechanism model and a boundary item are determined, the mechanism model describes the mechanism of the to-be-compressed data, the mechanism model comprises a partial differential equation, and the boundary item indicates the distribution boundary of the to-be-compressed data; inversion is carried out on the mechanism model based on the to-be-compressed data, parameters of the mechanism model are obtained, and the parameters comprise a background field of the partial differential equation or a branch set of the background field; and compressing the to-be-compressed data based on the background field or the support of the background field and the boundary term. According to the method, the model parameters of the mechanism model are obtained through inversion of the to-be-compressed data, the mechanism model can describe the mechanism of the to-be-compressed data, the to-be-compressed data are compressed by using the low-dimensional boundary item and the model parameters of the mechanism model, the compression rate of data compression is effectively reduced, and the storage overhead is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to a data compression and decompression method and device. BACKGROUND

[0002] High performance computing (HPC) data is an important part of unstructured data storage, and a large amount of time evolution data (space N dimensions + time dimensions) of various scalar fields is generated in scientific computing and physical field simulation processes. High-dimensional data storage costs are high, and general data compression algorithms have poor compression ratios. SUMMARY

[0003] Embodiments of the present application provide a data compression and decompression method and device, which can effectively reduce the compression ratio of data compression and reduce storage costs.

[0004] In a first aspect, the present application provides a data compression method, which comprises obtaining to-be-compressed data, the to-be-compressed data having a mechanism characteristic; determining a mechanism model and a boundary term, the mechanism model describing a mechanism of the to-be-compressed data, the mechanism model comprising a partial differential equation, and the boundary term indicating a distribution boundary of the to-be-compressed data; performing inversion on the mechanism model based on the to-be-compressed data to obtain a parameter of the mechanism model, the parameter comprising a background field of the partial differential equation or a support set of the background field; and compressing the to-be-compressed data based on the background field or the support set of the background field and the boundary term.

[0005] The present application obtains model parameters of a mechanism model constructed based on a partial differential equation (PDE) by inversion of to-be-compressed data, the mechanism model can describe a mechanism of the to-be-compressed data, and the to-be-compressed data is compressed by using a low-dimensional boundary term and the model parameters of the mechanism model, thereby effectively reducing the compression ratio of data compression and reducing storage costs.

[0006] In one possible implementation, a specific implementation of obtaining the parameter of the mechanism model based on the to-be-compressed data is that: initializing the background field based on an evolution trend of the to-be-compressed data to obtain an initial value of the background field; solving the partial differential equation based on the boundary term and the initial value of the background field to obtain a solution of the partial differential equation; determining a first residual based on the solution of the partial differential equation and the to-be-compressed data; and optimizing and adjusting the background field of the partial differential equation to obtain the background field, with the objective of minimizing the first residual.

[0007] By minimizing the residuals to optimize the background field of the partial differential equations, the partial differential equations can better approximate the mechanism by which the data to be compressed is described. The residuals are sparsity and have a small magnitude (i.e., low storage bits), resulting in low information entropy and improved compression performance.

[0008] In another possible implementation, a specific approach to solving the partial differential equation based on the boundary terms and the initial values ​​of the background field is as follows: the background field is regularized, with the regularization parameters determined based on the noise in the data to be compressed; the partial differential equation is then solved based on the boundary terms and the initial values ​​of the regularized background field to obtain the solution.

[0009] By regularizing the background field, the stability of the inversion process is ensured.

[0010] In another possible implementation, a specific method for inverting the mechanism model based on the data to be compressed to obtain the parameters of the mechanism model is as follows: Initialize the background field support based on the evolution trend of the data to be compressed to obtain initial values ​​for the background field support; solve the partial differential equation based on the boundary terms and the initial values ​​of the background field support to obtain the solution of the partial differential equation; determine the second residual based on the solution of the partial differential equation and the data to be compressed; adjust the background field support of the partial differential equation with the goal of minimizing the second residual to obtain the support of the background field.

[0011] Compression is achieved by using background field supports with lower dimensionality, further reducing the compression ratio of the data.

[0012] In another possible implementation, the support of the background field includes boundary terms and boundary value parameters of the background field. The boundary terms indicate the distribution boundary of the background field, and the boundary terms and boundary value parameters are used to determine the background field. The second residual is calculated based on the objective function, which includes residual terms and regularization terms. The value of the residual terms is determined based on the data to be compressed and the solution of the partial differential equation, and the value of the regularization terms is determined based on the regularization parameter and the boundary value parameters of the background field. A specific implementation of adjusting the support of the background field of the partial differential equation with the goal of minimizing the second residual is as follows: minimizing the function value of the objective function is the objective of optimizing and adjusting the boundary terms and boundary value parameters of the background field; the support of the background field is obtained based on the optimized and adjusted boundary terms and boundary value parameters of the background field.

[0013] In another possible implementation, the compression of the data to be compressed is also related to the target residual, which indicates the difference between the solution of the target partial differential equation and the data to be compressed. The solution of the target partial differential equation is obtained by solving the partial differential equation based on the boundary terms and the optimized background field.

[0014] In this possible implementation, lossless compression of data is achieved by preserving residuals during the compression process, thus ensuring the accuracy of data compression.

[0015] In another possible implementation, the data to be compressed is compressed based on the background field or the support of the background field and the boundary terms, and before that, the target residual is checked; if the mean of the target residual is less than a preset threshold, the check passes.

[0016] By examining the residuals, it is determined whether the inverted mechanistic model can well characterize the mechanism of the data to be compressed, so as to ensure the compression effect; otherwise, other compression methods are used for compression.

[0017] In another possible implementation, the parameters of the mechanism model also include the source terms of the partial differential equation. A specific implementation of inverting the mechanism model based on the data to be compressed to obtain the parameters of the mechanism model is as follows: the background field is initialized based on the evolution trend of the data to be compressed to obtain the initial value of the background field; the partial differential equation is solved based on the boundary terms and the initial value of the background field to obtain the solution of the partial differential equation; the source terms are determined based on the solution of the partial differential equation and the data to be compressed; the background field of the partial differential equation is optimized and adjusted with the goal of minimizing the source terms to obtain the background field.

[0018] Optimizing the background field of partial differential equations using source terms results in lower computational overhead and higher efficiency.

[0019] In another possible implementation, the compression of the data to be compressed is also related to the target source term, which indicates the difference between the solution of the target partial differential equation and the data to be compressed. The solution of the target partial differential equation is obtained by solving the partial differential equation based on the boundary terms and the optimized background field.

[0020] In another possible implementation, the accuracy of the parameters of the mechanistic model is determined based on the accuracy of the data to be compressed.

[0021] In another possible implementation, the process of acquiring the data to be compressed further includes: dividing the data to be compressed into blocks to obtain multiple data blocks to be compressed. Optionally, the data to be compressed can be divided into several data blocks by taking into account hardware processing capabilities and compression parameters, and then compression can be performed on each data block.

[0022] In another possible implementation, the boundary terms include initial values ​​and / or boundary values, where the initial values ​​indicate the data distribution of the data to be compressed at the initial moment, and the boundary values ​​indicate the spatial distribution boundaries of the data to be compressed.

[0023] Secondly, this application also provides a data decompression method, comprising: acquiring compressed data, wherein the compressed data is obtained by compressing the data to be compressed based on the data compression method described in the first aspect or any possible implementation thereof; parsing the compressed data to obtain parsing results, wherein the parsing results include parameters and boundary terms of a mechanistic model, the mechanistic model describing the mechanism of the data to be compressed, the mechanistic model including partial differential equations, the parameters including the background field or the support of the background field of the partial differential equations, and the boundary terms indicating the distribution boundary of the data to be compressed; solving the partial differential equations based on the background field or the support of the background field, and the boundary terms to obtain the solution of the partial differential equations; and obtaining decompressed data based on the solution of the partial differential equations.

[0024] The decompression method provided in this application can be used to quickly decompress and reconstruct the original data distribution (i.e., the data distribution before compression).

[0025] In one possible implementation, the analytical result also includes residuals, which indicate the difference between the solution to the partial differential equation and the data to be compressed. A specific implementation of obtaining decompressed data based on the solution to the partial differential equation is as follows: Decompressed data is obtained based on the solution to the partial differential equation and the residuals. Thus, the uncompressed original data is obtained losslessly through the solution to the partial differential equation and the residuals.

[0026] Thirdly, this application also provides a data compression device, including a first acquisition module, a determination module, an inversion module, and a compression module. The first acquisition module acquires data to be compressed, which has mechanistic characteristics. The determination module determines a mechanistic model and boundary terms. The mechanistic model describes the mechanism of the data to be compressed and includes partial differential equations. The boundary terms indicate the distribution boundaries of the data to be compressed. The inversion module inverts the mechanistic model based on the data to be compressed to obtain parameters of the mechanistic model, including the background field of the partial differential equations or the supports of the background field. The compression module compresses the data to be compressed based on the background field or the supports of the background field and the boundary terms.

[0027] In one possible implementation, the inversion module is specifically used to: initialize the background field based on the evolution trend of the data to be compressed, and obtain the initial value of the background field; solve the partial differential equation based on the boundary terms and the initial value of the background field, and obtain the solution of the partial differential equation; determine the first residual based on the solution of the partial differential equation and the data to be compressed; and optimize and adjust the background field of the partial differential equation with the goal of minimizing the first residual, and obtain the background field.

[0028] In another possible implementation, a specific approach to solving the partial differential equation based on the boundary terms and the initial values ​​of the background field is as follows: the background field is regularized, with the regularization parameters determined based on the noise in the data to be compressed; the partial differential equation is then solved based on the boundary terms and the initial values ​​of the regularized background field to obtain the solution.

[0029] In another possible implementation, the inversion module is specifically used to: initialize the background field support based on the evolution trend of the data to be compressed, and obtain the initial value of the background field support; solve the partial differential equation based on the boundary terms and the initial value of the background field support, and obtain the solution of the partial differential equation; determine the second residual based on the solution of the partial differential equation and the data to be compressed; and adjust the support of the background field of the partial differential equation with the goal of minimizing the second residual, and obtain the support of the background field.

[0030] In another possible implementation, the support of the background field includes boundary terms and boundary value parameters of the background field. The boundary terms indicate the distribution boundary of the background field, and the boundary terms and boundary value parameters are used to determine the background field. The second residual is calculated based on the objective function, which includes residual terms and regularization terms. The value of the residual terms is determined based on the data to be compressed and the solution of the partial differential equation, and the value of the regularization terms is determined based on the regularization parameter and the boundary value parameters of the background field. A specific implementation of adjusting the support of the background field of the partial differential equation with the goal of minimizing the second residual is as follows: minimizing the function value of the objective function is the objective of optimizing and adjusting the boundary terms and boundary value parameters of the background field; the support of the background field is obtained based on the optimized and adjusted boundary terms and boundary value parameters of the background field.

[0031] In another possible implementation, the compression of the data to be compressed is also related to the target residual, which indicates the difference between the solution of the target partial differential equation and the data to be compressed. The solution of the target partial differential equation is obtained by solving the partial differential equation based on the boundary terms and the optimized background field.

[0032] In another possible implementation, the data compression apparatus provided in this application further includes a verification module for verifying the target residual; if the mean of the target residual is less than a preset threshold, the verification passes.

[0033] In another possible implementation, the parameters of the mechanistic model also include the source terms of the partial differential equation; the inversion module is specifically used to: initialize the background field based on the evolution trend of the data to be compressed, and obtain the initial value of the background field; solve the partial differential equation based on the boundary terms and the initial value of the background field, and obtain the solution of the partial differential equation; determine the source terms based on the solution of the partial differential equation and the data to be compressed; optimize and adjust the background field of the partial differential equation with the goal of minimizing the source terms, and obtain the background field.

[0034] In another possible implementation, the compression of the data to be compressed is also related to the target source term, which indicates the difference between the solution of the target partial differential equation and the data to be compressed. The solution of the target partial differential equation is obtained by solving the partial differential equation based on the boundary terms and the optimized background field.

[0035] In another possible implementation, the accuracy of the parameters of the mechanistic model is determined based on the accuracy of the data to be compressed.

[0036] In another possible implementation, the data compression apparatus provided in this application further includes a block segmentation module, which is used to segment the data to be compressed into multiple data blocks.

[0037] In another possible implementation, the boundary terms include initial values ​​and / or boundary values, where the initial values ​​indicate the data distribution of the data to be compressed at the initial moment, and the boundary values ​​indicate the spatial distribution boundaries of the data to be compressed.

[0038] Fourthly, this application also provides a decompression apparatus, including a second acquisition module, an analysis module, a solution module, and a decompression module. The second acquisition module acquires compressed data, which is obtained by compressing the data to be compressed based on the data compression method described in the first aspect or any possible implementation of the first aspect. The analysis module analyzes the compressed data to obtain analysis results, which include parameters and boundary terms of a mechanistic model. The mechanistic model describes the mechanism of the data to be compressed and includes partial differential equations. The parameters include the background field or the support of the background field of the partial differential equations, and the boundary terms indicate the distribution boundaries of the data to be compressed. The solution module solves the partial differential equations based on the background field or the support of the background field, and the boundary terms to obtain the solution to the partial differential equations. The decompression module obtains decompressed data based on the solution to the partial differential equations.

[0039] In one possible implementation, the analytical result also includes residuals, which indicate the difference between the solution to the partial differential equation and the data to be compressed. The decompression module is specifically used to obtain decompressed data based on the solution to the partial differential equation and the residuals. Thus, the uncompressed original data is obtained losslessly through the solution to the partial differential equation and the residuals.

[0040] Fifthly, embodiments of this application provide a computing device, including a memory and a processor, wherein the memory stores instructions that, when executed by the processor, cause the data compression method described in the first aspect or any possible implementation of the first aspect, and / or the data decompression method described in the first aspect or any possible implementation of the first aspect to be implemented.

[0041] In a sixth aspect, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the data compression method described in the first aspect or any possible implementation thereof, and / or the data decompression method described in the first aspect or any possible implementation thereof, to be implemented.

[0042] In a seventh aspect, embodiments of this application also provide a computer program or computer program product, the computer program or computer program product including instructions that, when executed, cause a computer to perform the data compression method described in the first aspect or any possible implementation of the first aspect, and / or the data decompression method described in the first aspect or any possible implementation of the first aspect.

[0043] Eighthly, embodiments of this application also provide a chip including at least one processor and a communication interface, the processor being configured to execute the data compression method described in the first aspect or any possible implementation of the first aspect, and / or the data decompression method described in the first aspect or any possible implementation of the first aspect. Attached Figure Description

[0044] Figure 1 This diagram illustrates a dimensionality reduction approach for mechanistic data.

[0045] Figure 2 A schematic diagram illustrating the principle of the data compression algorithm provided in an embodiment of this application is shown;

[0046] Figure 3 A schematic diagram illustrating the principle of data compression based on background field inversion is shown.

[0047] Figure 4 This illustration shows a system architecture diagram of which the data compression method provided in this application embodiment can be applied;

[0048] Figure 5 This illustration shows an application scenario diagram of the data compression method provided in an embodiment of this application;

[0049] Figure 6 A schematic flowchart illustrating the data compression method provided in this application embodiment;

[0050] Figure 7 The relationships between the variables after dimensionality reduction are shown;

[0051] Figure 8 The diagram illustrates the compression process of a data compression method provided in an embodiment of this application.

[0052] Figure 9 A schematic flowchart of a data decompression method provided in an embodiment of this application;

[0053] Figure 10 This illustration shows a schematic diagram of the compression process of a data decompression method provided in an embodiment of this application;

[0054] Figure 11 The diagram shows the distribution of solutions to the passive transport equation and the distribution of streamlines in the background field.

[0055] Figure 12 A schematic diagram of the original data distribution and the reconstructed data distribution is shown;

[0056] Figure 13 A schematic diagram of the background field reconstruction results is shown;

[0057] Figure 14 A schematic diagram of the initial field distribution and background field distribution for translational transport is shown;

[0058] Figure 15 This is a schematic diagram of the reconstruction results for the translational transport scenario;

[0059] Figure 16 This is a schematic diagram showing the data distribution at the beginning and end of the original data.

[0060] Figure 17 This is a schematic diagram of the data distribution after reconstruction.

[0061] Figure 18 A schematic diagram summarizing and comparing the effects of the data compression methods provided for the implementation of this application;

[0062] Figure 19 This is a schematic diagram of the structure of a data compression device provided in an embodiment of this application;

[0063] Figure 20 This is a schematic diagram of the structure of a data decompression device provided in an embodiment of this application;

[0064] Figure 21 A schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation

[0065] The term "and / or" used in this article describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. The symbol " / " in this article indicates that the related objects are in an "or" relationship; for example, A / B means A or B.

[0066] The terms "first" and "second," etc., used in the specification and claims herein are used to distinguish different objects, not to describe a specific order of objects. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same properties in the description of embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such processes, methods, systems, products, or apparatus.

[0067] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0068] In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more, for example, multiple processing units means two or more processing units, multiple elements means two or more elements, etc.

[0069] Data compression methods in related technologies all have various problems. For example, traditional high-dimensional data compression methods such as filtering perform basis transformation on the entire data to remove spatial redundancy, and utilize the sparsity of frequency domain coefficients to divide the data into different bit planes for encoding, thereby saving storage space in existing scientific computing data compression schemes.

[0070] This method utilizes the idea of ​​dimensionality reduction and has good versatility, but it does not combine the inherent mechanism of the data (physical model), so it fails to fully realize the compression potential for mechanistic data compression.

[0071] The second related technology is data compression methods based on prediction and entropy coding. Common predictors include the Lorenzo predictor and interpolation predictor. The principle is to achieve data dimensionality reduction through neighbor prediction. When predicting a sample point, the Lorenzo predictor estimates the scalar value at that point using its processed neighbor points, assuming the data is stored in a conventional grid format. Both the compressor and decompressor traverse the data points in a linear order. For subsequent compression, this method only needs to compress and store the boundary points and the errors at each point.

[0072] For predictor-based methods, the use of fixed empirical models has limited applicability to compress scientific computing data in complex scenarios. While compression is effective for data conforming to one type of mechanism, it may result in significant model errors and limited compression for data conforming to another type of mechanism, due to the model's inability to accurately represent the physical laws. Therefore, there is considerable room for improvement. Furthermore, the characteristics of scientific computing data in complex scenarios often vary considerably, requiring the design of specific models tailored to these data characteristics to improve compression rates.

[0073] Furthermore, traditional predictors are often designed for two-dimensional image data. Designing a predictor for high-dimensional data is a challenge. Moreover, for high-dimensional spatiotemporal mechanism data, general predictors are unable to reflect the differences and correlations between spatiotemporal mechanisms. In addition, the well-stability of the computation process in high-dimensional cases is not trivial.

[0074] The third related technology is learning algorithms based on neural networks. Data-driven modeling based on deep neural networks (DNNs) and corresponding machine learning methods approximate data using neural networks, optimizing and storing network parameters and data residuals. The resulting model is in neural network form, fundamentally different from models in the form of differential equations or other governing equations. Furthermore, there is model learning and data compression based on physics-informed neural networks (PINNs). Essentially, this also approximates data using neural networks, with the model acting as a constraint but ultimately representing the data in network form.

[0075] For high-dimensional data models, the parameter scale and training overhead in neural networks increase accordingly, making it difficult to make good predictions (interpretations and generalizations). Due to the complexity and nonlinearity of various neural network models, their interpretability is weak, making theoretical analysis difficult, including the estimation and analysis of compression effects. These methods focus on general data approximation and do not involve physical laws governing equations. Therefore, for data scenarios with mechanistic aspects, they often cannot well reflect the differences in data characteristics and extract the inherent data supports. When the approximation degree is high, the model parameter storage cost is large, while when the approximation degree is low, the model bias (residual) storage volume is large, thus limiting the overall compression effect. At the same time, the compression and decompression computational complexity is high, and the computational overhead of typical deep learning methods is often very large, and the applicability of the black-box model trained is limited.

[0076] In data storage, a large amount of data possesses specific physical mechanisms due to its generation and acquisition mechanisms. HPC scientific computing data, such as numerical simulation data of various physical fields like turbulence, atmosphere, and ocean, correspond to well-defined model mechanisms (governing equations); various observational data, especially geophysical satellite remote sensing data, such as sea surface height, temperature, and atmospheric temperature and pressure, evolve in accordance with various conservation laws and also contain underlying mechanisms; other industrial process data, such as concentration and density, also possess specific mechanisms if they conform to inherent laws such as continuity and conservation.

[0077] For the aforementioned mechanistic data, if the mechanism can be characterized by a model, the compression ratio can be significantly improved by extracting the main mechanism. Mechanistic data can usually be characterized by differential equations. For high-dimensional data, this is a partial differential equation or a system of equations (PDE). The solution of a PDE is determined by appropriate initial and boundary conditions and source terms. Therefore, the entire physical field (model data) is determined by low-dimensional well-posed data, such as initial and boundary value data (data supports). See [link to relevant documentation]. Figure 1 Therefore, based on the idea of ​​PDE dimensionality reduction, it is hoped that effective data compression can be achieved.

[0078] For two-dimensional data, traditional methods such as the Lorenzo predictor can characterize the data mechanism and achieve mechanism compression. However, for high-dimensional spatiotemporal data, the data has multiple dimensions, with continuous changes in the time dimension and significant differences in the spatial dimension, but still exhibiting spatiotemporal correlations. Therefore, the ideas of traditional predictors are difficult to directly extend to high-dimensional cases.

[0079] In summary, compared to general compression methods based on empirical models, data compression methods based on mechanistic models have advantages for scientific computing data. Furthermore, with targeted model design, computational overhead can be reduced, facilitating theoretical analysis. Moreover, introducing model degrees of freedom and reconstructing specific models based on data can further optimize compression performance and generalization by making the data more closely match the underlying mechanisms. From a data storage perspective, we select suitable models to better reflect physical laws and reduce the storage overhead of model errors, while also selectively employing a class of widely applicable models. This means that while using PDEs for dimensionality reduction, we also ensure the stability and robustness of the aforementioned process.

[0080] The specific implementation of the data compression method and apparatus provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0081] General data compression algorithms and image compression algorithms in related technologies suffer from low data specificity, low compression ratio, and low performance when used for scientific computing data, high-precision remote sensing data, and other data with clearly defined mechanisms. These problems are even more pronounced for high-dimensional spatiotemporal data. Therefore, we need to provide a compression algorithm that is specifically designed for the characteristics of mechanistic data, fully extracts the data mechanism, reflects the differences and correlations in spatiotemporal dimensions, and achieves high compression ratio and high performance with a simple and easy-to-calculate model, low model bias, and few data supports. The technical problems to be solved include:

[0082] 1. Compression ratio requirements

[0083] By leveraging the unique characteristics and mechanisms of mechanistic floating-point data, such as scientific data and high-precision remote sensing data, compared to general data and image data, a dedicated compression algorithm is designed. This algorithm achieves a higher compression ratio than traditional dimensionality reduction compression methods such as filtering, and is more effective at mining the spatiotemporal mechanisms of the data than low-dimensional predictor methods such as the Lorenzo predictor.

[0084] 2. Adaptability and Reliability Requirements

[0085] It can adaptively obtain data-specific mechanistic models based on data, and the resulting models can be reversibly and stably reconstructed when used for data compression, avoiding problems such as low reliability and large computational load caused by complex models.

[0086] For mechanistic data, the data support is often very small, and the spatiotemporal global information can be uniquely determined by low-dimensional information or local information such as initial and boundary values. For such data, if the mechanism (model) can be reconstructed, the compression potential can often be fully explored.

[0087] Can the constructed model reconstruct the mechanism using only a portion of the data, enabling it to predict / extend information in a local approximation sense, and thus possess a generalization ability that is difficult to achieve using compression methods based on neural networks?

[0088] Figure 2 A schematic diagram illustrating the principle of the data compression algorithm provided in an embodiment of this application is shown. Figure 2 As shown, the data compression algorithm provided in this application works by reducing the dimensionality of high-dimensional mechanistic data based on a PDE model to achieve compression. For example, a mechanistic model constructed using partial differential equations describes the distribution mechanism of high-dimensional mechanistic data. By combining data supports (e.g., initial values ​​and boundary values) and other remainder terms (e.g., residuals, source terms, etc.), the distribution of high-dimensional mechanistic data can be determined. The amount of data in the partial differential equations, data supports, and other remainder terms is very small, which enables the compression of high-dimensional mechanistic data. Compared with traditional data compression methods, this greatly reduces the compression ratio and saves storage overhead.

[0089] Therefore, inverse PDEs can be used to effectively compress mechanistic data, and forward PDE calculations can be used to decompress the data.

[0090] For example, linear convection-diffusion partial differential equations are used to characterize the data mechanism for the basic forms of spatiotemporal evolution of physical fields, such as convection and diffusion. The partial differential equations are shown below:

[0091]

[0092] Where (v x (x,y),v y (x,y) represents the background transport field. Data compression is achieved by inversely solving the aforementioned partial differential equation, and data decompression is achieved by forward computation of the same partial differential equation.

[0093] It is easy to understand that mechanistic data refers to data with specific mechanisms, such as HPC scientific computing data, including numerical simulation data of various physical fields such as turbulence, atmosphere, and ocean, which correspond to clear model mechanisms (governing equations); various observational data, especially geophysical satellite remote sensing data, such as sea surface height, temperature, and atmospheric temperature and pressure, whose evolution satisfies various conservation laws, also contain mechanisms; other data in industrial processes, such as concentration and density, which, if they conform to inherent laws such as continuity and conservation, also have their specific mechanisms. Mechanistic data is often high-dimensional spatiotemporal data, which has a large storage overhead and needs to be compressed during storage. The following uses high-dimensional spatiotemporal mechanistic data as an example to introduce the data compression principle provided in the embodiments of this application.

[0094] Figure 3 A schematic diagram illustrating the principle of data compression based on background field inversion is shown. Figure 3 As shown, the spatiotemporal data to be compressed is first matched with the corresponding mechanism equation type. For example, if the spatiotemporal data is data on the spatiotemporal evolution of physical fields, then the corresponding mechanism equation type is a linear convection-diffusion partial differential equation. Then, the background field of the mechanism equation is obtained by inversion. If the residual passes the test, the high-dimensional data is compressed using the data support (i.e., boundary terms) and the background field.

[0095] It should be noted that data supports, also known as well-posed data, are data supports, such as the initial and boundary values ​​of the data, and the mechanistic model, which can determine the distribution of the data to be compressed.

[0096] To further improve the stability of background field inversion, this application also provides an algorithm for stabilizing the coefficients (background field) of the inversion equation based on regularization, that is, adaptively determining the model based on the current data.

[0097] After obtaining the background field through inversion, the data storage is converted into storage of the partial differential equation background field, boundary terms (e.g., initial values ​​and / or boundary values), and low-entropy source terms.

[0098] The data compression method provided in this application can be applied to various compression scenarios. For example, it can be used for compressing unstructured data storage and scientific computing floating-point data storage.

[0099] Figure 4 This diagram illustrates a system architecture to which the data compression method provided in this application can be applied. For example... Figure 4 As shown, the embodiments of this application can be applied to centralized or distributed storage systems as a compression algorithm for deep coding compression, providing a large-scale compression of scientific computing floating-point data.

[0100] Figure 5 This illustration shows an application scenario diagram of the data compression method provided in an embodiment of this application. For example... Figure 5 As shown, the data compression method provided in this application embodiment can be applied to effectively compress mechanistic data in unstructured data, such as scientific computing floating-point data, remote sensing image data, gene sequencing data, energy exploration data, and marine data.

[0101] In unstructured data, scientific computing floating-point data accounts for a large proportion. The data compression method provided in this application embodiment can compress data by a large proportion compared with traditional data compression algorithms, thereby effectively reducing storage costs.

[0102] The data compression method provided in this application can also be applied to compress various high-precision time scalar field data on a large scale, such as geophysical satellite remote sensing data, such as spatiotemporal data of sea surface height, temperature, atmospheric temperature and pressure, and concentration. Its spatiotemporal evolution is governed by various conservation laws, and typically exhibits background field transport and diffusion behaviors, with physical mechanisms and corresponding control equations. At the same time, it is high-precision floating-point data, and traditional image data and other general compression methods are not applicable or not targeted.

[0103] Figure 6This is a flowchart illustrating the data compression method provided in this application embodiment. This method can be executed by any device, equipment, platform, or cluster of devices with computing capabilities. This application embodiment does not specifically limit the specific computing device executing the method; a suitable computing device can be selected as needed. For example, it can be implemented on a terminal device; that is, the data compression method provided in this application embodiment can be deployed as a compressor plug-in in the storage system of the terminal device, executing the data compression method provided in this application embodiment when storing mechanistic data. It can also be implemented on a cloud device, providing mechanistic data compression services to users in the form of cloud services. For ease of description, the form of the executing entity will not be distinguished in the following text; all will be described as a data compressor. Figure 6 As shown, the data compression method provided in this application embodiment includes at least steps S601 to S604.

[0104] In step S601, the data to be compressed is obtained.

[0105] The computing device can receive uncompressed data sent by external devices (e.g., keyboard, camera, voice receiver, external sensors, other computing devices, etc.); or, uncompressed data generated by the computing device running applications.

[0106] The data to be compressed is data with mechanistic characteristics. For example, the data to be compressed can be HPC scientific calculation data, such as numerical simulation data of various physical fields such as turbulence, atmosphere and ocean; various observation data, especially geophysical satellite remote sensing data, such as sea surface height, temperature, atmospheric temperature and pressure, whose evolution satisfies various conservation laws and also contains mechanisms; other industrial process data such as concentration and density, if they conform to the inherent laws of continuity and conservation, also have their specific mechanisms.

[0107] The acquired data is used as the input to the data compressor. At this time, the data to be compressed can also be called the input data. The data compressor obtains the data to be compressed.

[0108] Optionally, after obtaining the input data, the input data is preprocessed. The preprocessing includes block processing, such as dividing the input data into multiple data blocks, and then compressing each data block.

[0109] For example, the data to be compressed is divided into blocks based on hardware processing capabilities and compression parameters, both spatially and temporally. Hardware processing capabilities refer to the amount of data that hardware, such as memory, can support in a single compression process. Compression parameters are based on the amount of data that can be effectively calculated at one time, such as the effective amount of mechanistic data that a mechanistic model can characterize; exceeding this amount will affect accuracy.

[0110] Preprocessing also includes data validity checks and noise reduction.

[0111] In another example, the preprocessing process includes setting coefficients and truncation points for the calculation process based on the precision of the input data (decimal / binary, significant digits / absolute precision). This means that the precision of subsequent calculations is truncated to maintain consistency with the precision of the input data.

[0112] In step S602, the mechanism model and boundary terms are determined.

[0113] In this step, based on the data distribution of the input data, a suitable mechanism model is determined. For example, if the input data is about the spatiotemporal evolution of a physical field, then a linear convection-diffusion partial differential equation can be used as the mechanism model. Or, if the distribution of the input data conforms to passive transport, then a passive transport partial differential equation can be used as the mechanism model.

[0114] The boundary terms are determined, including the initial value of the data to be compressed, i.e., the data distribution of the data block at the earliest time, and / or the boundary value, which refers to the boundary value of the spatial distribution of the data to be compressed. The following uses the boundary terms as the initial value and boundary value as an example (which can also be simply referred to as the initial boundary value) to introduce the scheme of the embodiment of this application.

[0115] In step S603, the mechanism model is inverted based on the data to be compressed to obtain the parameters of the mechanism model.

[0116] First, based on the evolution trend of the input data, the initial value of the background field of the partial differential equation is determined. Then, based on the boundary terms and the initial value of the background field, the partial differential equation is solved to obtain the solution of the partial differential equation. The solution of the partial differential equation and the data to be compressed are matched, and the residual or source term is minimized to optimize and adjust the background field to obtain the background field.

[0117] Therefore, the background field and boundary terms need to be encoded to achieve data compression.

[0118] In another example, to further reduce the compression ratio of the data compression, the initial value of the background field support of the partial differential equation can be determined first based on the evolution trend of the input data. Then, based on the boundary terms and the initial value of the background field support, the partial differential equation is solved to obtain the solution of the partial differential equation. The solution of the partial differential equation is matched with the input data, and the residual or source term is minimized to optimize and adjust the background field support to obtain the background field support.

[0119] It is understandable that the meaning of matching the solution to a partial differential equation with the input data is: the solution to the partial differential equation is spatiotemporally distributed data, and the input data is also spatiotemporally distributed data. The solution to the partial differential equation and the input data in the same spatiotemporal space are said to be matched. For example, the solution u`(x1,y1,t1) of the partial differential equation is matched with u(x1,y1,t1) in the input data.

[0120] In this way, the background field supports and boundary terms, which have smaller data volumes, are then encoded to achieve data compression.

[0121] In another example, after optimizing and obtaining the background field, the mechanism model needs to be verified. The residuals or source terms are calculated to verify the obtained background field and the verification residuals are calculated. If the residuals meet the preset conditions, such as the residual distribution being approximately a Gaussian distribution with a mean of 0, it means that the mechanism model has well described the distribution mechanism of the data to be compressed, and the verification is passed. Otherwise, the verification fails, and other compression methods are used for compression.

[0122] In other examples, in order to achieve lossless compression, after obtaining the background field through inversion, the mechanism model needs to be solved based on the background field and boundary terms. Then, based on the solution of the mechanism model and the corresponding data to be compressed, the residual or source term is calculated. Subsequently, the background field / background field support, residual or source term and boundary term are encoded to achieve lossless compression of the data.

[0123] For example, suppose the spatiotemporal data to be compressed Satisfies the passive transport equation:

[0124]

[0125] Where v x (x,y),v y (x, y) represents the two components of the velocity field (i.e., the background field), which are assumed to be spatially dependent. In practical applications, scalar fields such as temperature and density, and the evolution of dyes and floating pollutants in their corresponding fluids (such as air and water) constitute typical passive transport problems. The data (well-posed conditions) of the above model include initial values ​​and boundary values. The forward problem is: given the initial and boundary conditions, calculate u(x, y, t).

[0126] For actual structural data, Where (i,j) represents the spatial grid and n is the time layer, the difference equation obtained by discretization using finite difference is of the following form:

[0127]

[0128] That is, the data at time n+1 can be determined by the data at the previous time layer. n+1 =L UV u n Furthermore, since the values ​​at all interior points depend only on the values ​​of the outer boundary and the initial layer, in this case, only the boundary, the initial low-dimensional data, and the background field U need to be stored. i,j V i,j All internal data can be calculated directly.

[0129] If both passive transport and diffusion effects are considered, the model is as follows:

[0130]

[0131] In the formula, bc represents the boundary conditions (i.e., boundary values), and ic represents the initial conditions (i.e., initial values).

[0132] In general practical problems, if the data does not precisely satisfy the above ideal transport properties, i.e., the above homogeneous discrete equations, then

[0133] Method 1: Considering the residuals, let the solution to the above homogeneous equation be... The original data is u, stored When the model approximation is good, =u has sparsity and a small magnitude (low storage bits), corresponding to low information entropy, which can achieve compression.

[0134] Method 2: Consider the influence of the source term, i.e.:

[0135]

[0136] In this case, the corresponding difference equation is:

[0137]

[0138] At this point, in addition to storing the aforementioned low-dimensional well-posed data, it is also necessary to store the full-dimensional source term data. However, when the model approximation is good, f often has sparsity or low-value perturbation, corresponding to low information entropy. In this case, even if the source data of all dimensions is stored, the overhead is less than storing the original data, thus achieving the effect of compression.

[0139] Model inversion: As the above principle shows, the key to compression effectiveness lies in whether the model better approximates the data mechanism. The degrees of freedom of the above model are the two-dimensional background field U. i,j V i,j (When considering diffusion, the diffusion coefficient constant ν is also a parameter to be determined in the model). Therefore, we can determine the above background field (model parameters) by 1. prior setting, and 2. optimization (learning) based on data. The latter yields a model with better approximation and thus better compression. The "learning" of the above model is transformed into the following inversion problem:

[0140] Given u(x,t)on[0,t]×Ω, find v. x ,v y And predict information at time [t,T].

[0141] The prediction here refers to the fact that since the background field is independent of time, it can be determined using data from a portion of the time period, thus enabling prediction of data over a longer period and reducing the computational load of the compression process.

[0142] Generally speaking, when the model is assumed to be homogeneous, by You can get information about U i,j V i,j The system of linear algebraic equations, when the time level n≥3, forms an overdetermined system and can be solved using methods such as least squares. However, mathematically, this is an ill-posed problem; data perturbations and noise can make the process of solving the background field unstable, especially u. i,j In regions of minimal change, the aforementioned linear system becomes ill-conditioned, leading to significant biases in model reconstruction and greatly impacting prediction and compression performance. Therefore, this scheme further applies low-dimensional assumptions and regularization to the background field to ensure the stability of the inversion process. Let the background field be a potential flow (irrotational and incompressible):

[0143]

[0144] Then the velocity potential ψ satisfies Δψ=0, according to the properties of harmonic functions. That is, the velocity potential is determined by the single-layer potential density. Decision, among which Ω represents the background field region, and Φ(x,y) represents the fundamental solution of the harmonic function in free space. Furthermore, due to the unique extension of the harmonic function, it is only necessary to obtain the internal local data u(x,t)on[0,t]×ω. This determines the velocity potential ψ, and consequently the background field v. In summary, the background field inversion problem boils down to:

[0145] Given u(x,t)on[0,t]×Ω, find the boundary value μ of the velocity potential.

[0146] Figure 7 The relationship between the variables after dimensionality reduction is shown.

[0147] In specific implementation, for Method 1 (optimizing residuals): let the single-layer potential density of the background field be discretized as... {e i (x)} is Discrete basis functions on the upper surface, and denote θ = {c K}, then ψ| can be determined from θ. G Then determine U i,j V i,j Finally, the model solution is obtained by solving the difference equation in the forward direction. Let θ be the above. The reflection of Where u0, u bord Given the initial and boundary value data, the background field inversion calculation is transformed into solving the following optimization problem:

[0148]

[0149] in

[0150]

[0151] α is the regularization parameter, determined a priori by the noise level of the data itself. θ is then obtained. * Then, calculate the corresponding And calculate the residuals And save θ, u0, u bord This completes the compression. When applied, this method allows for the matching data to be local data because the resulting equation has extrapolative computational capabilities, thus saving algorithmic overhead.

[0152] During decompression, the difference equation is obtained by solving it in the forward direction. The original data is obtained by superimposing the residuals.

[0153] For method 2 (optimizing the source terms), based on the prior estimation of the above equations, a smaller source term modulus can control the modulus of the solution. Therefore, as described above, source terms can be equivalently optimized and stored to achieve compression. Similar to method 1, let... Given the background field determined by the boundary parameter θ, the objective function to be optimized here is:

[0154] F(u;θ)=||f(θ)|| 2 +α||θ|| 2

[0155] Where the source term f(θ)

[0156]

[0157] Storage source and θ,u0,u bord During decompression, the non-homogeneous difference equation is solved to obtain the data.

[0158] For the two optimization approaches mentioned above, Method 1 can directly optimize the residual so that the compression effect is directly reflected in the magnitude of the residual. The calculation of the objective function involves explicitly advancing the time equation. Method 2 involves explicitly calculating the source terms at each time level when calculating the objective function. The overall computational workload of the two approaches is roughly the same, but the latter is simpler to implement.

[0159] In step S604, the data to be compressed is compressed based on the background field or the support of the background field and the boundary terms.

[0160] For the steps above where only the background field is obtained, the background field and boundary terms are compressed and encoded to achieve lossy compression of the input data, resulting in compressed data.

[0161] For the background field support obtained in the above steps, the background field support and boundary terms are compressed and encoded to achieve lossy compression of the input data and obtain compressed data.

[0162] For the background field and residual obtained in the above steps, the background field, boundary terms and residual are compressed and encoded to achieve lossless compression of the input data and obtain compressed data.

[0163] For the background field support and residual obtained in the above steps, the background field support, boundary terms and residual are compressed and encoded to achieve lossless compression of the input data and obtain compressed data.

[0164] For the background field and source term obtained in the above steps, the background field, source term and boundary term are compressed and encoded to achieve lossless compression of the input number and obtain compressed data.

[0165] Figure 8 This diagram illustrates the compression process of a data compression method provided in an embodiment of this application. Figure 8 As shown, the input data is processed sequentially, including initializing the background field, optimizing the velocity field support (boundary potential), calculating and truncating the source terms, calculating and verifying the residuals, and encoding the velocity field, source terms, and boundary terms to obtain compressed data. For example, the boundary terms (i.e., initial boundary conditions), source terms, and velocity field (i.e., background field) are encoded, header information is saved, and written to the compressed file.

[0166] The data compression method provided in this application only requires a single current data to perform background field inversion, without relying on other training data, databases, large models, etc.

[0167] Model-based learning is employed. This involves optimizing the background field data support (low-dimensional) and explicitly solving the model equations, both of which can be parallelized. Due to the model's predictive power, the optimization process requires only a portion of the data, less than the original dataset.

[0168] It can adapt to various types of high-dimensional data with continuous spatiotemporal evolution, and obtain background fields / mechanisms, achieving data compression and obtaining physical mechanisms. The compression effect is outstanding when the data is scientific calculation data such as numerical solutions of differential equations, and it has the best effect for pure convection and thermal equations.

[0169] Lossless compression can be achieved by controlling the accuracy of source terms, or lossy compression with a mechanism can be achieved by making reasonable truncation based on well-posed theory of differential equations and the distribution of data source terms.

[0170] Therefore, the data compression method provided in this application is particularly suitable for storing data with physical mechanisms, i.e., data that satisfies evolution equations, such as atmospheric and oceanic observation data or industrial process data. Since the compression effect depends on the approximation degree of the model, it is not necessary to strictly satisfy the equations. It should be noted that the specific implementation of the above algorithm, including the form of "data support" and the uniqueness and stability of the background field reconstruction, is based on the relevant theory of differential equation inversion problems, and is therefore not trivial.

[0171] In addition to the data compression method provided in the embodiments of this application, the embodiments of this application also provide a data decompression method.

[0172] Figure 9 This is a flowchart illustrating a data decompression method provided in an embodiment of this application. This method can be executed by any device, equipment, platform, or cluster of devices with computing capabilities. This application does not specifically limit the specific computing device executing the method; a suitable computing device can be selected as needed. For example, it can be implemented on a terminal device; that is, the data decompression method provided in this application can be deployed as a decompression plugin in the storage system of the terminal device. For compressed data obtained using the data compression method provided in this application, the data decompression method provided in this application is executed to reconstruct the original data. Alternatively, it can be implemented on a cloud device, providing users with a decompression service for compressed data obtained using the data compression method provided in this application in the form of a cloud service. For ease of description, the form of the executing entity will not be distinguished below; all will be described as a data decompressor. Figure 9 As shown, the data compression method provided in this application embodiment includes at least steps S901 to S904.

[0173] In step S901, compressed data is acquired.

[0174] When it is necessary to decompress the compressed data, a data decompressor is invoked to decompress it and restore the original data. Here, the original data refers to the data before compression, i.e., the data to be compressed or the input data mentioned above. The compressed data here refers to the compressed data obtained by using the data compression method provided in the embodiments of this application.

[0175] Input the compressed data that needs to be decompressed into the data decompressor to obtain the compressed data.

[0176] In step S902, the compressed data is parsed to obtain the parsing results, which include the parameters and boundary terms of the mechanism model.

[0177] The compressed data is parsed to obtain the parameters and boundary terms of the mechanistic model. The content of the parsed results depends on the specific compression scheme used above.

[0178] For example, if compressed data is obtained by compressing and encoding the background field and boundary terms, then the parsing result includes the background field and boundary terms.

[0179] For example, if compressed data is obtained by compressing and encoding the background field supports and boundary terms, then the parsing result includes the background field supports and boundary terms.

[0180] For example, if compressed data is obtained by compressing and encoding the background field, boundary terms, and residuals, then the parsing result includes the background field support, boundary terms, and residuals.

[0181] For example, if compressed data is obtained by compressing and encoding the background field supports, boundary terms, and residuals, then the parsing result includes the background field supports, boundary terms, and residuals.

[0182] For example, if compressed data is obtained by compressing and encoding the background field, boundary terms, and source terms, then the parsing result includes the background field, boundary terms, and source terms.

[0183] For example, if compressed data is obtained by compressing and encoding the background field support, boundary terms, and source terms, then the parsing result includes the background field support, boundary terms, and source terms.

[0184] In step S903, the partial differential equation is solved based on the background field or background field support and boundary terms to obtain the solution of the partial differential equation.

[0185] After obtaining the analytical result, the partial differential equation solver is called to solve the partial differential equation and obtain the solution.

[0186] For example, by substituting the background field into the partial differential equation and solving the partial differential equation based on the boundary terms, the solution to the partial differential equation can be obtained.

[0187] For example, the complete background field is calculated based on the support of the background field. The background field is then substituted into the partial differential equation, and the partial differential equation is solved based on the boundary terms to obtain the solution of the partial differential equation.

[0188] In step S904, decompression data is obtained based on the solution of the partial differential equation.

[0189] For lossy compression, where residuals or source terms are not preserved during compression, the solution of the partial differential equation is directly used as the original data, i.e., the decompressed data.

[0190] For lossless compression, which preserves the residuals or source terms during compression, after obtaining the solution to the partial differential equation, the residuals or source terms are superimposed on the solution to obtain the original data, i.e., the decompressed data.

[0191] Figure 10 This diagram illustrates the compression process of a data decompression method provided in an embodiment of this application. Figure 10 As shown, the compressed data is parsed to obtain the background field support, initial boundary values, and source terms. Then, the background field is calculated based on the background field support. The background field and initial boundary values ​​are substituted into the model solver to obtain the solution of the model (i.e., the solution of the partial differential equation). The residuals are superimposed on the solution of the model to obtain the decompressed data.

[0192] The following specific example illustrates the practical application of the data compression and decompression methods provided in this application.

[0193] The compression process includes the following steps:

[0194] Step 1. Preprocess the input data. Divide the data into blocks, calculate the mean and offset based on hardware processing capabilities and compression parameters. Set coefficients and truncation points in the calculation process according to the data precision (decimal / binary; significant bits / absolute precision). Apply a preset uniform velocity field (v) in four directions. x ,v y ) = (±1,0), (0,±1) and the corresponding velocity potential function ψ(x,y) = ±x, ±y , Calculate and compare the residuals to determine the optimal initial value of the velocity potential function. Alternatively, initial values ​​for the input coefficients can be specified based on other prior tests.

[0195] Step 2. Background Field Inversion Process. Based on the partitioned data, the dimensionality-reduced background field (velocity potential) is optimized. The objective function is:

[0196]

[0197] In this process, i and j traverse the boundary space data points, and n traverses the time layer. The optimization process corresponds to minimizing the objective function. The calculation of the objective function only involves coefficient multiplication and algebraic summation, and explicit advancement only depends on the previous layer, thus the computational overhead is small, making it easy to extend to GPU deployment. The default optimization algorithm is the Levenberg-Marquardt algorithm to balance convergence speed and stability. The initial parameters are selected in step 1, and other optional parameters include the number of iterations, the number of times the objective function is called, and the tolerance. If global optimization is specified in the user settings, the SQP algorithm or other global optimization algorithms can be used, with the objective function remaining unchanged. The discrete components of the optimized velocity potential function, the boundary values ​​of the original block data, and the initial values ​​are retained as data supports.

[0198] Step 3. Calculate, verify, and save the residuals. 1. Based on the background field velocity potential boundary values ​​obtained in Step 2. Calculate the velocity field Then, by substituting the solutions into the transport equations and solving them forward, an approximate solution is obtained. Calculate residuals 1. To assess the effectiveness of the extraction mechanism, if the residual distribution approximates a Gaussian distribution with a mean of 0, it indicates that the mechanism has been extracted effectively. 2. Truncate the residual terms based on data precision. Compare the magnitudes of the residuals and the original data. If the residual magnitude is significantly lower than the original data, the compression effect is good. In lossless compression, the information entropy of the residual terms is also affected by data quality. If the data itself contains significant noise, the compression effect will be limited; if the data is white noise, there will be no compression effect.

[0199] Step 4. Encode and store the data. The data to be stored is the background field support (c) mentioned above. K k = 1, 2…m, i.e., m coefficients), residual term (if the original data is an N1×N2×N3 matrix, then the residual term is (N1-2)×(N2-2)×(N3-1), boundary data i = 0, N1 or j = 0, N2; n = 2, ..., N3, initial data (N1×N2 matrix). Then, the data to be stored is encoded and output.

[0200] The decompression process mainly includes the following specific steps:

[0201] Step 1. After parsing the compressed file, obtain the residual term, initial value, boundary value term, and background field support.

[0202] Step 2. Based on the background field support... Calculate the background field U i,j V i,j .

[0203] Step 3. Substitute into the difference equation Explicit advancement yields approximate solutions The final decompressed data is obtained by superimposing the residuals. The above equations are solved in an explicit format in the forward direction, which results in fast computation speed, low memory overhead, and parallel processing capability.

[0204] In order to verify the compression effect of the data compression method provided in this application, an experiment was conducted on the data compression method provided in the application.

[0205] This application selects several scientific computing data examples to illustrate that the data compression method provided in this application has the best performance when the data itself is a solution to a transport equation. It should be noted that in most cases, although the data satisfies a differential equation, it is not necessarily a linear equation, and the coefficients are not necessarily time-independent. To demonstrate that this algorithm still has good applicability, we examined the following scenarios:

[0206] 1. Numerical solution of the passive transport equation

[0207] We choose the initial distribution u0(x,y)=exp(-16((x-0.5)) 2 +(y-0.5) 2 (x,y)∈[0,1]×[0,1], discretized into 51×51×75 array blocks, with a time step dt=0.005, and a non-reflective boundary condition. The background field distribution is as follows:

[0208]

[0209] r1 = (x - x1) 2 +(y-y1) 2 r² = (x - x²) 2 +(y-y2) 2 ,x1=0,x2=1,y1=-0.1,y2=1.1

[0210] The original data was obtained by explicitly solving the difference equation using the initial and boundary value conditions and the difference equations mentioned above. The data precision is 16 bits, and the data distribution is as follows: Figure 11 As shown in (a), the distribution of streamlines in the background field is as follows: Figure 11 As shown in (b). It should be noted that for larger arrays, on the one hand, they can be divided into blocks and processed in parallel, with the calculation processes between blocks being independent of each other; on the other hand, from the perspective of local linearization, linear approximation in local regions is more effective for cases where the background field changes significantly.

[0211] The original binary file size is 1565K, which becomes 527K after zip encoding, with a compression rate of 33.7%. After applying the data compression method provided in this application, it can be further compressed by 60% to 318K, showing good results. The specific data is summarized in Table 1.

[0212] Example: PDE solving data original compression ratio zip 1565k 527k 33.7% Extraction mechanism + zip 1565k 318k 20.3%

[0213] Table 1: Comparison of Scientific Computing Data Compression Effects

[0214] Background field reconstruction, such as Figure 13As shown in the middle (Figure 1), due to the ill-posedness of the inversion problem itself and the further dimensionality reduction and discretization of the background field, the reconstruction result is relatively ideal. However, for the reconstruction of the final data field, the distributions of the original data and the reconstructed data at the final time step are as follows: Figure 12 As shown, the residual is as small as 10 -3 Since only the residual is actually stored, the actual number of valid bits in the stored data is effectively reduced, thus explaining its compression effect.

[0215] It should be noted that: 1. The above passive transport process is not trivial; compared with simple translation and rotation, it adds a deformation effect. The case of simple translation is a special case of this method and can be covered, such as... Figure 14-15 The results are shown. 2. Model reconstruction (intermediate quantities or parameters such as background field) is allowed to have some deviation, as long as the error of the final data field has a small magnitude, compression can be achieved. 2. By further modeling the background field, a stable reconstruction can be obtained, and the data support can be further reduced. 3. In the embodiments of this application, the reconstruction of the background vector field adopts the idea of ​​dimensionality reduction and the regularization method of bounded constraint to achieve stable reconstruction of the background field. If the inversion algorithm is not suitable, such as direct point-by-point reconstruction, the instability of the problem will have a significant impact on the results. Figure 13 As shown in the middle (right), if the background field is directly inverted without using the above regularization method, the background field reconstruction will have a large deviation, and it may even be impossible to reduce the residuals or make data predictions.

[0216] In addition, we explored the case of an analytic function, taking u(x,y,t) = sin(πt-x)cos / 2x + 2(yt)1e -0.01t The solution does not strictly correspond to the passive transport equation, and this method can still achieve good compression results, as shown in Table 2.

[0217] Example: Analytic Functions original compression ratio zip 1758k 1086k 61.8% Extraction mechanism + zip 1758k 655k 37.3%

[0218] Table 2: Comparison of Data Compression Effects of Analytic Functions

[0219] Corresponding to the technical problems to be solved in the embodiments of this application, the above examples solve the following key technical problems: First, for high-precision mechanistic data with spatiotemporal correlation, where it is difficult to compress using general compression methods, this method achieves effective compression by using a compression method based on the dimensionality reduction idea of ​​the PDE mechanistic model, achieving a compression rate of 60% or even lower than that of ordinary encoding. Second, this method has good adaptability to various types of mechanistic data, and can still achieve compression results even if the data itself does not strictly satisfy the control equations.

[0220] The data compression method provided in this application differs from related technologies in the following aspects: Improvement 1: Compared to general high-dimensional data compression methods, this method can effectively compress high-precision mechanistic data. Mechanistic data, especially solutions to differential equations, often lacks sparsity, with continuously changing variables, wide value distribution, and high information entropy. Therefore, traditional image compression or encoding methods struggle to achieve effective compression. Improvement 2: Compared to fixed equation models, this method introduces model degrees of freedom such as background field coefficients, effectively broadening the model's applicability. It can adaptively optimize different data to obtain a model that better fits the data, thereby reducing model bias and achieving better compression results. Furthermore, when lossy compression is permissible, this method further highlights its advantages by discarding random noise.

[0221] Actual observation data and remote sensing data often contain a certain level of random noise. Therefore, during the reconstruction process, noise will affect the compression effect. Pure white noise data cannot be effectively compressed using mechanistic models. Thus, the compression effect of these actual data will generally be reduced.

[0222] Furthermore, for cases involving diffusion mechanisms, to better align the mechanistic model with the data, this method considers introducing the diffusion term νΔu into the equation, where the diffusion coefficient ν is an undetermined model parameter, optimized together with the background field during calculation. It is important to note that after introducing diffusion, to avoid instability in the forward solution caused by backdiffusion, the direction of temporal evolution must first be determined based on the magnitudes of the data at both ends of the time dimension to correctly select the initial diffusion layer.

[0223] The following example uses the ISABEL-Hurricane dataset to illustrate the spatiotemporal evolution of atmospheric composition. A 150×150×10 floating-point array is used, and lossless compression is applied. The initial and final time-series data distributions are as follows: Figure 16 As shown, there is a clear attenuation phenomenon and noise is present. The data reconstructed through model inversion is as follows: Figure 17 As shown in the figure, the reconstructed residuals have been significantly reduced by an order of magnitude (note that precise reconstruction is not required here; as long as the order of magnitude of the residuals is effectively reduced, the compression effect can be achieved). The specific compression effect is shown in Table 3.

[0224] Example: Atmospheric data original compression ratio Raw data + zip 1758k 792k 45.1% Extraction mechanism + zip 1758k 748k 42.5%

[0225] Table 3: Comparison of Data Compression Effects

[0226] The effectiveness of the data compression method provided in this application is summarized and compared as follows: Figure 18 As shown, the data compression method provided in this application can effectively extract the data mechanism and still achieve compression effect under adverse conditions such as high noise.

[0227] The data compression method provided in this application is based on the idea of ​​model inversion. For data that generally follows a certain type of spatiotemporal evolution, by introducing model degrees of freedom such as background field and diffusion coefficient, the coefficients of the inversion equation are obtained through the current data, thereby obtaining the specific model / equation that best matches the data. It has outstanding effects on data with significant intrinsic mechanisms.

[0228] The data compression method provided in this application is based on differential equation theory and the idea of ​​local linearization. It can stably reconstruct the background field using single or even partial data, obtain the mechanism approximately satisfied by the data, has low overhead, and closely matches the laws of the data itself.

[0229] Compared to the two-dimensional cross-sectional mechanism model compression method, this method considers the correlation in the time direction, thus enabling the processing of multi-layer data at once. Furthermore, the model's degrees of freedom are limited to the equation coefficients, and the coefficients are also dimensionality-reduced, achieving a balance between computational efficiency and compression effect.

[0230] The mechanism model of this application has good predictive ability and generalization characteristics, and has potential applications in scenarios such as high-dimensional data imputation.

[0231] The algorithm for reconstructing the background field can be applied to practical engineering applications such as velocity field measurement.

[0232] The data compression method provided in this application is applicable to the compression of high-dimensional data with spatiotemporal continuity. It utilizes the ideas of model inversion, localization, and relevant theories of differential equations, proposing a model dominated by transport and diffusion mechanisms, along with a complete algorithm for obtaining optimal model coefficients and compression / decompression based on data. This method can support the superior compression needs of various scientific computing and measurement data, as well as high-precision remote sensing data in the era of massive data. For other types of data, it is also expected that the mechanism compression algorithm provided in this application can be tried after preprocessing or appropriate scale decomposition.

[0233] Furthermore, this method can be applied to more general models, such as diffusion coefficients and variable coefficients for the zeroth-order term. These coefficients can be assumed to be sparse or low-dimensional, thus not fundamentally increasing storage overhead but providing a better fit to the data and achieving better compression. For discrete formats, it can also be generalized to implicit or multi-layered formats in the time direction, and is compatible with the second-order time derivative case. The uniqueness and stability of the coefficient inversion of the above equations are guaranteed by partial differential equation inversion theory. This method represents a new attempt at the intersection of differential equation inversion theory and data compression.

[0234] Based on the same concept as the aforementioned data compression method embodiment, this application also provides a data compression apparatus 1900, which can compress mechanistic data, reduce the compression ratio, and reduce storage overhead. The data compression apparatus 1900 includes components for implementing... Figure 2-8The units or modules of each step in the data compression method shown.

[0235] Figure 19 This is a schematic diagram of a data compression device provided in an embodiment of this application. Figure 19 As shown, the data compression device 1900 includes a first acquisition module 1901, a determination module 1902, an inversion module 1903, and a compression module 1904. The first acquisition module 1901 acquires data to be compressed, which has mechanistic characteristics. The determination module 1902 determines the mechanistic model and boundary terms. The mechanistic model describes the mechanism of the data to be compressed and includes partial differential equations. The boundary terms indicate the distribution boundaries of the data to be compressed. The inversion module 1903 inverts the mechanistic model based on the data to be compressed to obtain parameters of the mechanistic model, including the background field of the partial differential equations or the supports of the background field. The compression module 1904 compresses the data to be compressed based on the background field or the supports of the background field and the boundary terms.

[0236] In one possible implementation, the inversion module 1903 is specifically used to: initialize the background field based on the evolution trend of the data to be compressed, and obtain the initial value of the background field; solve the partial differential equation based on the boundary terms and the initial value of the background field, and obtain the solution of the partial differential equation; determine the first residual based on the solution of the partial differential equation and the data to be compressed; and optimize and adjust the background field of the partial differential equation with the goal of minimizing the first residual, and obtain the background field.

[0237] In another possible implementation, a specific approach to solving the partial differential equation based on the boundary terms and the initial values ​​of the background field is as follows: the background field is regularized, with the regularization parameters determined based on the noise in the data to be compressed; the partial differential equation is then solved based on the boundary terms and the initial values ​​of the regularized background field to obtain the solution.

[0238] In another possible implementation, the inversion module 1903 is specifically used to: initialize the background field support based on the evolution trend of the data to be compressed, and obtain the initial value of the background field support; solve the partial differential equation based on the boundary terms and the initial value of the background field support, and obtain the solution of the partial differential equation; determine the second residual based on the solution of the partial differential equation and the data to be compressed; and adjust the support of the background field of the partial differential equation with the goal of minimizing the second residual, and obtain the support of the background field.

[0239] In another possible implementation, the support of the background field includes boundary terms and boundary value parameters of the background field. The boundary terms indicate the distribution boundary of the background field, and the boundary terms and boundary value parameters are used to determine the background field. The second residual is calculated based on the objective function, which includes residual terms and regularization terms. The value of the residual terms is determined based on the data to be compressed and the solution of the partial differential equation, and the value of the regularization terms is determined based on the regularization parameter and the boundary value parameters of the background field. A specific implementation of adjusting the support of the background field of the partial differential equation with the goal of minimizing the second residual is as follows: minimizing the function value of the objective function is the objective of optimizing and adjusting the boundary terms and boundary value parameters of the background field; the support of the background field is obtained based on the optimized and adjusted boundary terms and boundary value parameters of the background field.

[0240] In another possible implementation, the compression of the data to be compressed is also related to the target residual, which indicates the difference between the solution of the target partial differential equation and the data to be compressed. The solution of the target partial differential equation is obtained by solving the partial differential equation based on the boundary terms and the optimized background field.

[0241] In another possible implementation, the data compression device 1900 provided in this application further includes a verification module 1905, which is used to verify the target residual; if the mean of the target residual is less than a preset threshold, the verification passes.

[0242] In another possible implementation, the parameters of the mechanistic model also include the source terms of the partial differential equation; the inversion module 1903 is specifically used to: initialize the background field based on the evolution trend of the data to be compressed, and obtain the initial value of the background field; solve the partial differential equation based on the boundary terms and the initial value of the background field, and obtain the solution of the partial differential equation; determine the source terms based on the solution of the partial differential equation and the data to be compressed; optimize and adjust the background field of the partial differential equation with the goal of minimizing the source terms, and obtain the background field.

[0243] In another possible implementation, the compression of the data to be compressed is also related to the target source term, which indicates the difference between the solution of the target partial differential equation and the data to be compressed. The solution of the target partial differential equation is obtained by solving the partial differential equation based on the boundary terms and the optimized background field.

[0244] In another possible implementation, the accuracy of the parameters of the mechanistic model is determined based on the accuracy of the data to be compressed.

[0245] In another possible implementation, the data compression apparatus 1900 provided in this application further includes a block segmentation module 1906, which is used to segment the data to be compressed into blocks to obtain multiple data blocks to be compressed.

[0246] In another possible implementation, the boundary terms include initial values ​​and / or boundary values, where the initial values ​​indicate the data distribution of the data to be compressed at the initial moment, and the boundary values ​​indicate the spatial distribution boundaries of the data to be compressed.

[0247] The data compression apparatus 1900 according to the embodiments of this application can correspond to the execution of the methods described in the embodiments of this application, and the above and other operations and / or functions of each module in the data compression apparatus 1900 are respectively for implementing Figure 2-8 For the sake of brevity, the corresponding processes of each method in the code will not be elaborated here.

[0248] Based on the same concept as the aforementioned data decompression method embodiment, this application also provides a data decompression apparatus 2000. This data decompression apparatus 2000 can decompress compressed data compressed using the data compression method provided in this application embodiment, and reconstruct the original data. The data decompression apparatus 2000 includes components for implementing... Figure 9-10 The units or modules of each step in the data decompression method shown.

[0249] Figure 20 This is a schematic diagram of a data decompression apparatus provided in an embodiment of this application. Figure 20 As shown, the data decompression device 2000 includes a second acquisition module 2001, an analysis module 2002, a solution module 2003, and a decompression module 2004. The second acquisition module 2001 acquires compressed data, which is obtained by compressing the data to be compressed based on the data compression method described in the first aspect or any possible implementation thereof. The analysis module 2002 analyzes the compressed data to obtain analysis results, which include parameters and boundary terms of a mechanistic model. The mechanistic model describes the mechanism of the data to be compressed and includes partial differential equations. The parameters include the background field or the support of the background field of the partial differential equations, and the boundary terms indicate the distribution boundaries of the data to be compressed. The solution module 2003 solves the partial differential equations based on the background field or the support of the background field, and the boundary terms to obtain the solution to the partial differential equations. The decompression module 2004 obtains decompressed data based on the solution to the partial differential equations.

[0250] In one possible implementation, the analytical result also includes residuals, which indicate the difference between the solution to the partial differential equation and the data to be compressed; the decompression module 2004 is specifically used to obtain decompressed data based on the solution to the partial differential equation and the residuals. Thus, the uncompressed original data is obtained losslessly through the solution to the partial differential equation and the residuals.

[0251] The data decompression apparatus 2000 according to the embodiments of this application can correspond to the execution of the data decompression method described in the embodiments of this application, and the above and other operations and / or functions of each module in the data decompression apparatus 2000 are respectively for implementing Figure 9-10 For the sake of brevity, the corresponding processes of each method in the code will not be elaborated here.

[0252] This application embodiment also provides a computing device, including at least one processor, a memory, and a communication interface, wherein the processor is used to execute... Figure 2-10 The method described.

[0253] Figure 21 A schematic diagram of the structure of a computing device provided in an embodiment of this application.

[0254] like Figure 21 As shown, the computing device 2100 includes at least one processor 2101, a memory 2102, and a communication interface 2103. The processor 2101, memory 2102, and communication interface 2103 are communicatively connected, which can be achieved via a wired (e.g., bus) or wireless connection. The communication interface 2103 is used to send and / or receive data sent by other devices. The memory 2102 stores computer instructions, and the processor 2101 executes these instructions to perform the data compression method described in the foregoing method embodiments, compressing the mechanistic data with high quality to reduce the compression ratio and storage overhead, and to perform the data decompression method described in the foregoing method embodiments, decompressing the compressed data compressed using the data compression method provided in this application embodiment to reconstruct the original data.

[0255] It should be understood that in the embodiments of this application, the processor 2101 may be a central processing unit (CPU), and the processor 1801 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0256] The memory 2102 may include read-only memory and random access memory, and provides instructions and data to the processor 2101. The memory 2102 may also include non-volatile random access memory. Optionally, the random access memory may be, for example, high bandwidth memory (HBM).

[0257] The memory 2102 can be volatile memory or non-volatile memory, or it can include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0258] It should be understood that the computing device 2100 according to the embodiments of this application can execute the implementation of the embodiments of this application. Figure 2-10 The method shown is described in detail above, and will not be repeated here for the sake of brevity.

[0259] Embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, wherein when the computer instructions are executed by a processor, the aforementioned method is implemented.

[0260] An embodiment of this application provides a chip including at least one processor and an interface, wherein the at least one processor determines program instructions or data through the interface; the at least one processor is used to execute the program instructions to implement the method mentioned above.

[0261] Embodiments of this application provide a computer program or computer program product that includes instructions that, when executed, cause a computer to perform the methods mentioned above.

[0262] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0263] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented using hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0264] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A data compression method, characterized in that, include: Acquire data to be compressed, the data to be compressed having mechanistic characteristics; The mechanism model and boundary terms are determined. The mechanism model describes the mechanism of the data to be compressed and includes partial differential equations. The boundary terms indicate the distribution boundary of the data to be compressed. The mechanism model is inverted based on the data to be compressed to obtain the parameters of the mechanism model, including the background field of the partial differential equation or the support of the background field. The data to be compressed is compressed based on the background field or the support of the background field and the boundary terms.

2. The method according to claim 1, characterized in that, The step of inverting the mechanism model based on the data to be compressed to obtain the parameters of the mechanism model includes: Based on the evolution trend of the data to be compressed, the background field is initialized to obtain the initial value of the background field; Based on the boundary terms and the initial values ​​of the background field, the partial differential equation is solved to obtain the solution of the partial differential equation; Based on the solution of the partial differential equation and the data to be compressed, the first residual is determined; With the goal of minimizing the first residual, the background field of the partial differential equation is optimized and adjusted to obtain the background field.

3. The method according to claim 2, characterized in that, The process of solving the partial differential equation based on the initial values ​​of the boundary terms and the background field to obtain the solution includes: The background field is regularized, and the parameters of the regularization are determined based on the noise of the data to be compressed. Based on the boundary terms and the initial values ​​of the regularized background field, the partial differential equation is solved to obtain the solution of the partial differential equation.

4. The method according to claim 1, characterized in that, Based on the data to be compressed, the mechanism model is inverted to obtain the parameters of the mechanism model, including: Based on the evolution trend of the data to be compressed, the background field support is initialized to obtain the initial value of the background field support; Based on the initial values ​​of the boundary terms and the background field support, the partial differential equation is solved to obtain the solution of the partial differential equation; Based on the solution of the partial differential equation and the data to be compressed, the second residual is determined; With the goal of minimizing the second residual, the support of the background field of the partial differential equation is adjusted to obtain the support of the background field.

5. The method according to claim 4, characterized in that, The support of the background field includes the boundary terms and boundary value parameters of the background field. The boundary terms of the background field indicate the distribution boundary of the background field. The boundary terms and boundary value parameters of the background field are used to determine the background field. The second residual is calculated based on an objective function, which includes a residual term and a regularization term. The value of the residual term is determined based on the data to be compressed and the solution of the partial differential equation, and the value of the regularization term is determined based on a regularization parameter and a boundary parameter of the background field. The step of adjusting the support of the background field of the partial differential equation with the objective of minimizing the second residual to obtain the support of the background field includes: The objective is to minimize the function value of the objective function, and to optimize and adjust the boundary terms and boundary value parameters of the background field. Based on the optimized boundary terms and boundary value parameters of the background field, the support of the background field is obtained.

6. The method according to any one of claims 2-5, characterized in that, The compression of the data to be compressed is also related to the target residual, which indicates the difference between the solution of the target partial differential equation and the data to be compressed. The solution of the target partial differential equation is obtained by solving the partial differential equation based on the boundary terms and the optimized background field.

7. The method according to claim 6, characterized in that, The compression of the data to be compressed based on the background field or the support of the background field, and the boundary terms, further includes the following steps before: The target residual is examined; If the mean of the target residual is less than a preset threshold, the test is passed.

8. The method according to claim 1, characterized in that, The parameters of the mechanistic model also include the source terms of the partial differential equation; The step of inverting the mechanism model based on the data to be compressed to obtain the parameters of the mechanism model includes: Based on the evolution trend of the data to be compressed, the background field is initialized to obtain the initial value of the background field; Based on the boundary terms and the initial values ​​of the background field, the partial differential equation is solved to obtain the solution of the partial differential equation; Based on the solution of the partial differential equation and the data to be compressed, the source term is determined; With the goal of minimizing the source term, the background field of the partial differential equation is optimized and adjusted to obtain the background field.

9. The method according to claim 8, characterized in that, The compression of the data to be compressed is also related to the target source term, which indicates the difference between the solution of the target partial differential equation and the data to be compressed. The solution of the target partial differential equation is obtained by solving the partial differential equation based on the boundary term and the optimized background field.

10. The method according to any one of claims 1-9, characterized in that, The accuracy of the parameters of the mechanistic model is determined based on the accuracy of the data to be compressed.

11. The method according to any one of claims 1-10, characterized in that, The process of obtaining the data to be compressed further includes: The data to be compressed is divided into blocks to obtain multiple data blocks to be compressed.

12. The method according to any one of claims 1-11, characterized in that, The boundary terms include initial values ​​and / or boundary values, wherein the initial values ​​indicate the data distribution of the data to be compressed at the initial moment, and the boundary values ​​indicate the spatial distribution boundaries of the data to be compressed.

13. A data decompression method, characterized in that, include: Obtain compressed data, wherein the compressed data is obtained by compressing the data to be compressed based on the method described in any one of claims 1-12; The compressed data is parsed to obtain the parsing results. The parsing results include the parameters and boundary terms of the mechanism model. The mechanism model describes the mechanism of the data to be compressed. The mechanism model includes partial differential equations. The parameters include the background field of the partial differential equations or the support of the background field. The boundary terms indicate the distribution boundary of the data to be compressed. Based on the background field or background field support and boundary terms, the partial differential equation is solved to obtain the solution of the partial differential equation; Based on the solution of the partial differential equation, the decompression data is obtained.

14. The method according to claim 13, characterized in that, The analytical results also include residuals, which indicate the difference between the solution to the partial differential equation and the data to be compressed; Based on the solution of the partial differential equation, the decompression data is obtained, including: The decompression data is obtained based on the solution of the partial differential equation and the residual.

15. A data compression device, characterized in that, include: The first acquisition module is used to acquire data to be compressed, the data to be compressed having mechanistic characteristics; A determination module is used to determine the mechanism model and boundary terms. The mechanism model describes the mechanism of the data to be compressed and includes partial differential equations. The boundary terms indicate the distribution boundary of the data to be compressed. The inversion module is used to invert the mechanism model based on the data to be compressed to obtain the parameters of the mechanism model, the parameters including the background field of the partial differential equation or the support of the background field; A compression module is used to compress the data to be compressed based on the background field or the support of the background field and the boundary terms.

16. A data compression device, characterized in that, include: The second acquisition module is used to acquire compressed data, wherein the compressed data is obtained by compressing the data to be compressed based on the method described in any one of claims 1-12; The parsing module is used to parse the compressed data and obtain the parsing result. The parsing result includes the parameters and boundary terms of the mechanism model. The mechanism model describes the mechanism of the data to be compressed. The mechanism model includes partial differential equations. The parameters include the background field of the partial differential equation or the support of the background field. The boundary terms indicate the distribution boundary of the data to be compressed. The solution module is used to solve the partial differential equation based on the background field or background field support and boundary terms to obtain the solution of the partial differential equation. The decompression module is used to obtain decompressed data based on the solution of the partial differential equation.

17. A computing device, comprising a memory and a processor, characterized in that, The memory stores instructions that, when executed by a processor, cause the method described in any one of claims 1-14 to be implemented.

18. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it causes the method as described in any one of claims 1-14 to be implemented.