Ocean field data lossy compression method considering spatial-temporal heterogeneity
By dividing the ocean spatio-temporal field data into uniform data blocks and performing spatial and temporal feature modal decomposition, the problem of difficulty in determining compression parameters and poor compression effect in traditional methods is solved, and efficient ocean spatio-temporal field data compression is achieved.
Patent Information
- Application Number
- CN202510006816.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-11
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-30
AI Technical Summary
The traditional lossy compression method of ocean space-time data has problems such as high computational complexity, poor compression effect, and difficult to control compression errors, and it is especially difficult to effectively use a small number of principal components to preserve data details.
By dividing the ocean three-dimensional spatiotemporal field data into uniform data blocks in space, dimensionality reduction and feature modal decomposition are performed considering the spatial and temporal heterogeneity, and the compression error is adaptively controlled to realize the blocked feature extraction and compression of the data.
This method significantly improves the compression ratio and compression accuracy, and the compression error has a uniform time and spatial distribution, providing efficient compression effect under a given error requirement.
Smart Images

Figure CN120074537A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of marine data processing, and particularly to a lossy compression method for marine field data considering spatio-temporal heterogeneity. Background Art
[0002] Marine field data generally refers to a set of marine element data that is spatio-temporally continuous. Common types include temperature, salinity, sea surface height, etc. Due to characteristics such as data density, high dimensionality, and large scale, marine spatio-temporal field data often occupies a large amount of memory. Moreover, with the progress of marine observation technology and numerical simulation technology, the accuracy of marine spatio-temporal field data continues to improve, and the data volume increases rapidly. In order to facilitate the storage and organization of such data, data compression technology will play an important role.
[0003] Data compression technology can be divided into lossless compression and lossy compression. Lossless compression can ensure that no data is lost during the compression process, while lossy compression sacrifices a certain degree of accuracy in exchange for a higher compression ratio. The compression ratio of lossless compression usually does not exceed 2:1. Therefore, for the compression of massive marine field data, lossy compression methods are usually preferred. Commonly used lossy compression algorithms include the Discrete Wavelet Transform (DWT) algorithm, the Discrete Cosine Transform (DCT) algorithm, and the Empirical Orthogonal Function (EOF) algorithm, etc.
[0004] Due to the complexity of the biological effects, chemical processes, and kinetic processes in the marine system itself, the spatio-temporal heterogeneity of marine spatio-temporal field data is relatively strong. When traditional data lossy compression methods are applied to the compression of marine spatio-temporal field data, there are often problems such as high computational complexity, poor compression effect, and difficult control of compression error. For example, the method based on principal component analysis extraction has the problem of being difficult to retain data details with a small number of principal components, resulting in an unclear compression effect. Summary of the Invention
[0005] Object of the Invention: The present invention aims to provide a lossy compression method for marine field data that considers spatio-temporal heterogeneity, extracts block features from high-dimensional marine spatio-temporal field data, and realizes adaptive control of compression error.
[0006] Technical Solution: The lossy compression method for marine field data considering spatio-temporal heterogeneity according to the present invention includes the following steps:
[0007] (1) Divide the three-dimensional marine spatio-temporal field data into n spatially uniformly distributed field data blocks in space;
[0008] (2) Consider spatio-temporal heterogeneity to perform dimensionality reduction and feature mode decomposition on the data blocks, and organize and represent the decomposition results in the form of several spatio-temporal feature levels;
[0009] (3) Set the maximum compression error error_def, determine the number of main levels, where the main levels are the levels that contain the main change information of the original data blocks. Extract the main levels from the level representation, store the main level features and the corresponding time series in the form of objects, and at the same time record the chunking information of the data and the mean value of all spatial grid points over time to achieve data compression.
[0010] Furthermore, step (2) is specifically as follows:
[0011] After the dimensionality reduction of the data blocks, a two-dimensional spatio-temporal data matrix C is obtained i , i = 1, 2,..., n, C i is a two-dimensional matrix where the rows represent the time dimension and the columns represent the spatial dimension;
[0012] Center C i to obtain the centered matrix C' i . Determine the correlation matrix V and obtain its eigenvector matrix E based on the relationship between the number of rows l and the number of columns p of the centered matrix C'; among them, the eigenvectors in E are arranged in descending order according to the corresponding eigenvalues; i
[0013] According to the centered matrix C' i and the eigenvector matrix E, construct the transformation matrix TS, introduce the spatial function and the time function, and determine the result of the eigenmode decomposition;
[0014] Organize the result of the eigenmode decomposition into a row vector Bi, and the j-th column element of this vector is the j-th eigenlevel of the i-th data block; concatenate the vectors B i in sequence to obtain the total matrix B composed of all eigenlevels.
[0015] Furthermore, the centered matrix C' i is
[0016]
[0017] where 1 l = (1,…,1) T is used to unify the matrix dimensions, represents the mean value of the p-th column of C i .
[0018] Furthermore, when l > p:
[0019]
[0020] Otherwise:
[0021]
[0022] Furthermore, the result of the eigenmode decomposition is
[0023]
[0024] where t is the time position; s is the space position; u ij (s) is a space function, i.e., the j-th column of the eigenvector matrix E; is a time function, i.e., the j-th row of the transformation matrix TS; Mean i (t, s) represents the mean removed during centering; m i is the number of feature levels of the i-th data block.
[0025] Furthermore, the total matrix B is
[0026]
[0027] where the i-th row of the total matrix B represents the feature level of the i-th data block, and the j-th column represents the j-th feature level of all data blocks.
[0028] Furthermore, according to the number of main levels perform main level extraction on each row of the total matrix B respectively to obtain the main feature level matrix Q:
[0029]
[0030] Furthermore, the number of main levels satisfies:
[0031] Initially, lower = 1, upper = m i , where lower and upper represent the lower and upper limits of the number of main levels respectively;
[0032] Calculate the current compression root mean square error error. When error > error_def, then lower = r i , otherwise upper = r i , update the r i value: Loop until error is approximately equal to error_def to end the loop. At this time, the r i value is the final number of main levels.
[0033] Beneficial effects: Compared with the prior art, the significant advantages of the present invention are as follows: 1. The present invention evenly divides the complete spatio-temporal three-dimensional data in the spatial dimension, and respectively performs feature decomposition and error control within each independent data block, fully exploring the spatio-temporal variation characteristics of large-scale ocean spatio-temporal field data, and solving the problems such as difficult determination of compression parameters and poor compression effect in traditional methods; 2. The present invention organizes the data block information, the selected main spatial feature levels and the corresponding time variation sequences into a compression result file, realizing the dimensionality reduction compression of high-dimensional data; 3. The compression error of the present invention has a uniform time distribution and spatial distribution, providing efficient compression under a given error requirement; the compression ratio of direct compression = 23.35:1, R 2 = 98.31%, compared with direct compression, the compression performance after block organization is that the compression ratio = 50.2:1, R 2 = 99.34%, the compression ratio is increased by 115%, and the compression accuracy is increased by 1%. Description of the drawings
[0034] Figure 1 is the linear relationship between the compression error and the cumulative variance interpretation rate of the feature level;
[0035] Figure 2 is the time distribution of the compression error under different block numbers BC;
[0036] Figure 3 is the relationship between the actual compression error and the number of data blocks when the preset compression error is 0.03 m;
[0037] Figure 4 is the relationship between the compression ratio and the number of data blocks when the preset compression error is 0.03 m. Detailed implementation manners
[0038] The present invention will be further described below with reference to the drawings.
[0039] The compression effect evaluation indexes involved in the present invention are shown in Table 1.
[0040] Table 1 Compression effect evaluation indexes
[0041]
[0042] The lossy compression method for ocean field data considering spatio-temporal heterogeneity according to the present invention includes the following steps:
[0043] (1) Divide the ocean three-dimensional spatio-temporal field data into n spatially evenly distributed field data blocks in space.
[0044] Considering that for a given ocean spatio-temporal field data A ∈ R I×J×KThe distribution of (where I, J, and K represent the number of data points in the longitude, latitude, and time dimensions respectively) is unknown, and matrix decomposition operations are only applicable to regular grid data. In the present invention, A is spatially divided into several data blocks, and data blocks chunk(A,n) = {C 1 , C 2 , …, C n} are obtained, where the distribution of the attribute characteristics within each data block is more balanced compared to the original data. The number of grid points in each data block is In the actual compression process, a smaller spatial range often helps to fully explore and extract the spatio-temporal variation characteristics of the data.
[0045] (2) Considering spatio-temporal heterogeneity, dimensionality reduction and eigenmode decomposition are performed on the data blocks, and the decomposition results are organized and represented in the form of several spatio-temporal feature levels.
[0046] (21) Modal decomposition considering spatio-temporal heterogeneity
[0047] Based on considering spatio-temporal heterogeneity, the present invention realizes eigenmode decomposition of data blocks, and the related definitions and calculation processes are as follows:
[0048] For a three-dimensional spatio-temporal data block C i (t, x, y) with a relatively uniform spatial distribution obtained after being processed in step (1), it is converted into two-dimensional spatio-temporal data C i (t, s), where t represents the time point and s represents the spatial position, and the data values change with the change of the time point t and the spatial point s:
[0049]
[0050] The above C i is a two-dimensional matrix with rows representing the time dimension and columns representing the spatial dimension. x lp is the value of the field data C i at the time point t = l and the spatial position s = p. First, centralize C i to highlight the variation characteristics of the data within the current spatial range:
[0051]
[0052] In the above formula, 1 l = (1, …, 1) T is used to unify the matrix dimensions, represents the mean of the p-th column of C' i , that is, the average of the data of the p-th spatial grid point in the time dimension. Next, the matrix V reflecting the correlation between different data points is obtained to mine the dominant modes of the spatio-temporal field data. There are two cases for the calculation of the matrix V. When l > p:
[0053]
[0054] Otherwise:
[0055]
[0056] Continue to obtain the eigenvalues and eigenvectors of the correlation matrix V. The eigenvectors form a matrix E, and each column of E is an eigenvector of matrix V. These eigenvectors are the local feature modes considering spatial heterogeneity after restoring the two-dimensional structure. The eigenvalues are retained as indicators to measure the importance of each mode. Sort the eigenmodes according to the magnitudes of the eigenvalues. Considering that general sorting functions usually use ascending order, the obtained eigenvalues can be negated before sorting. Taking the case of Equation (6) as an example, the eigenvalues form an array ([λ 1 , λ 2 ,..., λ p ), denoted as ev. Taking the negation gives -ev = ([-λ 1 , -λ 2 ,...,-λ p ). Sort the array -ev in ascending order to obtain an index array idx of the array elements, and then use idx to index the second dimension of matrix E. Taking the Python programming language as an example, the calculation of idx and the sorting of the importance of eigenmodes can be expressed as: idx = np.argsort(-ev), E = E[:, idx]. Finally, calculate a transformation matrix TS, and the calculation method is: The j-th column of the sorted matrix E is the j-th largest spatial feature of the local data, and this feature distribution is only related to space and can be simply expressed as a spatial function u ij (s). The j-th row of the TS matrix represents the temporal variation of the contribution of the spatial feature represented by u ij (s) to the original data, and these variations are only related to time and are denoted as a temporal function The sum of the products of each spatial function and temporal function gives the original matrix C after removing the mean i . Use the modal decomposition method considering spatio-temporal heterogeneity to decompose the data matrix, and a single decomposition result can be expressed as:
[0057]
[0058] where t represents the time position, s represents the spatial position, Mean i (t, s) represents the mean removed during centering, and m i is the number of calculated eigenpairs.
[0059] That is, decompose each spatio-temporal data C i into multiple spatial functions And the corresponding time function A linear combination of Represents C i The spatial level part of the j-th important feature level, the time function Reflection For data block C i One-dimensional time series of the temporal variation of the contribution, m i The value represents the number of all feature levels calculated, that is, the number of all eigenvectors in the matrix decomposition, so there is In most cases, m of different blocks i The values are the same, but there may still be a difference of ±1, so it is recorded here as m i This is to distinguish them and also to facilitate the research of different block strategies in the future. i (t,s) is the temporal mean of the original data. It does not change in time, so only spatial changes need to be retained when storing. i (s). The complete spatiotemporal data A∈R after removing the mean I×J×K Transformed into the total matrix B:
[0060]
[0061] In formula (9), B only lists all the feature levels obtained by decomposition in the form of a matrix for intuitive display, and is not a matrix in the mathematical sense. The same is true for the main feature level matrix Q involved in step 3. The rows of the total matrix B represent the level of each data block, and the jth column of B represents the jth important feature level of all data blocks, that is, the jth significant change feature of the data block. Formula (9) means that the complete A is divided into a total of n×m feature levels, thus completing the spatiotemporal decomposition of the ocean multidimensional spatiotemporal field data block and the dimensionality reduction expression of the data.
[0062] (3) Set the maximum compression error error_def and determine the number of main layers, i.e., the layer containing the main change information of the original data block. Extract the main layers from the layer representation, store the main layer features and the corresponding time series in the form of objects, and record the data block information and the mean of all spatial grid points in time to achieve data compression.
[0063] Based on all the hierarchical data obtained from equation (9) in step 2, the main changes in the ocean spatiotemporal field data are hidden in a relatively small number of main feature hierarchies. For each data block, the number of main feature hierarchies to be retained and stored needs to be determined separately. That is, each row of formula (9) is retained separately Item, the main feature hierarchy matrix Q is obtained as follows:
[0064]
[0065] During actual compression it is often much smaller than all the feature levels Therefore, the volume sizes of these main feature levels are significantly reduced relative to the original data. The main level features of each data block and the corresponding time series are stored in the form of objects respectively, while recording the block information of the data and the mean value Mean of all spatial grid points over time i (s). The huge three-dimensional spatio-temporal data set is represented by a small number of main feature levels and the mean value of the original data over time, achieving compression
[0066] Determination of compression parameters under a given error
[0067] The key point of step 3 lies in the determination of The present invention relates to a method for automatically adjusting the number of main feature levels of each data block based on a given compression error by taking into account the feature level differences of different data blocks and relying on the good linear relationship between the root mean square error of compression and the number of main levels. This method is implemented based on binary search. Denote the lower and upper limits of the possible number of main levels as lower and upper respectively, that is initially, lower = 1, upper = m i . The method of binary search is adopted to continuously adjust the number of main levels used for reconstructing each data block. For example, at the beginning, use levels to reconstruct the data block. The reconstruction process of a single data block can be expressed as:
[0068]
[0069] where is the value of the i-th data block after reconstruction. After specifying the maximum error error_def, the current root mean square error error is calculated through formula (2)
[0070] When error > error_def, take lower = r i , otherwise take upper = r i , and update the r i value: Continuously loop this process until the calculated error is approximately equal to error_def to end the loop. At this time, the r i is what is required. The above operations are performed on the feature levels of all data blocks in sequence to complete the spatio-temporal feature level extraction and organization of the entire data. This step has a high complexity and is the bottleneck for improving the algorithm performance. Subsequently, all the reconstructed data blocks are stitched together in the spatial dimension according to the block information to complete data reconstruction
[0071] (4) Experimental verification
[0072] (41) Experimental configuration
[0073] In this verification example, the daily average absolute dynamic topography data obtained by satellite observation and dynamic interpolation is used as the experimental data. The original data format is NetCDF. After reading, it is stitched together in time to form a three-dimensional data with dimensions of longitude, latitude, and time, denoted as Adt ∈ R 455×159×1917 , and the data is stored in double-precision format. The original data memory size is approximately 1058MB. The data is evenly divided and organized in space, and the block information is recorded for subsequent compression operations and data reconstruction.
[0074] (42) Data compression and index calculation
[0075] The compression ratio is an important indicator to evaluate the data compression effect. A higher compression ratio is also the advantage of lossy compression methods, so it is selected as one of the core indicators for experimental verification. In addition to the root mean square error with dimensions selected in step 3, the dimensionless coefficient of determination (R-Squared) is also selected as the evaluation index for compression accuracy during the verification process, which is used to evaluate the information restoration degree after decompression of this lossy compression method.
[0076] Apply the method proposed in the present invention to the hierarchical extraction and organization of Adt data, retain a few main feature levels, data block information, and the time dimension mean as the compressed file, calculate the compressed file size and compression ratio, and then reconstruct the decompressed data from the compressed file and calculate the compression accuracy and the overall root mean square error of compression by comparing with the original data. To highlight the role of block processing in reducing compression errors, the number of feature levels r of each data block is fixed for reconstruction, and the variation of compression error with time is calculated, as Figure 2 shown. Then, an r-value determination algorithm based on error control is adopted to calculate the compression ratio and compression error.
[0077] (43) Comparative experiment
[0078] As Figure 2As shown, block processing can significantly reduce the compression error and obtain a more stable error time distribution. To further prove the improvement of block division for this compression method, 20 groups of block numbers (Blocks Count, BC), namely 1 (i.e., no block organization), 4, 9, 16, …, 400, were set for comparative experiments. The results were mainly compared with the traditional direct compression decomposition (BC = 1). The traditional direct compression method ignores the characteristic of uneven spatial distribution of large-scale spatio-temporal three-dimensional data, and directly compresses, making it difficult to capture the change information of data with a small number of feature levels, resulting in a low compression ratio. While this method adopts the method of block-based feature level extraction and organization to improve the compression ratio, and fully considers the feature differences of each data block to determine the compression parameters for different blocks respectively, realizing a more balanced error distribution. Specifically, the optimal r value was selected by controlling the root mean square error at 0.03m, and the results are as Figure 3 and Figure 4 shown. It can be found that when the value of the block number BC exceeds 64, the error will show large fluctuations. The reason for the error fluctuations caused by too many block numbers should be that the data blocks are too delicate, and there may be obvious differences in the hierarchical features between each block, and the error control capabilities are not the same, that is, there are large differences in the errors between each block. Therefore, the value of BC should not be too large. The results also show that as the value of BC increases, the compression ratio first rises and then gradually decreases, and BC = 144 is the demarcation point between increase and decrease. When the value of BC is between 64 and 196, a compression ratio of more than 50 can be guaranteed. To balance the compression ratio and the error control ability, it can be considered that 64 is a better block number for this set of Adt data.
Claims
1. A lossy compression method for ocean field data considering temporal and spatial heterogeneity, characterized in that: The following steps are involved: (1) Divide the ocean three-dimensional spatiotemporal field data into n spatially evenly distributed field data blocks; (2) Considering the spatiotemporal heterogeneity, the data blocks are reduced in dimension and decomposed into characteristic modes, and the decomposition results are organized and represented in the form of several spatiotemporal feature levels; (3) Set the maximum compression error error_def and determine the number of main layers, i.e., the layer containing the main change information of the original data block. Extract the main layers from the layer representation, store the main layer features and the corresponding time series in the form of objects, and record the data block information and the mean of all spatial grid points in time to achieve data compression.
2. The method for lossy compression of ocean field data considering temporal and spatial heterogeneity according to claim 1, characterized in that: Step (2) is as follows: After the data block is reduced in dimension, a two-dimensional spatiotemporal data matrix C is obtained. i , i=1,2,…,n,C i It is a two-dimensional matrix with rows representing time dimensions and columns representing space dimensions; C i Centralize and get the centralized matrix C' i , by the centralization matrix C' i The correlation matrix V is determined by the relationship between the number of rows l and the number of columns p and its eigenvector matrix E is obtained; wherein the eigenvectors in E are arranged in descending order according to the size of the corresponding eigenvalues; According to the centralization matrix C' i and the eigenvector matrix E, construct the transformation matrix TS, introduce the space function and time function, and determine the eigenmode decomposition result; The characteristic mode decomposition results are organized into a row vector Bi, the j-th column element of the vector is the j-th feature level of the i-th data block; the vectors B are concatenated in sequence i Get the total matrix B consisting of all feature levels.
3. The method for lossy compression of ocean field data considering temporal and spatial heterogeneity according to claim 2, characterized in that: The centralization matrix C' i for Among them, 1 l =(1,…,1) T , used to unify the matrix dimensions, Represents C i The mean of the pth column.
4. The method for lossy compression of ocean field data considering temporal and spatial heterogeneity according to claim 3 is characterized in that: When l>p: otherwise:
5. The method for lossy compression of ocean field data considering temporal and spatial heterogeneity according to claim 4, characterized in that: The characteristic mode decomposition result is Among them, t is the time position; s is the spatial position; u ij (s) is the spatial function, i.e., the jth column of the eigenvector matrix E; is the time function, i.e. the jth row of the transformation matrix TS; Mean i (t,s) represents the mean removed during centering; m i is the number of feature levels of the i-th data block.
6. The method for lossy compression of ocean field data considering temporal and spatial heterogeneity according to claim 5, characterized in that: The total matrix B is Among them, the i-th row of the total matrix B represents the feature level of the i-th data block, and the j-th column represents the j-th feature level of all data blocks.
7. The method for lossy compression of ocean field data considering temporal and spatial heterogeneity according to claim 6, characterized in that: According to the number of main levels Perform main level extraction on each row of the total matrix B to obtain the main feature level matrix Q:
8. The method for lossy compression of ocean field data considering temporal and spatial heterogeneity according to claim 7, characterized in that: Number of main levels satisfy: Initially, lower = 1, upper = m i , Among them, lower and upper represent the upper and lower limits of the number of main levels respectively; Calculate the current compression root mean square error error. When error>error_def, lower=r i , otherwise upper = r i , update r i value: The loop ends when error is approximately equal to error_def. i The value is the final number of primary levels.
Citation Information
Cited By
Two-dimensional wave field reconstruction method based on compressed sampling
CN121257611A