Intelligent compression and reconstruction method and system for large-scale scientific calculation data
By using multi-dimensional physical feature recognition and differentiated compression strategies, combined with reconstruction optimization under physical constraints, the problems of uneven data compression and artifacts in existing technologies are solved, achieving efficient and reliable scientific computing data compression and reconstruction.
Patent Information
- Application Number
- CN202610083375.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-22
- Publication Date
- 2026-02-17
AI Technical Summary
Existing technologies cannot effectively distinguish the information value and physical importance of different regions in large-scale scientific computing data compression. This results in the destruction of details in key physical structure regions or the failure to utilize the compression potential of smooth regions. Furthermore, the decompression process cannot guarantee that the reconstructed data follows physical laws, introducing artifacts and reducing scientific usability.
By identifying multi-dimensional physical features, a differentiated compression strategy is generated. Based on regional feature identifiers and quality weight configurations, a differentiated compression algorithm is used to process scientific computing data. In the reconstruction stage, physical constraints are introduced for consistency verification and iterative optimization to generate optimized reconstructed data.
While maintaining a high compression rate, it significantly improves the fidelity of core scientific information, corrects artifacts, ensures that the reconstructed data remains physically correct and credible, adapts to changes in data complexity, and provides stable and high-quality compression results.
Smart Images

Figure CN121547057A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of digital data processing technology and relates to a method and system for intelligent compression and reconstruction of large-scale scientific computing data. Background Technology
[0002] Large-scale scientific computing, such as numerical simulations in fields like weather forecasting, aerospace, energy exploration, and biomedicine, is a crucial tool in modern scientific research and engineering design. These computational processes typically generate massive amounts of multidimensional array data, which precisely describe the evolution and distribution of physical quantities in space and time. How to effectively store, transmit, and manage these enormous datasets has become a major challenge in the field of high-performance computing; therefore, data compression technology is of paramount importance in this context.
[0003] In existing technologies, compression methods for scientific computing data mainly rely on general-purpose lossy compression algorithms. These methods, such as those based on predictive coding, transform coding, or quantization techniques, typically treat the entire data field as a unified numerical matrix and use a single, global error tolerance or compression ratio as control parameters. During processing, the algorithm treats all data points in the data field equally, applying the same compression transform and quantization criteria, aiming to minimize the global mean square error or meet a preset peak signal-to-noise ratio.
[0004] However, the aforementioned existing technical solutions have significant technical shortcomings. The information value and physical importance carried by different regions in a scientific computing data field are extremely uneven. For example, the accuracy of data from shock wave or vortex regions in a flow field is crucial for characterizing the overall physical phenomenon compared to a stable laminar flow region. The indiscriminate compression strategy employed in existing technologies often results in severe damage to the details of key physical structural regions while ensuring the overall compression rate; or, in order to protect these key regions, the overall compression rate is reduced, thus failing to fully utilize the compression potential of smooth regions. Furthermore, the traditional decompression process is merely the mathematical inverse operation of compression, which cannot guarantee that the reconstructed data field still follows the original physical conservation laws, often introducing artifacts that do not conform to physical laws, thus reducing the scientific usability of the compressed data. Summary of the Invention
[0005] In view of this, in order to solve the problems mentioned in the background technology, a method and system for intelligent compression and reconstruction of large-scale scientific computing data is proposed.
[0006] The objective of this invention can be achieved through the following technical solution: The first aspect of this invention provides a method for intelligent compression and reconstruction of large-scale scientific computing data, including: S1, acquiring the original large-scale scientific computing data to be compressed.
[0007] S2. Perform multi-dimensional physical feature identification on the raw data of large-scale scientific computing and generate multi-dimensional physical features.
[0008] S3. Generate differentiated compression strategies based on multi-dimensional physical features.
[0009] S4. Compress the raw data for large-scale scientific computing according to the differentiated compression strategy to generate a compressed data stream.
[0010] S5. Obtain the compressed data stream and the physical constraints driving the reconstruction, and use the physical constraints to parse the compressed data stream to perform differentiated reconstruction and generate preliminary reconstruction data.
[0011] S6. Perform physical consistency verification and iterative optimization on the preliminary reconstructed data to finally generate optimized reconstructed data.
[0012] A second aspect of the present invention provides a large-scale scientific computing data intelligent compression and reconstruction system, comprising: a raw data acquisition module for acquiring the large-scale scientific computing raw data to be compressed.
[0013] The multi-dimensional physical feature generation module performs multi-dimensional physical feature recognition on the raw data of large-scale scientific computing and generates multi-dimensional physical features.
[0014] The differentiated compression strategy generation module generates differentiated compression strategies based on multi-dimensional physical features.
[0015] The compressed data stream generation module compresses the raw data of large-scale scientific computing according to a differentiated compression strategy to generate a compressed data stream.
[0016] The preliminary reconstruction data generation module obtains the compressed data stream and the physical constraints driving the reconstruction, and uses the physical constraints to parse the compressed data stream to perform differentiated reconstruction and generate preliminary reconstruction data.
[0017] The reconstruction data optimization module performs physical consistency verification and iterative optimization on the initial reconstruction data, and finally generates optimized reconstruction data.
[0018] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects: (1) By identifying multi-dimensional physical features in the original data and generating differentiated compression strategies, the present invention can intelligently allocate limited coding resources to high-value sensitive areas containing key physical phenomena, while implementing higher compression in areas where data changes are gradual, thereby greatly improving the fidelity of core scientific information while ensuring a high compression rate.
[0019] This invention introduces physical constraints during the reconstruction phase to parse and reconstruct the compressed data stream, and performs physical consistency verification and iterative optimization on the preliminary results. This effectively corrects artifacts or deviations that may be introduced by lossy compression and violate physical laws, ensuring that the final reconstructed data field not only approximates the original data numerically, but also maintains correctness and credibility in a physical sense.
[0020] The method proposed in this invention incorporates a closed-loop feedback mechanism of real-time error monitoring and dynamic strategy adjustment, which adaptively adjusts the compression strategy based on the actual performance of the data during the compression process. This makes the method highly robust to time-varying or non-uniform complex scientific data, maintaining near-optimal performance throughout the entire compression task and ensuring stable and high-quality compression results for large-scale datasets in both time and spatial dimensions. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of the method steps of the present invention.
[0023] Figure 2 This is a schematic diagram of the system structure connection of the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] Please see Figure 1 The first aspect of the present invention provides a method for intelligent compression and reconstruction of large-scale scientific computing data, including: S1, acquiring the original large-scale scientific computing data to be compressed.
[0026] S2. Perform multi-dimensional physical feature identification on the raw data of large-scale scientific computing and generate multi-dimensional physical features.
[0027] In a specific embodiment of the present invention, the specific steps of performing multi-dimensional physical feature identification on large-scale scientific computing raw data and generating multi-dimensional physical features include: performing physical field continuity analysis on large-scale scientific computing raw data to identify high-value sensitive regions and low-value smooth regions.
[0028] It should be noted that the process of identifying multi-dimensional physical features from large-scale scientific computing raw data and generating multi-dimensional physical features aims to assign different importance levels to different regions in the raw data field through intelligent analysis, thereby laying the foundation for subsequent differentiated compression. This process first performs physical field continuity analysis on the acquired large-scale scientific computing raw data to be compressed. Large-scale scientific computing raw data typically exists in the form of multi-dimensional arrays, representing the spatial or temporal distribution of one or more physical quantities. The core of physical field continuity analysis lies in quantifying the drastic changes in values within the data field. This can be achieved by calculating the gradient magnitude at each point in the data field. For example, for a three-dimensional scalar field data... Its point gradient at It can be represented as: ,in It is the gradient operator, which is usually approximated in discrete data fields by numerical methods such as central difference; The norm of the vector, i.e., its magnitude. The calculated gradient magnitude. A gradient field of the same size as the raw data from the large-scale scientific computation is created. By setting a gradient threshold, the raw data field can be divided into different regions. Regions with gradient amplitudes higher than the gradient threshold indicate drastic changes in physical quantities, such as shock surfaces, vortex cores, or areas of material stress concentration in fluid simulations; these are identified as high-value sensitive regions. Conversely, regions with gradient amplitudes lower than the gradient threshold represent gradual changes in physical quantities, such as stable laminar flow regions or regions of uniform material without stress; these are identified as low-value smooth regions.
[0029] It should also be noted that the gradient operator When applied to a scalar field, it generates a vector field representing the rate of change of the scalar field in each direction. For three-dimensional scalar field data... Its gradient is: In discrete data fields, partial derivatives need to be approximated by numerical methods.
[0030] Structural correlation analysis is performed on raw data for large-scale scientific computing to identify regions of coupling relationships between multiple physical variables.
[0031] It should be noted that after completing the continuity analysis, the method then performs structural correlation analysis on the raw data of large-scale scientific computing. Scientific computing data typically contains multiple interrelated physical variables, such as velocity, pressure, and temperature in fluid dynamics. The purpose of structural correlation analysis is to identify regions where these physical variables are tightly coupled. This can be achieved by calculating the correlation coefficients between different physical variable fields within the local neighborhood of the data. For example, selecting two physical variable fields... and At data points Calculate their Pearson correlation coefficient within a small window around them. , ,in, It is a physical variable field and covariance, and These are physical variable fields. and The standard deviation of the correlation coefficient is taken as the threshold. If the calculated Pearson correlation coefficient exceeds the preset correlation coefficient threshold, it indicates that there is a strong coupling relationship between the physical variables in that region, and the region is identified as a coupling region. This region usually corresponds to the key area of multiphysics interaction, and the accuracy of its data is crucial to the correctness of the overall simulation.
[0032] Regional feature identifiers are generated based on the recognition results.
[0033] It should be noted that, based on the identification results of the aforementioned physical field continuity analysis and structural correlation analysis, the system generates a regional feature identifier for each data point or data block in the data field. This identifier is a label or classification code used to summarize the physical characteristics of the region. For example, label "1" can be set to represent a high-value sensitive region, label "2" to represent a coupled region, label "3" to represent a composite region that is both sensitive and coupled, and label "0" to represent a low-value smooth region. This forms a feature identifier field corresponding to the original data space.
[0034] Multi-dimensional physical features are generated based on regional feature identifiers and quality weight configurations.
[0035] It should be noted that, finally, based on the regional feature identifiers generated in the previous step and the pre-set quality weight configuration, the final multi-dimensional physical features are generated. The quality weight configuration is a mapping rule that maps each regional feature identifier to a specific numerical weight. This weight directly reflects the fidelity requirement that the regional data should retain during compression. For example, a user can configure the weight of high-value sensitive regions to be 0.9, the weight of coupled relationship regions to be 0.8, and the weight of low-value smooth regions to be 0.2. By replacing each identifier in the regional feature identifier field with its corresponding weight value, a quantified importance distribution map is generated; this is the aforementioned multi-dimensional physical feature.
[0036] This method, through in-depth analysis of the physical field continuity and structural correlation of the original data, can intelligently identify regions with different physical meanings and importance in scientific data. Compared with traditional indiscriminate compression methods, this method can accurately allocate limited coding resources to high-value sensitive regions and coupled relationship regions that require the most fidelity, while applying a higher degree of compression to low-value smooth regions with low information entropy. This differentiated processing based on physical characteristics can significantly improve the fidelity of reconstructing key physical phenomena while ensuring the overall data compression rate. This allows the reconstructed data field to better maintain the original physical consistency and scientific value, greatly improving the usability of compressed data in subsequent scientific analysis and visualization applications.
[0037] S3. Generate differentiated compression strategies based on multi-dimensional physical features.
[0038] In a specific embodiment of the present invention, the specific steps of generating a differentiated compression strategy based on multi-dimensional physical features include: obtaining global error constraints that indicate the maximum allowable distortion.
[0039] It should be noted that the implementation process of generating differentiated compression strategies based on the generated multi-dimensional physical features is a key step in transforming the feature analysis results of the physical domain into specific compression execution instructions. This process begins with obtaining global error constraints, which are pre-set by the user or system to indicate the maximum tolerable distortion of the entire large-scale scientific computing raw data after compression, typically quantified by peak signal-to-noise ratio. We define this as the overall maximum allowable distortion. .
[0040] Based on the regional feature identifiers in the multi-dimensional physical characteristics, different regions are assigned corresponding compression accuracy levels.
[0041] It should be noted that, next, the system assigns corresponding compression precision levels to regions with different physical characteristics based on the regional feature identifiers generated for each part of the data field in the previous steps. For example, regions identified as high-value sensitive areas are assigned a "high precision" level, while regions identified as low-value smooth areas are assigned a "low precision" level. This step maps abstract physical importance to qualitative compression processing requirements.
[0042] Based on the compression accuracy level and global error constraints, the error allocation budget for each region is calculated.
[0043] It should be noted that the method then performs a core quantization process, which calculates the precise value for each region based on the compression accuracy level of each region and global error constraints. Each region. We quantify compression accuracy levels into quality weights. A region with a higher weight should have a lower error budget allocated to it. Error allocation budget It can be calculated using the following formula: In this formula, This is the first The specific error allocation budget calculated for each region; It is the total number of points in the raw data for large-scale scientific computing; It is the first The number of data points contained in each region; and They are the first The region and the first The quality weights of each region are provided by multi-dimensional physical features; This represents a summation over all regions. This formula ensures that the error budget allocated to each region is inversely proportional to its physical importance, while guaranteeing that the weighted average of the errors from all regions, calculated by their size, precisely satisfies the overall maximum allowable distortion. Requirements.
[0044] Based on the error allocation budget and regional feature identifiers for each region, a differentiated compression strategy containing compression algorithm type and compression parameters is generated.
[0045] It should be noted that, finally, based on the calculated error allocation budget for each region and the corresponding region feature identifiers, a final differentiated compression strategy is generated. This differentiated compression strategy is a specific implementation plan, specifying the compression algorithm type and compression parameters for each region. The choice of algorithm type depends on the data characteristics reflected by the region feature identifiers. For example, for low-value smooth regions with smooth data, computationally efficient prediction or transform coding algorithms can be selected; for high-value sensitive regions containing complex structures, compression algorithms based on deep learning or wavelet transforms that better preserve details can be selected. After selecting an algorithm, its compression parameters, such as quantization step size, error limits, or bit rate, will be precisely set to ensure that the actual error generated in that region after compression can be strictly controlled within the previously calculated error allocation budget. The final generated differentiated compression strategy provides customized compression instructions for each independent region in the data field for subsequent compression processing modules.
[0046] This method achieves refined and differentiated control over compression distortion by establishing a quantitative mapping from global error constraints to regional error allocation budgets. Instead of applying a single, coarse compression standard to the entire data field, it intelligently allocates compression resources based on the physical importance of each region. This strategy ensures that the limited error budget is used in regions with the least impact on scientific analysis, thereby achieving a high compression ratio while maximizing the preservation of data fidelity in high-value sensitive areas and key physical structures. This enables the reconstructed data to more accurately serve subsequent scientific discoveries and engineering decisions, significantly improving the effectiveness and scientific value of data compression.
[0047] S4. Compress the raw data for large-scale scientific computing according to the differentiated compression strategy to generate a compressed data stream.
[0048] In a specific embodiment of the present invention, the specific steps of compressing the raw data of large-scale scientific computing according to the differentiated compression strategy to generate a compressed data stream include: selecting a corresponding compression algorithm instance according to the compression algorithm type in the differentiated compression strategy.
[0049] It should be noted that, firstly, for each identified region in the raw data of large-scale scientific computing, the system selects and instantiates a corresponding compression algorithm instance based on the compression algorithm type specified for that region in the differentiated compression strategy. A compression algorithm instance refers to the concrete implementation object of a specific compression algorithm at the software level, such as a compressor for the SZ algorithm or a compressor for the ZFP algorithm. The selection is based on the data characteristics of the region and the applicability of the algorithm. For example, the strategy might specify a predictive coding type algorithm for low-value smooth regions containing smooth fields, while specifying transformation-based or machine learning-based algorithms for high-value sensitive regions containing complex textures.
[0050] Configure the running parameters of the compression algorithm instance based on the compression parameters.
[0051] It should be noted that after selecting a compression algorithm instance, the system will configure the running parameters of this algorithm instance based on the compression parameters specified for that region in the differentiated compression strategy. These compression parameters are crucial for fine-tuning the compression process, with the most important parameters typically being the error limit or quantization level, whose values are directly derived from the error allocation budget for that region calculated in previous steps. For example, the system will assign the region... Error allocation budget Set the target error parameter to the selected compression algorithm instance. This ensures that the compression process for each region will strictly adhere to its assigned fidelity requirements.
[0052] The configured compression algorithm instance is used to compress the regions identified by regional feature identifiers in the raw data of large-scale scientific computing, generating regional compressed data.
[0053] It should be noted that after configuration, the system uses the configured compression algorithm instance to perform actual compression operations on the corresponding region data blocks in the original data, identified by region feature identifiers. The compression algorithm instance reads the original values of the region, processes them according to its internal logic, and finally outputs a binary data block, which is the region compressed data. Due to the lossy or lossless nature of compression, the size of the region compressed data is usually much smaller than the original data block. This process will be executed sequentially for all regions in the original data field, with each region using its own dedicated algorithm instance and parameter configuration.
[0054] The compressed data from each region is packaged with the corresponding region feature identifier to generate a compressed data stream.
[0055] It should be noted that, finally, to ensure the compressed data can be correctly reconstructed, the system packages the region-compressed data generated in each region with its corresponding region feature identifier. The region feature identifier, as metadata, implicitly contains information needed to reconstruct the data for that region, such as the decompression algorithm to be used and the spatial location of the data in the original field. All the compressed data packets from all regions are organized in a predetermined order, ultimately forming a single, continuous bitstream, which is the compressed data stream.
[0056] This method achieves truly intelligent and refined processing of large-scale scientific computing data by implementing a differentiated compression strategy. It avoids the one-size-fits-all compression approach of traditional methods, ensuring that computational resources and encoded bits are prioritized for protecting critical physical structures essential to scientific analysis. The resulting compressed data stream has a clear structure and self-contains the metadata needed for reconstruction. This not only significantly reduces the cost of data storage and transmission but also greatly ensures the scientific usability of the reconstructed data due to its respect for and protection of physical characteristics, laying a high-quality data foundation for subsequent data analysis, visualization, and scientific discovery.
[0057] In a specific embodiment of the present invention, the following operation is also included: during the compression process, the error generated in each region during the actual compression process is monitored in real time.
[0058] It should be noted that this method introduces a closed-loop feedback control mechanism while performing compression processing, the implementation process of which is as follows. During the process of compressing each region identified by regional feature identifiers according to the differentiated compression strategy and generating region compressed data, the system monitors the distortion introduced by the process in real time. Specifically, after compressing a region, the system immediately performs a temporary, in-memory decompression operation on the generated region compressed data to obtain temporary reconstructed data. Subsequently, by comparing this temporary reconstructed data with the original data of the region point by point, the actual compression error is calculated. The calculation of the actual compression error typically uses the same metric as the error allocation budget, and its calculation formula is: In this formula, It is the first The actual compression error of each region; It is the first The total number of data points contained in each region; It is the first The first of the original data in the region The value of each data point; Then it is the first The corresponding region in the reconstructed data obtained through temporary decompression is the first region. The values of each data point.
[0059] The error generated during the actual compression process in each region is compared with the allocated error allocation budget to obtain the error comparison results.
[0060] It should be noted that, next, the system will calculate the actual compression error. The error allocation budget allocated to this region in the previous steps A comparison is performed. This comparison produces an error comparison result, which clearly indicates the deviation of the current compression strategy. For example, if... Greater than If the compression is too aggressive for this type of region, it indicates that the current compression is too aggressive, leading to excessive errors; conversely, if... Less than If the value is 0, it means that the compression is too conservative and there is still room to improve the compression ratio.
[0061] The differential compression strategy is dynamically adjusted based on the error comparison results.
[0062] In a specific embodiment of the present invention, the specific steps of dynamically adjusting the differentiated compression strategy based on the error comparison results include: when the actual compression error of a certain region exceeds its error allocation budget, the compression accuracy level of that region is increased.
[0063] It should be noted that the specific steps for dynamically adjusting the differentiated compression strategy based on error comparison results constitute the core closed-loop control logic for achieving the adaptive capability of this method. When the real-time monitoring system detects that the actual compression error generated during the compression process in a region identified by a regional feature identifier exceeds its allocated error budget, the adjustment process is triggered. The first step is to upgrade the compression accuracy level of the category to which the problematic region belongs. This operation does not directly modify an abstract level label, but rather specifically adjusts the quality weight associated with the regional feature identifier of that region.
[0064] To maintain the overall error constraints, the compression accuracy level of other regions is reduced accordingly.
[0065] It's important to note that in the second step, to maintain the overall error constraints, the system must compensate for the aforementioned accuracy improvements. This means that error budget must be "saved" in other regions. The system will strategically select to reduce the compression accuracy level of one or more other regions, typically those regions that currently have good compression performance and low physical importance, such as low-value smoothing regions. Specifically, this involves lowering the quality weights corresponding to these region types.
[0066] Based on the adjusted compression accuracy levels of all regions, the error allocation budget for each region is recalculated.
[0067] It should be noted that in the third step, based on the newly adjusted compression accuracy levels for all regions—that is, the updated quality weight set—the system recalculates the error allocation budget for each region. This calculation process uses the previously established mathematical model, but employs the updated weights as input. The new error allocation budget... This can be derived from the following formula: In this formula, This is the first The new error allocation budget is calculated for each region; and Then they represent the first The and the first The quality weights of each region are updated after the first two steps of adjustment.
[0068] Based on the recalculated error allocation budget, a new differentiated compression strategy is generated.
[0069] It should be noted that, finally, based on this recalculated error allocation budget, the system updates and generates a new differentiated compression strategy. In the new strategy, the compression parameters for each region, especially the error limits or quantization levels, will be set to their latest error allocation budget values. This updated differentiated compression strategy will take effect immediately, guiding the compression processing of subsequent time steps or subsequent data blocks, thus applying the adjustment effects to future compression tasks.
[0070] This method transforms the static compression process into an intelligent system with adaptive and self-correcting capabilities by introducing a closed-loop feedback mechanism of real-time error monitoring and dynamic strategy adjustment. Its technical advantage lies in significantly enhancing the robustness of the compression method to changes in data complexity. Scientific computing data, especially in time-series simulations, has internal characteristics such as the generation and movement of shock waves that are not static. This method can automatically detect the fluctuations in compression performance caused by these changes and proactively adjust the strategy to adapt to new data characteristics, thus avoiding the problems of severely reduced fidelity in local areas or poor overall compression efficiency that may result from using a fixed strategy. This dynamic optimization ensures that the system can continuously achieve near-optimal compression performance while satisfying global error constraints throughout the entire compression task, guaranteeing stable and high-quality compression results for large-scale datasets in both time and space dimensions.
[0071] S5. Obtain the compressed data stream and the physical constraints driving the reconstruction, and use the physical constraints to parse the compressed data stream to perform differentiated reconstruction and generate preliminary reconstruction data.
[0072] In a specific embodiment of the present invention, the specific steps of parsing the compressed data stream using physical constraints to perform differentiated reconstruction and generate preliminary reconstructed data include: parsing the compressed data stream and extracting regional compressed data and corresponding regional feature identifiers.
[0073] It should be noted that the process of parsing the compressed data stream using physical constraints to perform differentiated reconstruction and generate preliminary reconstructed data is the reverse operation of the compression process, aiming to recover the complete data field from the compact compressed data stream. First, the reconstruction system receives the compressed data stream and parses it. The compressed data stream is a structured bit sequence, internally packaging compressed data for each region and its corresponding region feature identifier. The parsing process involves separating these information units one by one from the data stream according to predetermined format rules, extracting the region's compressed data and its unique region feature identifier for each region. This region feature identifier is the key metadata guiding differentiated reconstruction.
[0074] Based on the region's feature identifiers, a reconstruction algorithm is selected that corresponds to the type of compression algorithm used in the region's compression process.
[0075] It should be noted that, next, the system selects a reconstruction algorithm for each region based on the extracted regional feature identifiers, corresponding to the compression algorithm type used in the region's compression process. Since the compression process is differentiated, the reconstruction process must also be differentiated accordingly. The regional feature identifier implicitly contains information about the compression algorithm type used for that region. Therefore, by consulting a pre-set algorithm mapping table, the system can accurately match the corresponding decompression or reconstruction algorithm based on the regional feature identifier.
[0076] Based on physical constraints and regional feature identifiers, the reconstruction parameters of the reconstruction algorithm are configured.
[0077] It's important to note that after selecting a reconstruction algorithm, the system needs to configure its parameters to ensure it can correctly process the received regional compressed data. The configuration is primarily based on the physical constraints driving the reconstruction and the regional feature identifiers. Physical constraints guide the reconstruction process, making the results more consistent with physical reality. For example, if physical constraints specify that the data field should maintain a certain continuity or smoothness, then the interpolation or filtering parameters in the reconstruction algorithm can be adjusted accordingly. Simultaneously, regional feature identifiers also provide clues for parameter configuration; for instance, they may contain information such as the quantization step size or error limits used during compression. This information can be directly used as input parameters for the reconstruction algorithm to ensure that the reconstruction accuracy matches the settings used during compression.
[0078] The configured reconstruction algorithm is used to reconstruct the compressed data of each region, generating regional reconstruction data.
[0079] It should be noted that after configuration, the system uses the configured reconstruction algorithm to decompress and reconstruct the corresponding region compressed data. The reconstruction algorithm reads the binary region compressed data and, based on its internal logic and configuration parameters, reverse-engineers the compression operation to generate numerical region reconstructed data. This process is performed independently for each region in the compressed data stream, with each region using a tailored algorithm and parameters.
[0080] The reconstructed data from each region are merged according to the original data organization structure to generate preliminary reconstructed data.
[0081] It should be noted that, finally, after all the regional reconstruction data has been generated, the system needs to recombine these scattered data blocks into a complete data field. Based on the original spatial location and index information contained in the regional feature identifiers, the system precisely places each regional reconstruction data point back into its original position within the data field. In this way, all regional reconstruction data are seamlessly stitched together, ultimately fusing to generate preliminary reconstruction data with the same dimensions and structure as the original data.
[0082] This method achieves efficient and high-fidelity parsing of compressed data streams through differentiated reconstruction. Instead of using a single, universal decompression procedure, it leverages regional feature identifiers embedded during compression as a blueprint, calling the most suitable reconstruction tools and parameters for each data region. In particular, by introducing physical constraints to guide the reconstruction process, the reconstruction results are no longer merely mathematical approximations, but rather physically approximate the true solution. This intelligent, physics-based reconstruction strategy significantly suppresses artifacts and maintains structural continuity, resulting in preliminary reconstructed data with higher overall physical consistency and scientific credibility, providing a high-quality foundation for subsequent verification, optimization, and final applications.
[0083] S6. Perform physical consistency verification and iterative optimization on the preliminary reconstructed data to finally generate optimized reconstructed data.
[0084] In a specific embodiment of the present invention, the specific steps of performing physical consistency verification and iterative optimization on the preliminary reconstructed data to finally generate optimized reconstructed data include: during the reconstruction process, performing physical consistency verification on the generated regional reconstructed data based on physical constraints to identify abnormal regions.
[0085] It should be noted that the process of performing physical consistency verification and iterative optimization on the preliminary reconstructed data is a quality improvement and correction step based on the preliminary reconstruction. This process is initiated during or after the reconstruction process. First, based on user-provided or system-built-in physical constraints, the physical consistency of the reconstructed data in each region of the generated preliminary reconstructed data is verified. Physical constraints are mathematical expressions describing the fundamental laws that the simulated physical system should follow, such as the conservation of mass in fluid dynamics or the zero divergence condition for incompressible flow. The specific operation of physical consistency verification is to use these mathematical expressions as verification operators and apply them to the preliminary reconstructed data. For example, for a vector field that should satisfy zero divergence, the system will calculate the numerical divergence of the preliminary reconstructed data in that region. If the calculated divergence value deviates significantly from zero and exceeds the preset physical tolerance, the region is judged to have physical consistency anomalies and is identified as an abnormal region.
[0086] When a physical consistency anomaly is detected, an anomaly identifier is generated for the abnormal region.
[0087] It should be noted that when the system detects one or more abnormal regions through the above verification, it will immediately generate an anomaly identifier for these regions. This anomaly identifier is a special metadata tag attached to the corresponding region data, explicitly indicating that the physical consistency of that region has failed the verification. This step enables the system to accurately pinpoint and track the range of data that requires further processing.
[0088] Based on the anomaly identifiers of the abnormal regions and the original regional feature identifiers, a physical constraint-enhanced reconstruction algorithm is selected.
[0089] It should be noted that, next, the system, based on the newly generated anomaly identifier and the original regional feature identifier of the anomaly region, jointly decides and selects the most suitable physical constraint-enhanced reconstruction algorithm. This selection is highly targeted because the type of anomaly and the data characteristics of the region itself jointly determine the optimal repair strategy. For example, if the anomaly manifests as a severe discontinuity at the region boundary, and the region itself is a low-value smooth region, the system may select an interpolation reconstruction algorithm with boundary smoothing constraints.
[0090] Physical constraint-enhanced reconstruction algorithms are used to reconstruct and optimize abnormal regions, generating optimized reconstruction data.
[0091] In a specific embodiment of the present invention, the specific steps for reconstructing and optimizing anomaly regions using the physical constraint-enhanced reconstruction algorithm include: obtaining reconstruction data of adjacent regions of the anomaly region.
[0092] It should be noted that the specific steps of using the physical constraint-enhanced reconstruction algorithm to optimize the reconstruction of anomaly regions constitute a refined process aimed at fundamentally repairing data physical consistency deviations. This process first requires acquiring the reconstruction data of the directly adjacent regions of the anomaly region. Because the data of these adjacent regions has already passed physical consistency verification, it is considered a "trustworthy" reference point. It will provide crucial boundary constraints for the repair of the anomaly region, ensuring that the repaired region can smoothly and continuously connect with its surrounding environment.
[0093] A physical field continuity constraint model is constructed based on physical constraints. This model is used to quantify the difference between the corrected data and the initial reconstructed data and incorporate boundary constraints.
[0094] It should be noted that the core step of the method then involves constructing a physical field continuity constraint model based on pre-defined physical constraints. This physical field continuity constraint model is essentially an optimization problem, aiming to find optimal corrected data that satisfies the physical constraints while approximating the original preliminary reconstructed data as closely as possible. This physical field continuity constraint model is typically expressed as a cost function containing multiple objective terms, which can be represented as follows: In this formula, It is the total cost that needs to be minimized; It is reconstructed data after optimization and correction; It is the initial reconstruction data that was initially generated within the abnormal region and has abnormal physical consistency; It is a mathematical operator defined according to physical constraints; for example, if the physical constraint is an incompressible flow, then... It is a divergence operator; It is a regularization coefficient used to balance the weight between the two objectives of maintaining the original form of data and satisfying physical constraints. Its typical value is usually selected through experiments or experience based on the physical characteristics of the specific problem, the level of data noise, and the strictness of the constraints. The norm squared is typically used to quantify the magnitude of the variance or residual. The model must also incorporate boundary constraints derived from neighboring regions, i.e., it requires... The value at the boundary of the outlier region must be equal to the value of the reconstructed data in the adjacent region at the corresponding location.
[0095] The preliminary reconstruction data of the abnormal region is corrected using a physical field continuity constraint model to generate corrected reconstruction data.
[0096] It should be noted that the system subsequently uses the constructed physical field continuity constraint model to correct the initial reconstructed data of the anomalous region. This process is typically solved using iterative numerical optimization algorithms to find the solution that satisfies the cost function. Minimized This solution process essentially involves fine-tuning and correcting the initial reconstructed data under the strong guidance of physical laws, leading to the final solution. This refers to the corrected and reconstructed data.
[0097] The corrected reconstructed data is smoothly fused with the acquired reconstructed data from adjacent areas to generate optimized reconstructed data.
[0098] It should be noted that, finally, to completely eliminate any subtle seams or discontinuities that may occur at the region boundaries, the system performs a smooth fusion of the corrected reconstructed data generated in the previous step with the acquired reconstructed data from adjacent regions. This operation applies a weighted blending function at the region boundaries to ensure that the values and gradients of the data transition smoothly when crossing the boundaries of the original anomalous regions. After this fusion process, the final optimized reconstructed data for that region is generated.
[0099] It should also be noted that when smoothing and fusing the reconstructed data with the reconstructed data from adjacent regions, the blending function used is typically a linear weighted gradient function, which takes the following form: At the boundary between the outlier region and the adjacent region, a transition band of width d is defined (e.g., d = 3 grid points), and the blending function... With distance (The normalized distance from the boundary of the anomalous region to the interior of the adjacent region, s∈[0,1]) linear transformation: The merged data value is: ,in, This represents the reconstruction data of adjacent areas.
[0100] This method significantly enhances the scientific credibility of the final reconstructed data field by introducing a closed-loop process of physical consistency verification and iterative optimization. It goes beyond simply achieving numerical approximation; it strives for accuracy in terms of physical laws. By precisely identifying and targeting physical inconsistencies introduced during compression and initial reconstruction, such as artifacts or violations of conservation laws, this method corrects critical scientific errors. This post-processing optimization mechanism ensures that the final reconstructed data field is not only visually smooth and continuous but, more importantly, strictly adheres to the underlying physical principles at the data level. This allows the compressed data to provide reliable and accurate input when used for subsequent quantitative analysis, scientific computation verification, or to drive new simulations, thus fundamentally guaranteeing the value of data compression throughout the entire scientific research chain.
[0101] Reference Figure 2The second aspect of the present invention provides a large-scale scientific computing data intelligent compression and reconstruction system, comprising: a raw data acquisition module, a multi-dimensional physical feature generation module, a differentiated compression strategy generation module, a compressed data stream generation module, a preliminary reconstruction data generation module, and a reconstruction data optimization module.
[0102] The original data acquisition module is connected to the multi-dimensional physical feature generation module, the multi-dimensional physical feature generation module is connected to the differentiated compression strategy generation module, the differentiated compression strategy generation module is connected to the compressed data stream generation module, the compressed data stream generation module is connected to the preliminary reconstruction data generation module, and the preliminary reconstruction data generation module is connected to the reconstruction data optimization module.
[0103] The raw data acquisition module acquires the raw data for large-scale scientific computing to be compressed.
[0104] The multi-dimensional physical feature generation module performs multi-dimensional physical feature recognition on the raw data of large-scale scientific computing and generates multi-dimensional physical features.
[0105] The differentiated compression strategy generation module generates differentiated compression strategies based on multi-dimensional physical features.
[0106] The compressed data stream generation module compresses the raw data of large-scale scientific computing according to a differentiated compression strategy to generate a compressed data stream.
[0107] The preliminary reconstruction data generation module obtains the compressed data stream and the physical constraints driving the reconstruction, and uses the physical constraints to parse the compressed data stream to perform differentiated reconstruction and generate preliminary reconstruction data.
[0108] The reconstruction data optimization module performs physical consistency verification and iterative optimization on the preliminary reconstruction data, and finally generates optimized reconstruction data.
[0109] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, and all such modifications and additions should fall within the protection scope of the present invention.
Claims
1. A method for intelligent compression and reconstruction of large-scale scientific computing data, characterized in that, The method comprises the following steps: S1, obtaining large-scale scientific computing original data to be compressed; S2, identifying multi-dimensional physical features of the large-scale scientific computing original data to generate multi-dimensional physical features; S3, generating a differentiated compression strategy based on the multi-dimensional physical features; S4, compressing the large-scale scientific computing original data according to the differentiated compression strategy to generate a compressed data stream; S5, obtaining the compressed data stream and a physical constraint condition for driving reconstruction, and using the physical constraint condition to analyze the compressed data stream to perform differentiated reconstruction to generate preliminary reconstruction data; S6, performing physical consistency verification and iterative optimization on the preliminary reconstruction data to finally generate optimized reconstruction data.
2. The method of claim 1, wherein, The specific steps of identifying multi-dimensional physical features of the large-scale scientific computing original data to generate multi-dimensional physical features comprise: performing physical field continuity analysis on the large-scale scientific computing original data to identify high-value sensitive areas and low-value smooth areas; performing structure correlation analysis on the large-scale scientific computing original data to identify coupling relationship areas among multiple physical variables; generating area feature identifiers based on the identification results; generating multi-dimensional physical features according to the area feature identifiers and quality weight configurations.
3. The method of claim 2, wherein, The specific steps of generating a differentiated compression strategy based on the multi-dimensional physical features comprise: obtaining a global error constraint condition indicating an overall allowed maximum distortion; allocating corresponding compression precision levels to different areas according to the area feature identifiers in the multi-dimensional physical features; calculating error allocation budgets of the areas based on the compression precision levels and the global error constraint condition; and generating a differentiated compression strategy containing compression algorithm types and compression parameters according to the error allocation budgets of the areas and the area feature identifiers.
4. The method of claim 3, wherein, The specific steps of compressing the large-scale scientific computing original data according to the differentiated compression strategy to generate a compressed data stream comprise: selecting corresponding compression algorithm instances according to the compression algorithm types in the differentiated compression strategy; configuring running parameters of the compression algorithm instances based on the compression parameters; compressing each area identified by the area feature identifiers in the large-scale scientific computing original data using the configured compression algorithm instances to generate area compressed data; packing the area compressed data and the corresponding area feature identifiers to generate a compressed data stream.
5. The method of claim 4, wherein, The method further comprises the following operations: monitoring errors generated in actual compression processes of the areas in real time during the compression process; comparing the errors generated in the actual compression processes of the areas with the allocated error allocation budgets to obtain error comparison results; dynamically adjusting the differentiated compression strategy based on the error comparison results.
6. The method of claim 5, wherein, The specific steps of dynamically adjusting the differentiated compression strategy based on the error comparison results comprise: when it is detected that the actual compression error of an area exceeds its error allocation budget, increasing the compression precision level of the area; correspondingly decreasing compression precision levels of other areas to maintain the overall error constraint condition; recomputing error allocation budgets of the areas according to the adjusted compression precision levels of all the areas; updating a new differentiated compression strategy based on the recomputed error allocation budgets.
7. The method of claim 4, wherein, The specific steps of using the physical constraint condition to analyze the compressed data stream to perform differential reconstruction to generate the preliminary reconstruction data include: Analyzing the compressed data stream to extract the region compressed data and the corresponding region feature identifier; Selecting a reconstruction algorithm corresponding to the compression algorithm type used in the compression process of the region according to the region feature identifier; Configuring the reconstruction parameters of the reconstruction algorithm based on the physical constraint condition and the region feature identifier; Reconstructing each region compressed data using the configured reconstruction algorithm to generate region reconstruction data; Fusing each region reconstruction data according to the original data organization structure to generate the preliminary reconstruction data.
8. The method of claim 7, wherein, The specific steps of performing physical consistency verification and iterative optimization on the preliminary reconstruction data to finally generate the optimized reconstruction data include: During the reconstruction process, performing physical consistency verification on the generated region reconstruction data based on the physical constraint condition to identify abnormal regions; When detecting the physical consistency abnormality, generating an abnormal identifier for the abnormal region; Selecting a physical constraint enhanced reconstruction algorithm based on the abnormal identifier of the abnormal region and the original region feature identifier; Using the physical constraint enhanced reconstruction algorithm to optimize the reconstruction of the abnormal region to generate the optimized reconstruction data.
9. The method of claim 8, wherein, The specific steps of using the physical constraint enhanced reconstruction algorithm to optimize the reconstruction of the abnormal region include: Obtaining the adjacent region reconstruction data of the abnormal region; Constructing a physical field continuity constraint model based on the physical constraint condition, which is used to quantify the difference between the corrected data and the initial reconstruction data and is included in the boundary constraint; Using the physical field continuity constraint model to correct the preliminary reconstruction data of the abnormal region to generate corrected reconstruction data; Smoothing and fusing the corrected reconstruction data and the obtained adjacent region reconstruction data to generate the optimized reconstruction data.
10. A large-scale scientific computing data intelligent compression and reconstruction system, characterized in that, It includes: An original data acquisition module acquires large-scale scientific computing original data to be compressed; A multi-dimensional physical feature generation module identifies multi-dimensional physical features of the large-scale scientific computing original data to generate multi-dimensional physical features; A differential compression strategy generation module generates a differential compression strategy based on the multi-dimensional physical features; A compressed data stream generation module compresses the large-scale scientific computing original data according to the differential compression strategy to generate a compressed data stream; A preliminary reconstruction data generation module acquires the compressed data stream and the physical constraint condition for driving reconstruction, and uses the physical constraint condition to analyze the compressed data stream to perform differential reconstruction to generate preliminary reconstruction data; and a reconstruction data optimization module performs physical consistency verification and iterative optimization on the preliminary reconstruction data to finally generate optimized reconstruction data.