A lossless data compression method and system based on orthogonal symbol partitioning

CN122820867APending Publication Date: 2026-09-25钟琦巧
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610854797.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-13
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]符号冲突带来的问题包括:1. 转义开销;2. 长距离重复丢失;3. 解码歧义

Benefits of technology

[0011]1. 零符号冲突;2. 超倍率压缩(实测886:1);3. 解码确定性;4. 自适应降级;5.格式无关;6. 解码效率提升(实测吞吐量较gzip提升5-8倍);7. 存储空间节省;8. 硬件实现简化(FPGA减少约12% LUT和8%功耗);9. 零转义语义自识别。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820867A_ABST
    Figure CN122820867A_ABST
Patent Text Reader

Abstract

The application discloses a lossless data compression method and system based on orthogonal symbol partition. By dividing the storage space into an original data area and a high-dimensional compressed symbol area, and the encoding addresses of the two areas not overlapping with each other, the repeated mode in the data is mapped to the high-dimensional character space, and N:1 replacement with zero symbol conflict is realized. When specific rule data is detected, the system further extracts the generation parameters describing the data and constructs a reconstruction script, realizing parameterized extreme compression. The application is especially suitable for highly regularized data, and can achieve a compression ratio of hundreds to thousands (886:1 for a 512*512 chessboard BMP image in the embodiment) under the premise of lossless, and has strong decoding certainty; for irregular data, the application can be automatically degraded to ensure no expansion, covering the whole data regularity spectrum.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention belongs to the field of data compression technology, and specifically relates to a method and system for achieving lossless compression through coding space partitioning. [Background Technology]

[0002] Current mainstream lossless compression algorithms (such as LZ77, LZ78, BWT, arithmetic coding, and entropy coding) all perform encoding operations within a single byte space (0x00-0xFF, a total of 256 code points). Within this limited space, the substitution symbols introduced during the compression process inevitably overlap with the original data bytes in code points, resulting in symbol conflicts.

[0003] The problems caused by symbol conflicts include: 1. Escape overhead; 2. Loss of long-distance repetition; 3. Decoding ambiguity. Taking a 512×512 pixel checkerboard BMP image as an example, using state-of-the-art compression tools such as zstd yields almost no compression effect, exposing the fundamental limitations of traditional single-space architectures on highly regularized data.

[0004] Therefore, there is an urgent need for a lossless compression method that eliminates symbol conflicts at the physical level and breaks through the limitation of single-space code points. [Summary of the Invention]

[0005] The core architecture of this invention is orthogonal symbol partitioning, which divides the storage / encoding space into two non-overlapping subspaces: the original data area (0x00-0xFF) and the compressed symbol area (high-dimensional characters). The compressed symbol area uses fixed-length encoding, and the high-dimensional characters are selected from the Unicode Private Use Area (PUA). Their byte boundaries are self-identified by the prefix bit of the first byte of UTF-8, and the decoder can directly determine the symbol length without additional escaping.

[0006] Based on this, the present invention provides a rule optimization path and a general algorithm path. The rule optimization path extracts generation parameters to construct a reconstruction script when a specific rule pattern is detected, achieving extreme compression; the general algorithm path discovers repeating patterns through a suffix array and maps them to high-dimensional characters to complete an N:1 permutation.

[0007] [Source of the Invention Concept]

[0008] Over more than a decade of engineering practice, the inventor has internalized tools and methodologies such as PDCA closed-loop management, FMEA potential failure mode analysis, 5 Whys root cause tracing, fishbone diagram system analysis, and lean improvement into engineering intuition for solving anomalies in complex systems.

[0009] During the conceptualization of this invention, the inventors used cutting-edge AI models as auxiliary tools to conduct experiments and verify computational limits. However, the core inventive concept of this invention—transferring the industrial domain's methodology of partitioning, simplification, and closed-loop verification to the field of data compression, forming a complete technical solution of 'orthogonal symbol partitioning + global indexing + high-dimensional character mapping'—stems from the inventors' long-term accumulated engineering experience and cross-domain systematic thinking, rather than the direct output of AI.

[0010] [Beneficial Effects]

[0011] 1. Zero symbol collisions; 2. Super-rate compression (886:1 measured); 3. Deterministic decoding; 4. Adaptive degradation; 5. Format independence; 6. Improved decoding efficiency (5-8 times higher throughput than gzip measured); 7. Storage space saving; 8. Simplified hardware implementation (FPGA reduces LUTs by approximately 12% and power consumption by 8%); 9. Zero-escape semantic self-recognition. [Attached Image Description]

[0012] Figure 1 : Schematic diagram of orthogonal symbol partitioning architecture. This illustrates the complete architecture of the compression encoding of this invention: the original file is searched by a global retrieval module to find all repeating substrings; the profit sorting module arranges candidate patterns in descending order of profit score; the orthogonal partitioning mapping module maps replacement symbols to the high-dimensional character space (Unicode private usage area) of the compressed symbol area; the lossless substitution encoding module performs N:1 symbol replacement; and the decoding and restoration module achieves lossless restoration by querying the dictionary mapping table.

[0013] Figure 2 Rule optimization path flowchart

[0014] Figure 3 General Algorithm Path Flowchart

[0015] Figure 4 Decoder logic structure diagram

[0016] Figure 5 Comparison table of measured data of various types

Detailed Implementation Methods

[0017] Example 1: Rule-optimized path, 512×512 checkerboard BMP, compression ratio 886:1, MD5 passed.

[0018] Example 2: General algorithm path, based on repeating pattern discovery, greedy selection, high-dimensional mapping and permutation of suffix array + LCP array.

[0019] Example 3: Medium regularity data compression (2.8:1).

[0020] Example 4: Low regularity / random data compression (~1:1, no expansion).

[0021] Example 5: Decoder logic structure.

[0022] Summary table of measured data of various types (see the original instruction manual).

[0023] [Industrial Applicability]

[0024] It can be applied to scenarios such as screen recording and compression, industrial visual inspection image storage, key frame archiving of surveillance videos, and cloud storage big data transmission.

Claims

1. A lossless data compression method, characterized in that, include: The storage space is divided into a raw data area and a compressed symbol area, and the encoded addresses of the two areas do not overlap. The repetition pattern in the data to be compressed is mapped to a high-dimensional character in the compression symbol area; the repetition pattern is replaced with the high-dimensional character to generate compressed data; wherein the code points in the compression symbol area have no intersection with the code points in the original data area, thus achieving zero symbol collision.

2. The method according to claim 1, characterized in that: The high-dimensional characters are selected from the Unicode standard private use area code points (U+E000 to U+F8FF).

3. The method according to claim 1, characterized in that: Repeating patterns in the data to be compressed are discovered by constructing a suffix array and a longest common prefix array (LCP array).

4. The method according to claim 1, characterized in that: When the data to be compressed is detected to conform to a specific rule pattern, the generation parameters describing the rule pattern are extracted, a reconstruction script is constructed based on the generation parameters, and the reconstruction script is used to replace the high-dimensional character replacement to achieve parameterized compression.

5. The method according to claim 4, characterized in that: The specific rule pattern is a periodic binary alternation pattern, and the generation parameters include: image width, image height, grid size, first color value, second color value, and arrangement rule.

6. The method according to claim 1, characterized in that: The repeating patterns are sorted by their returns and then greedily selected. The returns are calculated based on the length and frequency of occurrence of the repeating patterns.

7. The method according to claim 1, characterized in that: The remaining data that was not replaced by the high-dimensional character is subjected to secondary compression, which uses an existing general compression algorithm (such as zlib).

8. A lossless data decoding method, characterized in that, include: Receive compressed data, the compressed data containing single-byte characters located in the original data area and high-dimensional characters located in the compressed symbol area; The single-byte character is output directly; The high-dimensional characters are restored to their corresponding original repeating patterns according to the dictionary mapping table, and the restoration result is output; wherein, the dictionary mapping table is generated during the compression process and stored together with the compressed data.

9. A lossless data compression system, characterized in that, include: The encoding space management module is used to divide the storage space into a raw data area and a compressed symbol area where the encoding addresses do not overlap. A repeating pattern detection module is used to identify repeating patterns in the data to be compressed; a symbol mapping module is used to map the repeating patterns to high-dimensional characters in the compressed symbol area. The replacement module is used to replace the repeating pattern with the high-dimensional character to generate compressed data.

10. The system according to claim 9, characterized in that, Also includes: The rule detection module is used to determine whether the data to be compressed conforms to a specific rule pattern; the parameter extraction module is used to extract the generation parameters describing the rule pattern. The reconstruction script generation module is used to construct a reconstruction script based on the generation parameters to achieve parameterized compression.

11. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of any one of claims 1 to 8.