Plane image compression method

By generating a task value density map and uncertainty, and combining invariant feature side channels with task consistency indication information for compression and decoding correction, the problem of key structural damage caused by planar image compression in existing technologies is solved, and readability and measurability consistency are achieved under limited bit rate.

CN121000883AActive Publication Date: 2025-11-21NANJING TECH UNIV

Patent Information

Application Number
CN202511536644.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2025-11-21
Estimated Expiration
2045-10-27

AI Technical Summary

Technical Problem

Existing image compression methods, when processing planar images containing text, tables, barcodes, or regular boundaries, result in thinner strokes, broken connectivity, decreased local contrast, curvature shift, and unstable boundary positions, affecting the performance of downstream tasks such as recognition and structured analysis.

Method used

The task value density map and its uncertainty are generated by feature extraction, quantization step size and bit allocation weight are generated, the image is compressed and transformed to generate main bit stream, sub bit stream and constraint bit stream, and topological consistency correction and constraint projection are performed at the decoding end to ensure the readability and measurability of key structures.

Benefits of technology

Under limited bitrate conditions, ensure the quality of text strokes, table lines, barcode modules, and regular geometric boundaries, avoid problems such as stroke breakage, contrast attenuation, and structural distortion, and achieve consistency between subjective visual quality and machine recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121000883A_ABST
    Figure CN121000883A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image compression, in particular to a plane image compression method, and provides the following scheme: obtaining a to-be-compressed plane image; carrying out feature extraction on the plane image to obtain image features, and calculating a task value density map and uncertainty corresponding to the task value density map; generating a quantization step size and a bit allocation weight according to the task value density map and the uncertainty, and performing compression transformation on the plane image to obtain a main code stream; further generating invariant feature side channel and task consistency indication information according to the image features, and performing compression to obtain an auxiliary code stream and a constraint code stream; and sending the main code stream, the auxiliary code stream and the constraint code stream to a decoding end. According to the method, the readability and testability of key structures such as character strokes, table grids and boundary curvatures can be guaranteed under the condition of limited code rates, and the consistency of subjective visual quality and machine recognition performance is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image compression, in particular to a planar image compression method. BACKGROUND

[0002] Most of the existing image compression methods take improving the subjective visual quality as the goal, and usually adopt optimization strategies based on pixel error or perception indicators. However, when the compression object is a planar image containing text, table, barcode or regular boundary, the processing results of the existing technology often have problems such as thinning of strokes, breaking of connected relationship, decrease of local contrast, deviation of curvature form and instability of boundary position. Further, after multiple rounds of transcoding, scaling or re-encoding, these problems will be amplified, so that although the image can still maintain high clarity in overall visual effect, the performance in downstream tasks such as recognition, structured analysis and automatic measurement will be significantly reduced, which is manifested as decrease of OCR recall rate, increase of table segmentation error rate, increase of barcode decoding failure rate and insufficient accuracy of key geometric measurement. Therefore, the existing technology has the defect of inconsistency between perception quality and task usability in the compression scene of planar images, and it is difficult to maintain machine readability and measurability while ensuring limited code rate.

[0003] To solve the above problems, the present application designs a planar image compression method. SUMMARY

[0004] The technical problem to be solved by the present application is to provide a planar image compression method to solve the problems of the prior art. The planar image compression method comprises the following steps: acquiring a planar image to be compressed; performing feature extraction on the planar image to obtain image features, and calculating a task value density map and its corresponding uncertainty; generating a quantization step and a bit allocation weight according to the task value density map and the uncertainty, performing compression transformation on the planar image to obtain a main code stream; further generating an invariant feature side channel and task consistency indication information according to the image features, and compressing to obtain a secondary code stream and a constraint code stream; and sending the main code stream, the secondary code stream and the constraint code stream to a decoding end. After the decoding end performs inverse entropy decoding and inverse quantization on the main code stream to obtain an initial reconstructed image, topological consistency correction is performed in combination with the secondary code stream, and constraint projection and connected compensation are performed using the parameter information of the constraint code stream, and finally a recovered image meeting the task consistency requirement is output. The present application can guarantee the readability and measurability of key structures such as character strokes, table grids and boundary curvature under the condition of limited code rate, and realize the consistency of subjective visual quality and machine recognition performance.

[0005] To achieve the above-mentioned purpose, the present application provides the following technical solutions:

[0006] A planar image compression method is applied to the encoding end of an image compression system, the image compression system further including a decoding end, the decoding end being used to reconstruct the image based on the bitstream output by the encoding end, the method comprising:

[0007] Obtain the planar image to be compressed;

[0008] Feature extraction is performed on the planar image to obtain image features, and a task value density map and the uncertainty corresponding to the task value density map are calculated based on the image features.

[0009] Based on the task value density map and the uncertainty corresponding to the task value density map, a quantization step size and bit allocation weight are generated to compress and transform the planar image to obtain the main bitstream;

[0010] The main bitstream is sent to the decoding end.

[0011] The image features are used to characterize the edges, strokes, connectivity, curvature, texture complexity, and confidence of text candidate regions in a planar image.

[0012] The method further includes:

[0013] Based on the image features, calculate the invariant feature side channel, wherein the invariant feature side channel includes at least one of the following: stroke skeleton and connected topology, line segments and arc segments and their geometric tolerances, text line direction statistics, key points and adjacency relationships, table grid structure descriptor and region shape;

[0014] Based on the image features, task consistency indication information is generated, wherein the task consistency indication information includes at least one of stroke width, local contrast, connectivity retention rate, curvature deviation upper limit and boundary offset upper limit;

[0015] The invariant feature side channel and the task consistency indication information are compressed to obtain the corresponding sub-bitstream and constraint bitstream;

[0016] The secondary bitstream, the constraint bitstream, and the main bitstream are sent to the decoding end.

[0017] Calculate the task value density map and the corresponding uncertainty based on the image features, including:

[0018] The planar image is divided into pixel-level and sub-block-level candidate units according to a preset multi-scale partitioning rule, and multiple sub-features are calculated for each candidate unit based on the image features to characterize the importance of downstream tasks.

[0019] The multiple sub-features are fused according to the preset mapping rules to obtain the task contribution score of each candidate unit. The task contribution score is then corrected for connectivity through spatial regularization and topological consistency constraints to generate a task value density map with the same coordinate system dimension as the planar image.

[0020] The uncertainty is obtained by performing an uncertainty assessment on the task value density map.

[0021] The uncertainty assessment includes at least one of the following:

[0022] Evaluation of result dispersion;

[0023] Stability assessment based on input perturbations, wherein the input perturbations include scaling, noise, and color perturbations;

[0024] Evaluation of the entropy value of the output probability distribution;

[0025] Observational confidence assessment of texture complexity.

[0026] The generation of quantization step size and bit allocation weights includes:

[0027] Compressed sensing is performed on the planar image to obtain the corresponding transform coefficients;

[0028] Based on the uncertainty corresponding to the task value density map, the transformation coefficients of the planar image are subjected to weighted sparse modeling to obtain a sparse signal representation for characterizing the task value weights and confidence constraints.

[0029] The sparse signal representation is projected onto the observation sequence matrix, and the reconstruction error of each candidate unit in the task value density map is calculated in the projection domain, wherein the reconstruction error is used to characterize the recovery sensitivity of the candidate unit under the condition of satisfying the global code rate constraint.

[0030] Based on the reconstruction error, the candidate units are prioritized to obtain high-priority, medium-priority, and low-priority regions.

[0031] For high-priority regions, the quantization step size is set within a preset tightening range and the corresponding bit share is allocated; for medium-priority regions, the quantization step size is set within a preset baseline range and the corresponding bit share is allocated; for low-priority regions, the quantization step size is set within a preset loosening range and the corresponding bit share is allocated.

[0032] Compressed sensing is performed on the planar image to obtain the corresponding transform coefficients, including:

[0033] The planar image is divided into multiple candidate regions according to a preset segmentation rule, and the image signal is sparsely represented in each candidate region using a sparse dictionary to obtain a set of sparse coefficients.

[0034] Construct an observation sequence matrix that satisfies the compressed sensing constraints corresponding to the planar image, and use the observation sequence matrix to perform linear observations on the sparse coefficient set to obtain the corresponding observation values;

[0035] Based on the observed values ​​and the preset sparse priors, an optimization solution is performed to obtain the approximate coefficient representation of the planar image in the sparse domain.

[0036] The approximation coefficients are represented as the transformation coefficients.

[0037] The invariant feature side channel and the task consistency indication information are compressed to obtain the corresponding sub-bitstream and constraint bitstream, including:

[0038] The invariant feature side channel is topology-preserving encoded to obtain an invariant feature symbol sequence. The topology-preserving encoding includes arranging the spatial position points of the stroke skeleton, line segments, and arc segments according to a preset scanning order; performing differential encoding on the coordinate differences between adjacent points in the invariant feature side channel; representing the connectivity between adjacent points through an adjacency list and encoding the adjacency list index; setting topology marker bits for closed regions, intersections, and bifurcation points, and embedding them into the encoding sequence in Boolean symbol form.

[0039] The task consistency indication information is segmented and encoded to obtain a task symbol sequence. The segmentation encoding includes dividing the task consistency indication information into multiple regions based on the spatial distribution of the task consistency indication information in the planar image. Each region corresponds to a parameter set. The parameter sets are fitted to obtain the range label and priority label corresponding to each parameter set.

[0040] Using the task value density map and the uncertainty corresponding to the task value density map as context conditions, entropy coding is performed on the invariant feature symbol sequence and the task symbol sequence to obtain a basic bit stream and an enhanced bit stream, wherein the basic bit stream includes skeleton topology and forced constraints, and the enhanced bit stream includes geometric refinement and proposed constraints.

[0041] The sub-bitstream and the constraint bitstream are output based on the base bitstream and the enhanced bitstream.

[0042] A planar image compression method is applied to the decoding end of an image compression system. The image compression system further includes an encoding end, which is used to compress the planar image and transmit the main bitstream, sub-bitstream, and constraint bitstream corresponding to the compression result to the decoding end. The method includes:

[0043] Receive the main bitstream, the secondary bitstream, and the constraint bitstream;

[0044] The main bitstream is subjected to anti-entropy decoding and dequantization to obtain the first reconstructed image and the residual distribution of the first reconstructed image;

[0045] The sub-stream is reconstructed, and the reconstructed invariant feature side channel is spatially aligned with the first reconstructed image, and topological consistency correction is performed to obtain the second reconstructed image.

[0046] The constrained bitstream is reconstructed, the reconstructed task consistency indication information is constrained and projected onto the residual distribution, and the second reconstructed image is connected to compensate according to the constrained projection result to obtain the restored image.

[0047] The topology consistency correction includes:

[0048] The first reconstructed image is bridged and repaired based on adjacency relationships according to the broken position of the stroke skeleton in the invariant feature side channel.

[0049] Based on the boundary position of the line segment or arc segment in the invariant feature side channel, the first reconstructed image is geometrically distorted based on parameterized tolerance.

[0050] The first reconstructed image is adjusted based on vertex difference according to the offset boundary of the closed region in the invariant feature side channel.

[0051] Compared with the prior art, the beneficial effects of this application are:

[0052] This application introduces joint modeling of task value density and uncertainty at the encoding end, and combines invariant feature side channel and task consistency indication information for topology correction and constraint projection at the decoding end. This can prioritize the restoration quality of key areas such as text strokes, table lines, barcode modules and regular geometric boundaries under limited bit rate conditions, thereby avoiding the problems of stroke breakage, contrast attenuation and structural distortion that occur in the prior art. Attached Figure Description

[0053] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0054] Figure 1 An exemplary application scenario diagram provided for an embodiment of this application;

[0055] Figure 2 A schematic flowchart illustrating a planar image compression method provided in an embodiment of this application;

[0056] Figure 3 A flowchart illustrating the method for generating sub-streams and constraint streams provided in this application embodiment;

[0057] Figure 4 This is a schematic flowchart of another planar image compression method provided in an embodiment of this application. Detailed Implementation

[0058] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0059] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0060] In actual business operations, planar images usually carry highly structured information rather than natural scenes in the general sense.

[0061] For example, bills, contracts, receipts, equipment nameplates, logistics waybills, inspection reports, forms and partial drawings, etc., in the same frame usually include fine lines and regular grids, character strokes and low-texture backgrounds, which are visually simple and statistically heterogeneous.

[0062] For institutions that conduct automated acceptance, reconciliation, filing, and auditing for these types of platforms (including but not limited to government electronic archives acceptance platforms, centralized financial bill warehousing systems, cross-border customs document image review platforms, and drug and special equipment compliance verification platforms), a joint process of restricted transmission, platform integration, and long-term retention has been established between the data collection and archiving links.

[0063] The front-end devices are diverse and widely distributed, requiring centralized backhaul via dedicated lines or encrypted tunnels with controlled bandwidth. To reduce storage and retrieval costs, uniform transcoding, sampling, and scaling are performed to meet archiving specifications. The back-end then directly performs OCR, table parsing, key field location, signature / signature verification, barcode / QR code decoding, and layout consistency verification on the images.

[0064] The core contradiction here lies not in the data task of planar images, but in how to stably guarantee the lower limit of machine-readable indicators to avoid rework and review, given that controlled bitrates and unified transcoding are unavoidable.

[0065] Understandably, the inherent characteristics of planar images determine the vulnerability of compression: the key information of the task is often concentrated in the geometry and topology, such as the minimum line width of character strokes, the local contrast of adjacent strokes, the connectivity of grids, the curvature tolerance of polylines or arcs, and the smallest resolvable unit of barcode modules. Once these high-frequency textures are flattened or broken by compression strategies driven by conventional perception as sacrificial elements, even if the subjective quality is still considered clear, the downstream recognition, localization and measurement will experience a disproportionate performance collapse.

[0066] Furthermore, real-world scenarios involve multiple rounds of uncertain processing, such as automatic sharpening and noise reduction at the shooting end, unified scaling and color gamut conversion on the platform side, and re-encoding during long-term storage. This makes it difficult for traditional code control methods that only target PSNR / SSIM to provide any guarantee for task indicators. Research and development work by those skilled in the art has long faced the dilemma of mismatch between perception indicators and task indicators.

[0067] In a representative implementation, financial institutions collect and centrally store cross-branch bills and contract attachments via mobile devices.

[0068] The image contains red overlaid text, fine grid tables and diagonal lines, localized highlights caused by reflections from embossed metal nameplates, and barcodes / QR codes with adjacent characters. Due to the limited encrypted bandwidth quotas from the office park to the data center and the centralized data entry window, the front-end needs to control the bitrate before uploading. The data entry platform, for unified retrieval and long-term archiving, will standardize the image size and re-encode it. If the conventional compression path is still used, even slight line width reduction, disruption of stroke connectivity, and a decrease in local contrast are enough to cause a drop in OCR recall, missegmentation of table cells, and misjudgment by the barcode module, thus triggering manual review, resubmission, and delays in the auditing process.

[0069] To address the objective constraints of this type of scenario, the method in this application performs online evaluation of image regions at the encoding end based on task value density and uncertainty. It uses minimum stroke width, local contrast, connectivity, and curvature deviation as rigid boundaries or soft constraints for resource allocation, prioritizing the regions most critical to downstream tasks within a given bitrate budget. At the decoding end, it combines the structure-side channel carried in the sub-bitstream and the task consistency indicator carried in the constraint bitstream to perform topological consistency correction and constraint projection restoration on the initial reconstruction result. This counteracts the uncontrollable disturbances introduced by the platform transcoding, ensuring consistency between subjective visibility and machine readability within the same bitstream.

[0070] Based on the same concept, this application method can be seamlessly transferred to scenarios such as government electronic archives acceptance, cross-border trade document review, and compliance record archiving for pharmaceutical production and equipment maintenance.

[0071] The common thread in the aforementioned scenarios is:

[0072] Uplink bandwidth is controlled or there is time slice competition; the platform must perform standardized transcoding; the value of the image is mainly carried by structured and extractable strokes / boundaries / grids; and the downstream task indicators have the minimum available line with legal or compliance constraints.

[0073] Under this constraint system, the implementation method of this application provides an engineering path for planar image compression centered on task consistency, avoiding the common defect of traditional strategies that have good perception indicators but fail in task indicators, and ensuring the readability / measurability reliability of cross-platform, cross-link and cross-time period processing without limiting specific devices, shooting postures or single-category documents.

[0074] refer to Figure 1 , Figure 1 This is an exemplary application scenario diagram provided for an embodiment of this application.

[0075] Figure 1 The application scenarios shown include the encoding end, the transmission end, and the decoding end.

[0076] The encoding end is used for feature extraction, task value density calculation, and bitstream generation of the planar image to be processed. The encoding end includes:

[0077] The data acquisition module is used to acquire the planar image to be compressed and perform preprocessing.

[0078] The feature extraction module is used to perform feature analysis on the planar image to obtain image features that characterize edges, strokes, connectivity, curvature, texture complexity, and confidence of text candidate regions, and to generate a task value density map and corresponding uncertainty based on the image features.

[0079] The encoding module is used to generate quantization step size and bit allocation weight based on the task value density map and uncertainty, compress and transform the planar image to obtain the main bitstream, and further generate invariant feature side channel and task consistency indication information. After compression, the sub-bitstream and constraint bitstream are obtained, and finally the main bitstream, sub-bitstream and constraint bitstream are output.

[0080] The transmission end includes a transmission path for transmitting the main bitstream, the secondary bitstream, and the constraint bitstream between the encoding end and the decoding end.

[0081] The decoding end is used to decode and restore the received bitstream. The decoding end includes:

[0082] The decoding module is used to perform anti-entropy decoding and dequantization on the main bitstream to obtain the initial reconstructed image, and to recover the invariant feature side channel and task consistency indication information based on the sub-bitstream and constraint bitstream;

[0083] The data restoration module is used to fuse the invariant feature side channel and task consistency indication information with the initial reconstructed image, perform topological consistency correction and constraint projection reconstruction, and obtain the target reconstructed image that meets the task consistency lower limit requirement.

[0084] Next, with reference to the accompanying drawings, a planar image compression method provided by an embodiment of this application will be further described. Figure 2 The method shown is applied to the encoding end of an image compression system, which also includes a decoding end. The decoding end is used to reconstruct the image based on the bitstream output by the encoding end. The method includes:

[0085] S1: Obtain the planar image to be compressed;

[0086] In this embodiment, the planar image to be compressed can be a document image, contract attachment, scanned form, logistics waybill, or equipment nameplate photo used for automatic processing. Their common feature is that the image contains heterogeneous elements such as text strokes, table lines, stamp edges, and low-texture backgrounds.

[0087] Those skilled in the art will understand that the source of the planar image is not limited to a scanner or mobile terminal; any two-dimensional image capable of generating structured content is applicable, and this application does not limit it.

[0088] S2: Extract features from the planar image to obtain image features, and calculate the task value density map and the uncertainty corresponding to the task value density map based on the image features;

[0089] In this embodiment, image features include, but are not limited to, edge strength, stroke energy, connectivity index, curvature variation, texture complexity, and confidence of candidate text regions. Based on these features, the image is divided into pixel-level and sub-block-level units using a multi-scale partitioning method, and their task contribution scores are calculated for each. These are then corrected using spatial regularization and topological consistency constraints to finally generate a task value density map. Furthermore, the reliability of the task value density map is estimated by combining multi-model inference results, input perturbation testing, and probability distribution entropy values, thus forming uncertainty.

[0090] S3: Based on the task value density map and the uncertainty corresponding to the task value density map, generate a quantization step size and bit allocation weight to compress and transform the planar image to obtain the main bitstream;

[0091] In this embodiment, compressed sensing processing is first performed on the planar image to obtain transform domain coefficients. Then, based on the joint distribution of task value density and uncertainty, weighted sparse modeling is performed on each candidate unit, and the reconstruction error is calculated by projecting the observation sequence matrix. According to the superposition effect of reconstruction error and task value, the region is divided into high-priority, medium-priority, and low-priority regions, corresponding to the quantization step sizes of the tightening interval, the baseline interval, and the loosening interval, respectively. Simultaneously, a higher bit share than the baseline is allocated to the high-priority region, a lower bit share than the baseline is allocated to the low-priority region, and the medium-priority region maintains an average level. This differentiated allocation strategy can stably guarantee the minimum recovery requirements of character stroke width, table connectivity, and curvature constraints within a given bit rate budget, avoiding the loss of machine readability caused by traditional perceptual compression.

[0092] S4: Send the main bitstream to the decoding end;

[0093] It is understood that the image compression system in this application specifically refers to a functional assembly formed to realize the aforementioned planar image compression and reconstruction process, including an encoding end for executing encoding method steps, a transmission path for carrying the main bitstream / sub-bitstream / constraint bitstream transmission, and a decoding end for performing decoding and constraint restoration; its physical form can be a standalone device, an edge-end integrated machine, a distributed cluster, or a cloud-edge collaborative device, or it can be implemented by a processor executing program instructions in a storage medium, or implemented through a programmable logic device / application-specific integrated circuit; the transmission path can be a wired or wireless network, a leased line or a virtual private network, and intermediate transcoding, scaling, or archiving links are allowed; the encoding end and the decoding end are only logically divided and are not limited to independent hardware units, and can be integrated with the acquisition / storage / business processing module or deployed separately as needed; the terminology is not limited to a specific brand, interface, or protocol stack, as long as it can output a constraint bitstream containing a main bitstream, a sub-bitstream carrying invariant feature side channels, and task consistency indication information, and complete reconstruction and consistency correction at the decoding end accordingly, it should be considered to fall within the application scope of the image compression system described in this application.

[0094] Before detailing the specific technical aspects of the steps, this application's embodiments need to reiterate:

[0095] To facilitate understanding of the technical logic of this application, the information organization method of planar images in engineering compression scenarios will be explained first.

[0096] In this application, the planar image is not a homogeneous signal; its effective information is highly concentrated in geometric and topological elements that can be formally characterized, such as the line width and connectivity of character strokes, the row and column relationships of grids, the boundaries and corners of condition markings, and the smallest modular unit of barcodes / QR codes.

[0097] In practical compression, if only pixel error or perceptual measurement is considered as the target, these factors are easily misjudged as high-frequency components that can be sacrificed, leading to nonlinear failures in downstream recognition and measurement performance even when visual quality is acceptable. To avoid such mismatches, this embodiment first establishes a division between a fidelity-preserving region and a distortion-tolerant region for task-related information at the encoding end. All subsequent parameter generation is constrained by this division, aiming to concentrate bitrate consumption on the key structures that determine machine readability, thereby achieving consistency between bitrate control and task metrics.

[0098] In the method logic on the encoding side, the first step is to form a joint characterization of the degree of impact on the task and the stability of the evaluation based on the local statistics and semantic cues of the image content. These two can be understood as value and confidence, respectively.

[0099] Furthermore, instead of directly using joint characterization for thresholding, it is transformed into a parameterized description of the set of permissible distortions:

[0100] For regions with high value and high confidence, quantization perturbations are only allowed under conditions that meet the following requirements: minimum stroke width, minimum local contrast, no disruption of connectivity, and controlled curvature deviation. For regions with low value or low confidence, more relaxed distortion boundaries are allowed, but structural consistency with adjacent regions is still required. The method in this application is equivalent to setting a feasible solution space for different regions, allowing subsequent quantization and bit allocation to find the best solution within the constrained solution set, thereby avoiding situations where the perception score increases but the task fails.

[0101] In addition to the main stream, this embodiment also constructs two sets of additional information to enhance recoverability after cross-platform propagation.

[0102] One set of invariant feature side channels uses a compact representation such as topology-preserving graph structure, geometric element parameters and regular lattice description to independently preserve structural information that is easily damaged in recoding, such as connection relationships, curvature shape, grid fundamental frequency and offset, and achieves spatial alignment with the main bitstream through anchor points and block-level indexes.

[0103] The other set is task consistency indication information, which segments and parameterizes thresholds such as stroke width, contrast, connectivity, curvature deviation, and boundary offset by region. Only the segment parameters, applicable scope, and priority markers are recorded, so that the hierarchy of mandatory constraints and suggested constraints can facilitate the decoding end to perform differentiated recovery.

[0104] The two sets of information are encoded by context adaptive entropy to form a sub-bitstream and a constraint bitstream, respectively. The purpose is to provide the decoding end with a verifiable and projectable structural benchmark and constraint basis without significantly increasing the total bit rate.

[0105] It is worth noting that the aforementioned process is not limited to specific shooting media, acquisition postures, or business categories. Value and confidence can be obtained through different features and criteria, including but not limited to edge strength, stroke energy, connectivity, curvature change, text candidate confidence, key point density, etc.; the granularity of segmentation parameters can also be adjusted according to storage and real-time requirements.

[0106] Those skilled in the art will understand that as long as a feasible solution space for the task-sensitive structure can be formed on the encoding side, and topological correction and constraint projection can be performed on the decoding side accordingly, the technical effect of ensuring task consistency with a finite bit rate can be achieved.

[0107] Next, the technical content of the image feature method in this application will be further elaborated.

[0108] In this embodiment, the image features are used to characterize the edges, strokes, connectivity, curvature, texture complexity, and confidence of text candidate regions of a planar image.

[0109] It is understood that feature extraction can be achieved through classical image processing procedures (including but not limited to directional filtering, structural tensor analysis, thinning / skeletonization, connected component labeling, local fitting and parameterized description), or through learning-based feature networks (including but not limited to lightweight semantic segmentation, text detection and thin-line enhancement subnetworks), or by a combination of both; the running platform can be a general-purpose processor, graphics processor, programmable logic device or special-purpose circuit, and parameter training can be completed based on public data or internal samples, which will not be elaborated here.

[0110] Next, the technical content of image feature processing in this application will be further elaborated.

[0111] It is understood that the task value density map and the uncertainty corresponding to the task value density map are not independent endpoints in this application, but rather serve as a pre-input for subsequent compressed sensing processing. This is used to characterize the differences and reliability of different regions of the image in terms of task semantics before signal projection and sparse modeling. The basic principle is as follows:

[0112] The reconstruction accuracy of compressed sensing depends on two core factors: the sparsity of the signal in the sparse domain and the resource allocation method for sparse coefficients during observation and reconstruction. If homogeneous sparse modeling and observation are applied to all image regions indiscriminately at the encoding end, the limited projection dimension and bit rate will often be evenly distributed among a large number of low-value regions, resulting in irreversible breaks and weakening of key strokes, connected boundaries, and grid lines during reconstruction.

[0113] In this embodiment, by pre-calculating the task value density map, a weighted model can be established for the sparse signal before compressed sensing projection: high-value regions are assigned higher preservation weights, and low-value regions are assigned lower preservation weights. This guides observation energy to be concentrated in key task regions when constructing the observation sequence matrix and performing sparse representation, allowing the limited observations to carry more task-related structural information. Simultaneously, in conjunction with the uncertainty map, protective redundancy can be appropriately increased in high-value but unstable evaluation regions, while reducing code rate input in low-value but stable evaluation regions. In other words, value density provides a ranking of importance, and uncertainty provides a boundary of reliability; their combination forms a weighted sparse signal representation that can be directly utilized by the compressed sensing process.

[0114] In one example, calculating the task value density map and the corresponding uncertainty based on the image features includes:

[0115] S2.1: Divide the planar image into pixel-level and sub-block-level candidate units according to the preset multi-scale division rules, and calculate multiple sub-features for each candidate unit based on the image features to characterize the importance of downstream tasks;

[0116] Specifically, structures carrying task semantics in planar images often appear at sub-pixel or pixel scales and are not naturally aligned with the block grid used in subsequent encoding. If only a single block-level granularity is used for characterization, information leakage will occur at stroke intersections and fine line bends, while pixel-level granularity is susceptible to noise and contrast fluctuations. To address this, a candidate unit set is constructed that operates in parallel at both the pixel and sub-block levels, and aliasing effects and block boundary artifacts are suppressed through multi-scale partitioning, making the characterization of fine lines and connectivity both sensitive and stable.

[0117] In this embodiment, pixel-level candidates are represented by a full-pixel grid; sub-block-level candidates are represented by a grid aligned with the coding block, with block side lengths of 8×8 or 16×16, and half-block overlap is set to alleviate boundary truncation.

[0118] For each candidate unit, sub-features are calculated: For edges and strokes, a directionally adjustable filter bank is used to obtain gradient magnitude and main direction, and high-frequency textures are filtered out by combining thin-line enhancement response and local contrast constraints; responses satisfying the thin-line assumption are refined to obtain a skeleton, and the approximate line width is estimated along the main direction using distance transformation; For connectivity, connected components are labeled in the binarized foreground, recording component area, aspect ratio, and hole count, and an undirected graph is constructed on the skeleton map to obtain endpoints, bifurcation points, and their adjacency; For curvature, broken lines and arcs are extracted from the skeleton and significant edges, and local robust fitting is used to estimate the curvature at vertices and the curvature change trend, and the parameterized description and tolerance band of suspected regular geometry are recorded; For texture complexity, the proportion of structural tensor eigenvalues, bandpass energy density, local entropy, and spectral energy concentration are calculated to distinguish low-texture backgrounds, regular repetitive textures, and irregular textures; For text candidates, the outputs of traditional candidates and lightweight detection branches on the downsampled feature map are combined to obtain candidate text line direction, line spacing, character anchor points, and region confidence. The aforementioned sub-features are output directly at the pixel level and at the sub-block level as regional statistics for subsequent fusion.

[0119] S2.2: The multiple sub-features are fused according to the preset mapping rules to obtain the task contribution score of each candidate unit. The task contribution score is then corrected for connectivity through spatial regularization and topological consistency constraints to generate a task value density map with the same coordinate system dimension as the planar image.

[0120] Specifically, the contributions of multi-source sub-features to the task are non-homogeneous:

[0121] Stroke width and local contrast are more sensitive to OCR, connectivity and grid patterns are more critical to layout analysis, and curvature and its tolerance are more important for maintaining regular geometry. If simply summarizing linearly, high-frequency textures are easily misjudged as high value or fine lines are underestimated in low-contrast backgrounds.

[0122] In this embodiment, each sub-feature is first normalized and its dynamic range is compressed. A lookup table and piecewise linear mapping are used to establish a monotonic relationship with the task relevance. For example, pixels with line width close to the minimum readable threshold are given higher weights, the risk of breakage of connected components and the density of skeleton endpoints are given an enhancement term, and the tolerance band deviation of regular geometry is given a protection term.

[0123] Furthermore, initial task contribution scores are calculated at both the pixel and sub-block levels, and edge-guided regularization is performed spatially:

[0124] The differences between adjacent units are widened at strong edges and tightened in the same text line or table line direction to avoid breakpoints along the structural direction;

[0125] Furthermore, the embodiments of this application also include topological consistency constraints, which use skeleton graphs and mesh graphs as carriers to propagate scores along the graph edges, ensuring the continuity of scores for the entire text line and the entire table line, and preventing low scores appearing only at inflection points or weak segments from breaking the structure; for barcode and QR code areas, using module size and alignment angle as priors, contribution peaks are retained at module boundaries while being moderately suppressed inside the module, in order to meet the decoding algorithm's dependence on boundary clarity.

[0126] S2.3: Perform uncertainty assessment on the task value density map to obtain the uncertainty;

[0127] Specifically, relying solely on task value density for resource allocation can lead to overconfidence in boundary conditions, low-contrast regions, or regions with similar structures, resulting in limited bitrate being allocated to regions that do not offer stable returns. The purpose of introducing uncertainty assessment is to measure the robustness of value assessment, ensuring that resource preferences are driven not only by importance but also by assessment reliability constraints, thereby maintaining a recoverable margin on the decoding side under unknown processing chains (such as scaling, re-encoding, and color transformation) and unknown task sets.

[0128] In this embodiment, the uncertainty assessment includes at least one of the following:

[0129] Evaluation of result dispersion;

[0130] Stability assessment based on input perturbations, wherein the input perturbations include scaling, noise, and color perturbations;

[0131] Evaluation of the entropy value of the output probability distribution;

[0132] Observational confidence assessment of texture complexity;

[0133] In one example, the result dispersion assessment is achieved through repeated computations under multiple models or multiple parameter settings.

[0134] Specifically, for the same planar image, different parameter combinations (such as different threshold settings, different scale filter banks, or different skeleton extraction algorithms) can be used in the feature extraction stage to generate multiple candidate task value density maps. Statistical analysis of the contribution scores of each candidate unit is then performed. If the results differ significantly among multiple models, it indicates that the value judgment of that region is unstable and should be assigned a higher degree of uncertainty; conversely, if the results are highly consistent, the uncertainty is lower. This method can reveal the risks brought about by model bias and parameter dependence, allowing bitrate allocation to avoid over-investment in regions with inconsistent models.

[0135] In another example, stability assessment based on input perturbations is accomplished by simulating potential distortions in the transmission link.

[0136] Specifically, a slight perturbation is applied to the original image, such as scaling the resolution, adding low-intensity noise, adjusting color components or brightness range, and then the task value density map is recalculated. The distribution of contribution scores before and after the perturbation is compared. If the distribution shifts significantly under the perturbation condition, it indicates that the feature representation of that region is sensitive to link perturbations, and the subsequent reconstruction uncertainty is high. If the results remain basically consistent after the perturbation, it indicates that the region has good robustness and the bitrate can be appropriately reduced during allocation.

[0137] In another example, the entropy evaluation of the output probability distribution is primarily applicable to candidate regions based on detection or classification models.

[0138] Specifically, when the text detection, table detection, or barcode detection module provides a probability output for a certain area, if the probability distribution is concentrated and the maximum value is significantly higher than the second highest value, it indicates that the model is relatively certain in its judgment of the area and the uncertainty is low; if the probability distribution is dispersed and the maximum value is close to the second highest value, it indicates that the model has ambiguity or vagueness regarding the area and should be judged as having high uncertainty.

[0139] The detection or classification model can be implemented through traditional image processing algorithms, such as rule detectors based on connected component analysis, edge enhancement and geometric constraints, or through deep learning networks, such as convolutional neural networks, recurrent neural networks or object detection frameworks based on attention mechanisms, or through lightweight semantic segmentation networks or multi-task learning networks. This application does not limit the implementation of these methods.

[0140] Those skilled in the art will understand that the choice of model can be adjusted according to the specific deployment environment, computing resources, and the accuracy requirements of the target task, as long as it can output the regional probability distribution and meet the basic requirements for entropy calculation.

[0141] In another example, the observation confidence assessment of texture complexity is calculated by analyzing local signal-to-noise ratio, local contrast, and structural tensor features.

[0142] Specifically, in low-contrast backgrounds or highly textured areas, the difference between strokes and background is insufficient, and boundary information is easily masked by noise or compression loss, thus reducing task readability. In this case, the signal-to-noise ratio and contrast ratio of the local region are calculated, and the directional features of the structure tensor are combined to determine whether the information is stable enough; if the region has complex texture and low signal-to-noise ratio, high uncertainty is assigned; if the region has simple texture and high signal-to-noise ratio, low uncertainty is assigned.

[0143] Furthermore, the aforementioned multiple evaluation results are weighted and fused, and smoothed through spatial regularization and topological consistency constraints, so that the uncertainty is continuously distributed near the structural boundary without isolated noise points or large-area unreasonable abrupt changes, thus obtaining the corresponding uncertainty.

[0144] Next, the technical content of the compression transformation method in this application will be further elaborated.

[0145] In one example, the generation of the quantization step size and bit allocation weights includes:

[0146] S3.1: Perform compressed sensing on the planar image to obtain the corresponding transformation coefficients;

[0147] Specifically, the planar image is first divided into blocks according to the grid aligned with the coding block plus half-block overlap, and based on the text / line segment direction obtained in the previous steps, a sparse representation domain matching its structure is selected for each block.

[0148] In one example, the basic principle of compressed sensing of the planar image to obtain transform coefficients is to utilize the compressibility of the signal in a specific sparse domain to approximate the recovery of high-dimensional information through low-dimensional observation, thereby reducing the amount of redundant data and improving coding efficiency.

[0149] Specifically, the planar image is first divided into several candidate regions according to a preset segmentation rule. This is done because local regions usually have stronger sparsity under a certain sparse basis, which is convenient for subsequent modeling. Within each region, the image signal is sparsely represented using a sparse dictionary (such as a wavelet dictionary, a discrete cosine dictionary, or an overcomplete dictionary obtained through training) to obtain a sparse coefficient set. Most elements in the sparse coefficient set are close to zero, and only a small number of coefficients carry the main information.

[0150] Furthermore, an observation sequence matrix that satisfies the compressed sensing constraints is constructed. The observation sequence matrix is ​​usually required to maintain low correlation with the sparse dictionary to ensure that the observation process retains sufficient information. Commonly used design methods include Gaussian random matrices, Bernoulli matrices, or structured Hadamard matrices. The sparse coefficient set is then linearly observed through the observation sequence matrix to obtain the observed values.

[0151] Furthermore, based on observations and sparse prior knowledge, an optimization method is used to recover the approximate coefficient representation of the planar image in the sparse domain. This approximate coefficient representation is then used as transform coefficients for subsequent quantization and bit allocation. In this process, sparse representation refers to reconstructing the original signal using as few non-zero coefficients as possible within a high-dimensional set of basis functions. The observation sequence matrix is ​​a linear operator that maps the high-dimensional sparse signal to a low-dimensional observation space, and the optimization method is a computational method for recovering the sparsest solution from an underdetermined system of equations. Through this process, without significantly increasing computational overhead, a low-dimensional projection approximation can be used to replace the high-dimensional original data, achieving effective compression of the planar image. Simultaneously, it provides a stable foundation of transform coefficients for differentiated resource allocation guided by subsequent task value.

[0152] S3.2: Based on the uncertainty corresponding to the task value density map, perform weighted sparse modeling on the transformation coefficients of the planar image to obtain a sparse signal representation for characterizing the task value weights and confidence constraints;

[0153] Specifically, based on task value density and uncertainty, two types of weights are assigned to each candidate unit and its corresponding frequency band:

[0154] One type is the retention weight, which is used to increase the selection priority of the corresponding atom in the unit;

[0155] Another type is the confidence boundary, which is used to limit the acceptable approximate error range of the cell.

[0156] The retention weights monotonically increase with value density, and the confidence boundaries monotonically tighten with uncertainty. Both diffuse spatially in the direction of text lines / table lines, making the constraints of the entire structure consistent. In the frequency band dimension, directional atoms near the main direction of the stroke and their harmonic bandpass atoms receive higher retention weights, while low background frequencies and high frequencies that are irrelevant to the task are downgraded.

[0157] In this embodiment, a weighted sparse prior is applied to the transformation coefficient tensor, and group structure modeling is used to fit the real geometry:

[0158] Atoms within the same skeleton neighborhood form a group, the same grid row / column forms a group, and the barcode module boundary forms a group; the activation of atoms within a group is increased by the retention weight, while that outside the group is suppressed.

[0159] In some optional implementations, the reconstruction process also includes a dual-track shutdown rule: when the approximation error of a candidate cell reaches its lower confidence boundary, further refinement of that cell is stopped, and the budget is transferred to high-value cells that have not yet reached the boundary; when the global bit rate trial reaches the reservation threshold, the process switches to a mode that performs minor refinement only on high-value, high-uncertainty cells, in order to form a sparse signal representation that is selectively retained and has a margin of suppression.

[0160] S3.3: Project the observation sequence matrix onto the sparse signal representation, and calculate the reconstruction error of each candidate unit in the task value density map in the projection domain, wherein the reconstruction error is used to characterize the recovery sensitivity of the candidate unit under the condition of satisfying the global code rate constraint.

[0161] Specifically, the weighted sparse representation is forward-projected onto the same observation sequence matrix as S3.1 to obtain simulated observations, which are then compared with actual observations. The projection residuals of the two reflect the degree to which the current representation satisfies the consistency with the observations.

[0162] In this embodiment, two stabilization methods are introduced for error calculation.

[0163] One method is residual extrapolation for iterations that stop early:

[0164] Error calculations are performed on each cell using only a few iterations, and extrapolation is performed using the slope of the error-iteration curve from the cell's history, thus avoiding a comprehensive and costly reconstruction.

[0165] Another approach is structural consistency penalty, which amplifies errors that conflict with the skeleton / mesh structure in the substream, causing subsequent priority allocation to favor units that contribute more to structural consistency.

[0166] It is understandable that there are various ways to calculate the reconstruction error of each candidate unit in the task value density map within the projection domain.

[0167] For example, a residual energy allocation method can be used, which calculates the residual energy in the projection domain based on the difference between the observed value and the reconstructed value, and then back-allocates the residual energy to the corresponding candidate cell through atomic response or basis function mapping, thereby estimating its reconstruction error;

[0168] Alternatively, an iterative reconstruction difference-based approach can be used, recording the convergence trajectory of candidate units at different iteration rounds during the sparse reconstruction process, and inferring the reconstruction error through the convergence rate and stability.

[0169] Alternatively, a task-loss approximation method can be used to correlate the residuals of candidate units in the projection domain with the deviations of task features, thereby indirectly deriving the task-related reconstruction error of the candidate unit.

[0170] S3.4: Based on the reconstruction error, the candidate units are prioritized to obtain high-priority, medium-priority, and low-priority regions;

[0171] It is understood that high-priority, medium-priority, and low-priority regions can be specifically divided by pre-set segmentation thresholds. The segmentation thresholds can be determined by a large number of experiments conducted by those skilled in the art. The specific threshold can be a range, and the reconstruction error within the corresponding range is the corresponding priority. Alternatively, it can be multiple specific values, and the corresponding priority is determined by the numerical magnitude relationship. This application does not impose any further limitations here.

[0172] S3.5: For high-priority regions, set the quantization step size within a preset tightening range and allocate the corresponding bit share; for medium-priority regions, set the quantization step size within a preset baseline range and allocate the corresponding bit share; for low-priority regions, set the quantization step size within a preset loosening range and allocate the corresponding bit share.

[0173] Specifically, a reference quantization step size and upper and lower limit range are preset for each frequency band, and the corresponding range is selected according to the regional priority: a tightened range is used for high priority, a relaxed range is used for low priority, and a reference range is used for medium priority.

[0174] In this embodiment, the interval boundary is jointly determined by value density and uncertainty. The higher the uncertainty, the lower the upper bound of the tightened interval and the higher the lower bound of the loosened interval, thus increasing protection redundancy. Bit shares are allocated on a region-by-region basis, first satisfying the basic needs of high-priority regions, and then supplementing according to decreasing sensitivity. Within a region, a linear or piecewise gradual step transition is adopted along the text line / table line direction to avoid quantization steps along the structural direction.

[0175] Furthermore, the quantization step size and bit share settings can be linked with the entropy model to perform rapid evaluation of trial quantization-trial entropy coding for each region. If the cumulative bit rate exceeds the budget, priority is given to reclaiming the supplementary allocation share for low-priority regions; if there is still a budget surplus, a small-scale refinement quota is added to the high-priority regions with high uncertainty. To ensure consistency across block boundaries, a nearest-tightest merging strategy is adopted for the step size of adjacent blocks within the boundary band; to ensure the executability of structural constraints, a structural guard threshold is added within the tightening interval.

[0176] If trial quantization results in the minimum stroke width or minimum contrast prediction falling below the threshold, the step size is automatically reduced until the threshold is met or the bit share of that region is increased with a very small increment.

[0177] In one example, the method of this application also includes:

[0178] S5: Generate a sub-bitstream and a constraint bitstream based on the image features.

[0179] It is understandable that after generating the sub-stream and constraint stream, the sub-stream, constraint stream, and main stream can be sent to the decoding end.

[0180] refer to Figure 3 , Figure 3 This is a flowchart illustrating the method for generating sub-streams and constraint streams according to an embodiment of this application.

[0181] In one example, the specific steps of S5 are as follows:

[0182] S5.1: Calculate the invariant feature side channel based on the image features, wherein the invariant feature side channel includes at least one of the following: stroke skeleton and connected topology, line segments and arc segments and their geometric tolerances, text line direction statistics, key points and adjacency relationships, table grid structure descriptor and region shape;

[0183] Specifically, to provide a structural benchmark for the reconstructed results after multiple rounds of transcoding or scaling at the decoding end, it is necessary to separate low-dimensional structural quantities from the planar image that are stable for task judgment and weakly correlated with perceptual details. These low-dimensional structural quantities remain relatively invariant under resampling and mild denoising, and can serve as anchor points for topological correction and constrained projection. Therefore, descriptors related to boundaries, skeletons, regular geometry, and typographic order are extracted as independent side channels and reversibly aligned with the main bitstream to prioritize the recoverability of character connectivity, grid alignment, and curvature morphology under limited bitrates.

[0184] In this embodiment, skeleton nodes are numbered sequentially according to the space-filling curve, node coordinates are quantized and meshed relative to block anchor points, and adjacency relationships are recorded using a sparse adjacency list and Boolean markers for endpoints / bifuzzy points. Line segments and arc segments are obtained through robust fitting of edge point sets. Straight lines are parameterized by origin-direction-length, and arcs are parameterized by center direction-radius-arc length. The tolerance band is estimated based on local curvature distribution and imaging scale. Text line direction statistics are obtained by jointly calculating the direction histogram and projection spectrum, outputting the main direction, line spacing, and direction bucket index. After stability screening, key points are used to establish adjacency relationships using k-nearest neighbors and angle consistency, and cross-structure short circuits are eliminated. The table grid structure is represented by row and column fundamental frequencies and offset sequences, and sparse masking of intersection points. The region shape is represented by polygon chain code and vertex difference.

[0185] S5.2: Based on the image features, generate task consistency indication information, wherein the task consistency indication information includes at least one of stroke width, local contrast, connectivity retention rate, curvature deviation upper limit, and boundary offset upper limit;

[0186] Specifically, individual structural anchor points can only be used for morphological alignment and are difficult to constrain the impact of quantization perturbations on the readable / measurable thresholds. In order to ensure that the decoding end can maintain the lower limit of task performance under unknown processing chains, the acceptable distortion range needs to be solidified into a projectable constraint set in a parameterized form, covering key thresholds such as line width, contrast, connectivity, curvature and boundary position, so as to provide a clear feasible area for local adjustments during the reconstruction stage.

[0187] In this embodiment, the stroke width threshold is obtained statistically from the distance transformation from the skeleton centerline to the boundary, and a minimum retention value is given for different font sizes / line width segments in a piecewise constant or piecewise linear form; the local contrast threshold is output statistically based on the robust quantile difference and the contrast within the direction window to avoid the influence of global brightness drift; the connectivity retention rate is calculated based on the edge set of the skeleton graph and the risk of weak contrast breakpoints, and is defined as the minimum continuous proportion that the structural edges should maintain after reconstruction, and a priority order is given for bridging repair positions; the curvature deviation upper limit is determined jointly by the tolerance band of the arc fitting and the sampling scale, and a near-zero tolerance is given for the equivalent curvature of straight line segments, and a tolerance upper limit that increases with the radius is given for small radius arc segments; the boundary offset upper limit is represented by the vertex offset boundary equivalent to the Hausdorff distance, and tighter offset boundaries are given segmentally on the long side to control the cumulative error. The spatial allocation of parameters is performed on a structural region basis: values ​​are assigned to text lines, table lines are assigned by row / column segments, and isolated graphics are assigned by shape region, with an applicable range mask and priority marker attached.

[0188] S5.3: Compress the invariant feature side channel and the task consistency indication information to obtain the corresponding sub-bitstream and constraint bitstream;

[0189] Specifically, storing structures and constraints in pixel-by-pixel form would create an unnecessary burden and make it difficult to maintain topological invariance. Therefore, a joint compression method combining topological preservation and parametric modeling is adopted to ensure that the symbol sequence is sufficiently deredundant before entropy encoding and to record spatial anchoring information with the main bitstream in the bitstream, ensuring that the decoding end can access it randomly and use it incrementally.

[0190] In this embodiment, the invariant feature side channel uses three types of encoding:

[0191] One approach is graph structure encoding, where skeleton nodes are sorted according to the space-filling curve, node coordinates are differentially expressed with block anchor points and encoded in variable length, the adjacency list is recorded using node index differences, and endpoints / forks / loops are marked with Boolean bits.

[0192] The second method is geometric element encoding, where line / arc parameters are orthogonalized and then component quantized. Adjacent element parameters are differentiated and run-length encoding is used for repetitive patterns. Tolerance band quantization is used to convert the data into level codewords.

[0193] Thirdly, grid and text direction encoding are used. The row and column fundamental frequencies and offset sequences are encoded using prediction residuals. The intersection sparse mask is encoded using a combination of bit plane and run-length encoding. The text direction bucket index and row spacing are encoded using entropy.

[0194] Furthermore, the task consistency indication information is encoded using a segmented model:

[0195] Segmented boundary indexes, parameter vectors, applicable masks, and priority markers are written by region. Boundary trajectories are represented by chain codes or polygon vertex differences. Parameter vectors are rotated and quantized according to dimensional correlation. Forced constraints and suggested constraints are entered into the base layer and enhancement layer, respectively.

[0196] In one example, the invariant feature side channel and the task consistency indication information are compressed to obtain the corresponding sub-bitstream and constraint bitstream, including:

[0197] The invariant feature side channel is topology-preserving encoded to obtain an invariant feature symbol sequence. The topology-preserving encoding includes arranging the spatial position points of the stroke skeleton, line segments, and arc segments according to a preset scanning order; performing differential encoding on the coordinate differences between adjacent points in the invariant feature side channel; representing the connectivity between adjacent points through an adjacency list and encoding the adjacency list index; setting topology marker bits for closed regions, intersections, and bifurcation points, and embedding them into the encoding sequence in Boolean symbol form.

[0198] The task consistency indication information is segmented and encoded to obtain a task symbol sequence. The segmentation encoding includes dividing the task consistency indication information into multiple regions based on the spatial distribution of the task consistency indication information in the planar image. Each region corresponds to a parameter set. The parameter sets are fitted to obtain the range label and priority label corresponding to each parameter set.

[0199] Using the task value density map and the uncertainty corresponding to the task value density map as context conditions, entropy coding is performed on the invariant feature symbol sequence and the task symbol sequence to obtain a basic bit stream and an enhanced bit stream, wherein the basic bit stream includes skeleton topology and forced constraints, and the enhanced bit stream includes geometric refinement and proposed constraints.

[0200] The sub-bitstream and the constraint bitstream are output based on the base bitstream and the enhanced bitstream.

[0201] Next, with reference to the accompanying drawings, another planar image compression method provided in this application will be further described. Figure 4 The method shown is applied to the decoding end of an image compression system, which also includes an encoding end. The encoding end is used to compress the planar image and transmit the main bitstream, sub-bitstream, and constraint bitstream corresponding to the compression result to the decoding end. The method includes:

[0202] A1: Receive the main bitstream, the secondary bitstream, and the constraint bitstream;

[0203] Specifically, upon receiving the data, the encapsulation headers of the three bitstreams are parsed to read the version identifier, block-level index, spatial anchor point, scale level, random access index, and integrity verification field. Out-of-order reordering and packet loss detection are performed based on timestamps or sequence numbers. If an enhancement layer is detected as missing, only the base layer is used in subsequent processes. The header information of the main bitstream is used to recover regionalized quantization parameters, context model numbers, and block grid alignment relationships. For the secondary bitstreams, spatial indices of structural symbols such as skeletons / geometry / grids are recovered. For the constrained bitstreams, the applicable range mask for segmentation parameters and mandatory / suggested priority markers are recovered, and a one-to-one mapping table with the main bitstream block grid is established.

[0204] A2: Perform anti-entropy decoding and inverse quantization on the main bitstream to obtain the first reconstructed image and the residual distribution of the first reconstructed image;

[0205] Specifically, context-adaptive inverse entropy decoding is performed on the main bitstream according to the segment index to restore the transform domain symbol and regionalized quantization step size. Inverse quantization and inverse transformation are then performed based on the block grid and scale level to obtain the initial reconstruction result. To characterize the local recoverability differences caused by quantization and inverse transformation, the pixel / sub-block level residual distribution is estimated based on the decoded quantization step size, coefficient amplitude, and inter-block boundary continuity. The residual distribution records the locally allowed correction amplitude and direction preference (such as easier repair along the skeleton normal direction) in the form of a repairable margin map.

[0206] In this embodiment, the inverse quantization parameters are loaded with boundary information of three types of intervals: tightened, reference, and relaxed, according to regional priority. A transition band is set at the block boundary, and the quantization steps are smoothed by linear or piecewise interpolation. To reduce ringing and block effects, a light artifact removal preprocessing with structure preservation is adopted without destroying the frequency band energy distribution. Equalization filtering is applied only in low-texture areas, and the original frequency band relationship is maintained near fine lines and boundaries. At the same time, the difference before and after processing is written into the residual distribution as the preprocessing residual, which provides a reference for subsequent constrained projection.

[0207] A3: Reconstruct the sub-stream, spatially align the reconstructed invariant feature side channel with the first reconstructed image, and perform topological consistency correction to obtain the second reconstructed image;

[0208] Specifically, after the sub-stream is de-entropy decoded, the coordinates and adjacency relationships of skeleton nodes, parameters of line segments and arc segments and their tolerance bands, statistics of text line directions and line spacing, fundamental frequencies and offset sequences of table rows and columns, sparse masks of intersection points, key points and their adjacency graphs, etc. are reconstructed. Then, based on the block index, intra-block offset, scale level and spatial anchor point, the above structural quantities are aligned with the first reconstructed image. During the alignment process, cross-block connectors are used to splice segments into chains for long cross-block structures, and sub-pixel interpolation is performed on diagonal structures according to the direction vector.

[0209] In one example, the topology consistency correction includes:

[0210] The first reconstructed image is bridged and repaired based on adjacency relationships according to the broken position of the stroke skeleton in the invariant feature side channel.

[0211] Based on the boundary position of the line segment or arc segment in the invariant feature side channel, the first reconstructed image is geometrically distorted based on parameterized tolerance.

[0212] The first reconstructed image is adjusted based on vertex difference according to the offset boundary of the closed region in the invariant feature side channel.

[0213] It is understood that the topological consistency correction can be implemented based on existing image processing and computer vision methods. For example, when there are broken positions in the stroke skeleton recorded in the invariant feature side channel, the broken areas can be connected and compensated at the pixel level in the first reconstructed image by using the adjacency relationship of the skeleton nodes and the shortest path connection strategy, combined with thinning algorithms or morphological bridging methods, to restore the continuity of the strokes. When the invariant feature side channel provides the boundary positions of line segments or arc segments, parametric fitting methods can be used, combined with a preset geometric tolerance range, to reverse the offset boundaries in the first reconstructed image, so that the edge shape is consistent with the geometric model in the side channel. When the invariant feature side channel contains the offset boundaries of closed regions, the boundary positions can be corrected vertex by vertex using polygon vertex difference representation and chain code matching methods to ensure that the overall shape and structure of the closed regions remain stable.

[0214] These processing methods are all common topology preservation and geometric correction techniques in this field, and can be accomplished through existing image inpainting, contour adjustment, or graph-based optimization methods. They will not be elaborated upon here.

[0215] A4: Reconstruct the constrained bitstream, perform constrained projection on the reconstructed task consistency indication information and the residual distribution, and perform connectivity compensation on the second reconstructed image based on the constrained projection result to obtain the restored image;

[0216] Specifically, after decoding the constrained bitstream, the lower limits of stroke width, local contrast, connectivity preservation, curvature deviation, and boundary offset, along with their applicable masks and priorities, are obtained for each region. This parameter set is then combined with the current residual distribution to construct a feasible solution description for each target region: the allowed adjustment directions for pixels, the allowed amplitude limits, and the structural thresholds that must be satisfied. Constrained projection is performed region by region.

[0217] In stroke regions, restricted thickening or thinning is performed along the skeleton normal until the lower limit of line width is reached; in low-contrast regions, contrast stretching is limited to a local window while maintaining the background mean; in curvature-constrained regions, the offset of boundary points is constrained to ensure the equivalent curvature does not exceed the limit; in boundary position-constrained regions, vertex displacement is limited to not exceed the offset upper limit. Each local adjustment is based on minimizing the increment of the residual distribution, and the adjusted difference is written back to the projected residual.

[0218] In this embodiment, connectivity compensation is performed on the graph structure according to connectivity preservation rate and bridging priority:

[0219] For edges in the skeleton graph that fall below the target retention rate, constrained morphological bridging is performed along the shortest connection path. The bridging strength is determined by both constraint priority and residual margin. For table grid breakpoints, intersection backfilling is performed after alignment based on row / column fundamental frequencies, and differential smoothing is applied at the boundaries to avoid step jumps. For barcode / QR code boundaries, fine-grained contrast stretching is preferentially performed at module boundaries, and excessive modification within modules is prohibited to maintain decoding robustness. After completing region-level connectivity compensation, a consistency marker map is generated. The mandatory constraint regions are verified region by region. If the threshold is still not met and a corresponding fragment exists in the enhancement layer, local backfilling is triggered and projection is performed again until the mandatory constraint passes verification or the enhancement layer is exhausted.

[0220] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A planar image compression method, applied to the encoding end of an image compression system, characterized in that, The image compression system further includes a decoding end, which is used to reconstruct the image based on the bitstream output by the encoding end. The method includes: Obtain the planar image to be compressed; Feature extraction is performed on the planar image to obtain image features, and a task value density map and the uncertainty corresponding to the task value density map are calculated based on the image features. Based on the task value density map and the uncertainty corresponding to the task value density map, a quantization step size and bit allocation weight are generated to compress and transform the planar image to obtain the main bitstream; The main bitstream is sent to the decoding end.

2. The planar image compression method according to claim 1, characterized in that, The image features are used to characterize the edges, strokes, connectivity, curvature, texture complexity, and confidence of text candidate regions in a planar image.

3. The planar image compression method according to claim 1, characterized in that, The method further includes: Based on the image features, calculate the invariant feature side channel, wherein the invariant feature side channel includes at least one of the following: stroke skeleton and connected topology, line segments and arc segments and their geometric tolerances, text line direction statistics, key points and adjacency relationships, table grid structure descriptor and region shape; Based on the image features, task consistency indication information is generated, wherein the task consistency indication information includes at least one of stroke width, local contrast, connectivity retention rate, curvature deviation upper limit and boundary offset upper limit; The invariant feature side channel and the task consistency indication information are compressed to obtain the corresponding sub-bitstream and constraint bitstream; The secondary bitstream, the constraint bitstream, and the main bitstream are sent to the decoding end.

4. The planar image compression method according to claim 1, characterized in that, Calculate the task value density map and the corresponding uncertainty based on the image features, including: The planar image is divided into pixel-level and sub-block-level candidate units according to a preset multi-scale partitioning rule, and multiple sub-features are calculated for each candidate unit based on the image features to characterize the importance of downstream tasks. The multiple sub-features are fused according to the preset mapping rules to obtain the task contribution score of each candidate unit. The task contribution score is then corrected for connectivity through spatial regularization and topological consistency constraints to generate a task value density map with the same coordinate system dimension as the planar image. The uncertainty is obtained by performing an uncertainty assessment on the task value density map.

5. The planar image compression method according to claim 4, characterized in that, The uncertainty assessment includes at least one of the following: Evaluation of result dispersion; Stability assessment based on input perturbations, wherein the input perturbations include scaling, noise, and color perturbations; Evaluation of the entropy value of the output probability distribution; Observational confidence assessment of texture complexity.

6. The planar image compression method according to claim 1, characterized in that, The generation of quantization step size and bit allocation weights includes: Compressed sensing is performed on the planar image to obtain the corresponding transform coefficients; Based on the uncertainty corresponding to the task value density map, the transformation coefficients of the planar image are subjected to weighted sparse modeling to obtain a sparse signal representation for characterizing the task value weights and confidence constraints. The sparse signal representation is projected onto the observation sequence matrix, and the reconstruction error of each candidate unit in the task value density map is calculated in the projection domain, wherein the reconstruction error is used to characterize the recovery sensitivity of the candidate unit under the condition of satisfying the global code rate constraint. Based on the reconstruction error, the candidate units are prioritized to obtain high-priority, medium-priority, and low-priority regions. For high-priority regions, the quantization step size is set within a preset tightening range and the corresponding bit share is allocated; for medium-priority regions, the quantization step size is set within a preset baseline range and the corresponding bit share is allocated; for low-priority regions, the quantization step size is set within a preset loosening range and the corresponding bit share is allocated.

7. The planar image compression method according to claim 6, characterized in that, Compressed sensing is performed on the planar image to obtain the corresponding transform coefficients, including: The planar image is divided into multiple candidate regions according to a preset segmentation rule, and the image signal is sparsely represented in each candidate region using a sparse dictionary to obtain a set of sparse coefficients. Construct an observation sequence matrix that satisfies the compressed sensing constraints corresponding to the planar image, and use the observation sequence matrix to perform linear observations on the sparse coefficient set to obtain the corresponding observation values; Based on the observed values ​​and the preset sparse priors, an optimization solution is performed to obtain the approximate coefficient representation of the planar image in the sparse domain. The approximation coefficients are represented as the transformation coefficients.

8. The planar image compression method according to claim 3, characterized in that, The invariant feature side channel and the task consistency indication information are compressed to obtain the corresponding sub-bitstream and constraint bitstream, including: The invariant feature side channel is topology-preserving encoded to obtain an invariant feature symbol sequence. The topology-preserving encoding includes arranging the spatial position points of the stroke skeleton, line segments, and arc segments according to a preset scanning order; performing differential encoding on the coordinate differences between adjacent points in the invariant feature side channel; representing the connectivity between adjacent points through an adjacency list and encoding the adjacency list index; setting topology marker bits for closed regions, intersections, and bifurcation points, and embedding them into the encoding sequence in Boolean symbol form. The task consistency indication information is segmented and encoded to obtain a task symbol sequence. The segmentation encoding includes dividing the task consistency indication information into multiple regions based on the spatial distribution of the task consistency indication information in the planar image. Each region corresponds to a parameter set. The parameter sets are fitted to obtain the range label and priority label corresponding to each parameter set. Using the task value density map and the uncertainty corresponding to the task value density map as context conditions, entropy coding is performed on the invariant feature symbol sequence and the task symbol sequence to obtain a basic bit stream and an enhanced bit stream, wherein the basic bit stream includes skeleton topology and forced constraints, and the enhanced bit stream includes geometric refinement and proposed constraints. The sub-bitstream and the constraint bitstream are output based on the base bitstream and the enhanced bitstream.

9. A planar image compression method, applied to the decoding end of an image compression system, characterized in that, The image compression system further includes an encoding end, which is used to compress the planar image and transmit the main bitstream, sub-bitstream, and constraint bitstream corresponding to the compression result to the decoding end. The method includes: Receive the main bitstream, the secondary bitstream, and the constraint bitstream; The main bitstream is subjected to anti-entropy decoding and dequantization to obtain the first reconstructed image and the residual distribution of the first reconstructed image; The sub-stream is reconstructed, and the reconstructed invariant feature side channel is spatially aligned with the first reconstructed image, and topological consistency correction is performed to obtain the second reconstructed image. The constrained bitstream is reconstructed, the reconstructed task consistency indication information is constrained and projected onto the residual distribution, and the second reconstructed image is connected to compensate according to the constrained projection result to obtain the restored image.

10. The planar image compression method according to claim 9, characterized in that, The topology consistency correction includes: The first reconstructed image is bridged and repaired based on adjacency relationships according to the broken position of the stroke skeleton in the invariant feature side channel. Based on the boundary position of the line segment or arc segment in the invariant feature side channel, the first reconstructed image is geometrically distorted based on parameterized tolerance. The first reconstructed image is adjusted based on vertex difference according to the offset boundary of the closed region in the invariant feature side channel.

Citation Information

Patent Citations

  • Learning image compression method and device for image sparse mask window attention

    CN118368431A

  • Machine and human vision-oriented image coding and decoding method and compression method

    CN119180874A

  • Multi-thread picture compression method and device and storage medium

    CN120676162A

  • Machine print, hand print, and signature discrimination

    US20150347836A1

Cited By

  • Multi-mode communication method, system and equipment for fusion terminal

    CN121463107A