PDF (Portable Document Format) file browsing method based on image compression and rapid decompression technology

By hierarchical archive and encapsulation compression of PDF files on the source side, and reverse decompression on the sink side to generate pre-browsing files, the problems of slow transmission and insufficient browsing fluency caused by the large amount of PDF file image data are solved, and a more efficient data transmission and browsing experience is achieved.

CN120123296AActive Publication Date: 2025-06-10BEIJING GUANGLIANDA YUNTU DREAM TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510608656.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-10
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

In the prior art, the large amount of image data of PDF files leads to slow transmission, and it is difficult to balance efficiency and quality for compression and decompression, resulting in insufficient data transmission efficiency and browsing fluency.

Method used

By calling the compression processing module on the source side for forward activation, scanning the PDF file for file data interpretation, hierarchical archive and encapsulation based on the data value and hot and cold coefficient, decoupling and compression are performed, and compressed files are determined. Then, the compression processing module is called on the sink for reverse activation, and reverse decompression based on compression logic is performed. Pre-browsing files are generated on the display interface, and the decompression constraints are used to pre-decompression of high-value data and progressive decompression of image data.

Benefits of technology

It has achieved the improvement of data transmission efficiency and user browsing experience, solved the problems of slow transmission and insufficient browsing fluency, and achieved more efficient compression and decompression effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123296A_ABST
    Figure CN120123296A_ABST
Patent Text Reader

Abstract

The invention discloses a PDF (Portable Document Format) file browsing method based on an image compression and fast decompression technology, and relates to the technical field related to data processing, the method comprises the following steps: at an information source end, calling a compression processing module and performing forward activation; scanning the PDF file, carrying out file data interpretation, carrying out hierarchical archiving and packaging according to a data value and a cold and hot coefficient, determining a plurality of packaging blocks, executing decoupling compression, and determining a compressed file; and calling a compression processing module, performing reverse activation, executing reverse decompression based on compression logic, generating a pre-browsed file on a display interface, and taking pre-decompression of high-value data and progressive decompression of image data as decompression constraint conditions. The technical problems that in the prior art, due to the fact that the PDF file image data size is large, transmission is slow, compression and decompression are difficult to balance efficiency and quality, and data transmission efficiency and browsing fluency are insufficient are solved, and the technical effect of improving data transmission efficiency and user browsing experience is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a PDF file browsing method based on image compression and fast decompression technology. Background Art

[0002] Due to its cross-platform nature, format stability, and content integrity, PDF files have become a widely used electronic document format in fields such as office work, education, and publishing. However, PDF files usually contain a large amount of text and image data, and face many challenges during data transmission and browsing. On the one hand, traditional file compression technologies, such as ZIP, RAR, etc., have limitations in processing PDF files. Although they can reduce the file size to a certain extent, for the text data in PDF files, they fail to fully exploit the redundant information at the semantic and character levels and cannot achieve deep optimization compression. For image data processing, they cannot effectively distinguish geometric structures and texture information and are difficult to achieve efficient compression, resulting in the compression ratio and compression efficiency being unable to meet the transmission and storage requirements. On the other hand, when browsing PDF files, existing decompression technologies cannot perform differential processing according to data value and user browsing needs, and the entire file needs to be completely decompressed before the content can be presented. This not only consumes a large amount of time and computing resources but also causes users to be unable to quickly obtain key information at the initial stage of browsing. Especially in the case of mobile devices or poor network environments, the browsing experience is severely limited; at the same time, it affects the data transmission efficiency and browsing response speed of PDF files.

[0003] In the current related technologies, there are technical problems such as slow transmission due to the large amount of image data in PDF files, and it is difficult to balance efficiency and quality in compression and decompression, resulting in insufficient data transmission efficiency and browsing fluency. Summary of the Invention

[0004] By providing a PDF file browsing method based on image compression and fast decompression technology, this application solves the technical problems in the prior art, such as slow transmission due to the large amount of image data in PDF files, and it is difficult to balance efficiency and quality in compression and decompression, resulting in insufficient data transmission efficiency and browsing fluency, and achieves the technical effect of improving data transmission efficiency and user browsing experience.

[0005] This application provides a PDF file browsing method based on image compression and fast decompression technology. The method includes: at the source end, calling the compression processing module and performing forward activation; scanning the PDF file, performing file data interpretation, hierarchically archiving and encapsulating the file data according to data value and cold-hot coefficient, determining multiple encapsulated blocks and performing decoupled compression, and determining the compressed file, where the decoupled compression is semantic-character two-way compression for text data and geometric-texture two-way compression for image data; as the compressed file is received at the sink end, calling the compression processing module and performing reverse activation, performing reverse decompression based on the compression logic, and generating a pre-view file on the display interface, where the pre-decompression of high-value data and the progressive decompression of image data are used as decompression constraint conditions.

[0006] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology further performs the following processing: using data archiving and encapsulation as the first pre-processing node and adaptive loss compression based on multi-threading as the second compression processing node to construct the compression processing module; deploying the compression processing module on a cloud processor; at the physical data end, calling and activating the compression processing module from the cloud processor.

[0007] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology further performs the following processing: performing forward learning on the compression processing module to determine the first compression logic; performing reverse learning on the compression processing module to determine the second decompression logic; and performing directional activation on the compression processing module based on the first compression logic and the second decompression logic with the file status as the standard.

[0008] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology further performs the following processing: constructing a multi-threaded compression port according to the redundancy ratio-compression lossy coefficient, where the redundancy ratio is determined by balancing data value and cold-hot coefficient; for the multi-threaded compression port, using format splitting-two-way decoupled parallel compression as the processing logic to perform compression learning on each thread compression port; and using the multi-threaded compression port after parallel integrated learning as the second compression processing node.

[0009] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology further performs the following processing: calling the compression processing module to the physical data end; temporarily determining the end type of the physical data end by identifying the file status of the PDF file; and determining the activation direction of the compression processing module according to the end type, including forward activation and reverse activation, where forward activation performs compression processing and reverse activation performs decompression processing.

[0010] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology further performs the following processing: If the file status is the compressed state, the end type is identified as the sink end, and decompression is used as the processing method; if the file status is the uncompressed state, the end type is identified as the source end, and compression is used as the processing method.

[0011] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology further performs the following processing: According to the first pre-processing node, through file data interpretation, the content architecture system of the PDF file is determined; for the content architecture system, primary hierarchical archiving is performed based on data value to determine the first archiving result; secondary hierarchical archiving is performed based on the cold-hot coefficient to determine the second archiving result; archiving reorganization under weighted calculation is performed on the first archiving result and the second archiving result to determine multiple archiving layers and perform data encapsulation to determine multiple encapsulation blocks, where each encapsulation block corresponds to a redundancy ratio.

[0012] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology further performs the following processing: Determine a first encapsulation block, where the first encapsulation block is any one of the multiple encapsulation blocks; for the first encapsulation block, determine a first redundancy coefficient based on the archiving layer of the first archiving result; determine a second redundancy coefficient based on the archiving layer of the second archiving result; determine a first redundancy ratio according to the first redundancy coefficient and the second redundancy coefficient.

[0013] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology further performs the following processing: Import the multiple encapsulation blocks into the second compression processing node, activate the multi-threaded compression port by identifying the redundancy ratio, and import the target compression port, where the target compression port includes at least one; the target compression port identifies the imported encapsulation block, performs identification of the data structure within the encapsulation block to determine the first text data and the second image data; performs semantic-character decoupling on the first text data to determine the first compression round trip; performs geometric-texture decoupling on the second image data to determine the second compression round trip; performs concurrent compression processing according to the first compression round trip and the second compression round trip to determine the compressed encapsulation block; integrates the compressed encapsulation blocks output by each target compression port as the compressed file.

[0014] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology further performs the following processing: perform pre-decompression deployment according to the data value gradient to determine the first decompression condition; use progressive decompression as the image data decompression method to determine the second decompression condition; reverse-activate the compression processing module, introduce the first decompression condition and the second decompression condition, perform decompression processing on the compressed file, and perform progressive restoration of the interface of the PDF file in the display interface execution timing to determine the pre-viewed file.

[0015] It is intended to propose a PDF file browsing method based on image compression and fast decompression technology through this application. At the source end, call the compression processing module and perform forward activation; scan the PDF file, perform file data interpretation, perform hierarchical archiving and encapsulation according to the data value and cold-hot coefficient, determine multiple encapsulated blocks and perform decoupled compression to determine the compressed file; call the compression processing module and perform reverse activation, perform reverse decompression based on the compression logic, generate a pre-viewed file on the display interface, and use the pre-decompression of high-value data and the progressive decompression of image data as the decompression constraint conditions. This solves the technical problems in the prior art that the large amount of image data in PDF files leads to slow transmission, and it is difficult to balance efficiency and quality in compression and decompression, resulting in insufficient data transmission efficiency and browsing fluency, and achieves the technical effect of improving data transmission efficiency and user browsing experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings of the embodiments of the present invention will be briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the operations before or below do not necessarily need to be executed precisely in sequence. On the contrary, according to needs, they can be executed in reverse order or processed simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.

[0017] Figure 1 It is a schematic flowchart of a PDF file browsing method based on image compression and fast decompression technology provided by an embodiment of the present application.

[0018] Figure 2 It is a schematic flowchart of constructing a compression processing module in a PDF file browsing method based on image compression and fast decompression technology provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The above description is only an overview of the technical solutions of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specifically lists the detailed implementation manners of this application.

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0021] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict. The terms "first / second" merely distinguish similar objects and do not represent a specific order for the objects. The terms "comprising" and "having", and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, product, or server that includes a series of steps does not necessarily have to be limited to those steps clearly listed, but may include other steps not clearly listed or inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application.

[0022] The embodiments of this application provide a PDF file browsing method based on image compression and fast decompression technology, as Figure 1 shown. The method includes: Step S100, at the source end, call the compression processing module and perform forward activation.

[0023] Preferably, the source end, i.e., the source of information, refers to the end that generates or owns the original PDF file and needs to process it for transmission or storage. For example, when a user creates a PDF file containing a large number of images and texts on their own computer, this computer is the source end. Or, if a server stores PDF files to be sent to other devices, the server is regarded as the source end. Call the compression processing module and perform forward activation. That is, at the source end, start the compression processing module through specific instructions and interfaces to make it start working. For example, in PDF processing software, when the user clicks "Compress File", corresponding code is executed to call the compression processing module and perform data transmission processing on the currently open PDF file data. Among them, the compression processing module is a functional module used to compress PDF file data and can perform compression processing on different data types (such as text data, image data, etc.) in the PDF file. For example, use semantic analysis and character compression to process text data, and use geometric feature extraction and texture compression to process image data. Forward activation means being consistent with the normal process direction of compression processing, that is, starting from the original uncompressed PDF file data, passing through the compression processing module to execute the processing logic, and finally obtaining the compressed file.

[0024] Further, step S100 further includes step S110, using data archiving and encapsulation as the first pre-processing node, using adaptive loss compression based on multi-threading as the second compression processing node, and constructing the compression processing module; step S120, deploying the compression processing module on a cloud processor; step S130, at the physical data end, call and activate the compression processing module from the cloud processor.

[0025] Preferably, complex PDF file data includes various types such as text, images, and charts, and has different data values and cold-hot coefficients (i.e., data heat or usage frequency). Use data archiving and encapsulation as the first pre-processing node. Data archiving and encapsulation means interpreting the PDF file data, analyzing its data characteristics and attributes, hierarchically dividing the file data according to the data value and cold-hot coefficient, and then separately encapsulating the data at different levels to form multiple encapsulated blocks. Specifically, first comprehensively analyze the data in the PDF file, identify different types of data such as text, images, graphics, tables, etc., and extract relevant metadata such as font information, image resolution, color mode, etc. At the same time, analyze the structure and organization method of the data, such as the paragraph structure of the text, the position of the image on the page, etc. Then, according to the importance and usage frequency of the data, evaluate its value and cold-hot coefficient by analyzing the semantics of the data, its position in the document, its association with other data, etc. For important text content such as titles and key conclusions, assign a higher value and hot coefficient, and for auxiliary images or less important text, assign a relatively lower value and cold coefficient.

[0026] Preferably, the data is hierarchically divided according to value and the cold-hot coefficient. For example, it is divided into three levels: high, medium, and low. Data with high value and high hot coefficient are classified as the high level, such as core text content, important images, etc. Data with medium value and usage frequency are classified as the middle level, and data with low value and low usage frequency are classified as the low level, such as some decorative images or less critical annotations, etc.; then the data at the same level is encapsulated to form independent encapsulation blocks. For text data, it is encapsulated according to paragraphs or chapters. For image data, it is grouped and encapsulated according to its position or logical relationship in the page, and preprocessing is performed during the encapsulation process, such as character encoding conversion for text, format conversion or resolution adjustment for images, etc., which helps to improve the compression efficiency and quality.

[0027] Preferably, the adaptive loss based on multi-threading is compressed into the second compression processing node. During the compression process, multi-threading can be used to parallelly perform compression operations on different encapsulation blocks, thus significantly improving the compression speed. Adaptive loss compression is an intelligent compression strategy that can automatically adjust the compression degree according to the type and importance of the data. For data with high quality requirements, such as important text content, the loss is minimized during compression to ensure the integrity and accuracy of the data; for data with relatively low quality requirements, such as fine image textures, the compression loss is appropriately increased on the premise of not affecting the overall visual effect to obtain a larger compression ratio. Specifically, the archived and encapsulated data is divided into multiple data blocks according to data type, data quality, data volume size, or encapsulation block level, and then one or more threads are assigned to each data block according to the complexity and estimated processing time of the data block; then the data importance is evaluated, and corresponding loss thresholds are determined for different types and importance of data. For example, for high-resolution images, the resolution is allowed to be appropriately reduced on the premise of not affecting the visual effect, and for text data, the proportion of character replacement or deletion is restricted; finally, according to the data type and loss threshold, a suitable compression algorithm is selected. For example, for text, Huffman coding, run-length coding, etc. can be used, and for images, compression formats such as JPEG and PNG can be used, and the compression parameters are adjusted according to the loss threshold; and multiple threads simultaneously perform compression processing on the data blocks assigned to them, and each thread selects a suitable compression algorithm and parameters for compression according to the pre-established adaptive loss strategy.

[0028] Preferably, the first pre - processing node and the second compression processing node are integrated to construct a compression processing module, which is deployed on a cloud processor. The cloud processor is a remote computing resource based on cloud computing technology. That is, the entire compression processing function is migrated to the cloud. Users do not need to install complex compression software and computing hardware on local devices. They only need to connect to the cloud through the network to perform compression processing. The cloud processor can dynamically adjust computing resources according to actual usage requirements. When a large number of users request compression services simultaneously, the cloud can automatically allocate more computing resources to ensure the efficient operation of the service.

[0029] Preferably, then the compression processing module is called and activated at the physical data end. The physical data end refers to the device that actually stores or uses PDF file data, such as a personal computer, a mobile device, etc. Users send requests to the cloud processor through the network to call and activate the compression processing module deployed on the cloud processor. Specifically, when a user needs to compress a certain PDF file, the application program at the physical data end uploads the data of the file to the cloud processor and sends a call instruction. After receiving the instruction and data, the cloud processor activates the compression processing module and, according to the pre - configuration, first performs data archiving and encapsulation, then performs adaptive loss compression based on multi - threads, and finally returns the compressed file to the physical data end.

[0030] Further, as Figure 2 shown, step S110 further includes step S111, performing forward learning on the compression processing module to determine the first compression logic; step S112, performing reverse learning on the compression processing module to determine the second decompression logic; step S113, taking the file status as a standard, performing directional activation on the compression processing module based on the first compression logic and the second decompression logic.

[0031] Preferably, forward learning is performed on the compression processing module, that is, the original uncompressed PDF file data is processed, including evaluating the characteristics of the data of a large number of actual processing cases, the distribution of different types of data, the associations between data, etc., learning how to compress the data more efficiently, and then summarizing a set of best compression processes, that is, the first compression logic, which includes compression strategies for different types of data (such as text, images). For example, for text data, what encoding method is used for two-way semantic-character compression, and for image data, how geometric-texture two-way compression is performed, as well as the specific operation steps and parameter settings in each link such as data archiving encapsulation and multi-threaded adaptive loss compression. Reverse learning is performed on the compression processing module, that is, the compressed data is analyzed and processed to efficiently and accurately restore the original data, and the second decompression logic is determined, which stipulates the order and method for processing the compressed data during decompression. For example, if semantic analysis and character encoding are performed on text data during compression, during decompression, the opposite steps need to be followed, first decoding the encoded characters and then restoring them according to the semantic information; it also includes progressive decompression of image data.

[0032] Preferably, taking the file status as the standard, according to the first compression logic and the second decompression logic, the compression processing module is directionally activated. Here, the file status refers to the current stage of the PDF file, which is mainly divided into the to-be-compressed state and the decompression-browsing state. The to-be-compressed state indicates that the file is original and has not been compressed, and needs to be compressed for storage or transmission; the decompression-browsing state indicates that the file is already a compressed file, and the user needs to decompress it and browse it on the display interface. According to the different states of the file, the compression processing module specifically activates the corresponding logic. That is, when the file is in the to-be-compressed state, the first compression logic is activated, and the compression processing module compresses the file according to the best compression process determined by forward learning, converting the original PDF file into a compressed file; when the file is in the decompression-browsing state, the second decompression logic is activated, and the module decompresses the compressed file according to the decompression order determined by reverse learning, generating a pre-browsing file on the display interface to meet the user's browsing needs; thus, the compression processing module can more intelligently and efficiently complete the compression and decompression tasks of PDF files.

[0033] Furthermore, step S110 further includes step S114 of constructing a multi-threaded compression port according to the redundancy ratio - compression loss coefficient, where the redundancy ratio is determined by balancing the data value and the cold-hot coefficient; step S115 of performing compression learning on each thread compression port with the parallel compression under format segmentation - two-way decoupling as the processing logic for the multi-threaded compression port; and step S116 of integrally learning the multi-threaded compression port in parallel and using it as the second compression processing node.

[0034] Preferably, the redundancy ratio is determined by balancing the data value and the hot-cold coefficient. Among them, the data value reflects the importance of the data in the entire PDF file. For example, key text content, core images, etc. have a higher data value. The hot-cold coefficient reflects the usage frequency of the data. Frequently accessed data is hot data, and vice versa. By comprehensively considering the data value and the hot-cold coefficient, the redundancy degree existing in the data can be evaluated; the compression loss coefficient represents the degree of data volume or quality loss allowed during the compression process. Different types of data can be set with different compression loss coefficients. For data with high quality requirements, such as important text, this coefficient should be set smaller. For some data with relatively low quality requirements, such as some details of images, this coefficient can be appropriately increased to obtain a higher compression ratio. Then, according to the redundancy ratio and the set compression loss coefficient, a multi-threaded compression port is constructed, that is, multiple parallel compression channels are constructed. Each channel can independently perform data compression processing. By allocating data according to data types, the redundancy ratio of the data, etc., to different ports for processing, the computing power of the multi-core processor is fully utilized to improve the data compression efficiency.

[0035] Preferably, the PDF file contains various types of data formats, such as text, images, graphics, tables, etc. Different formats of data have different characteristics and compression requirements. Format segmentation means separating different data formats (such as text, images, tables, etc.) in the PDF file, and then performing compression processing separately, that is, performing parallel compression under two-way decoupling, including semantic-character two-way compression and geometric-texture two-way compression. Among them, semantic compression is processed from the meaning level of the text, such as removing redundant expressions, merging similar semantic information, etc. Character compression is to optimize the encoding of the characters themselves, such as using Huffman coding and other methods to reduce the storage space of the characters; Geometric compression mainly processes the geometric features of the image, such as the shape, size, etc. of the image, such as adjusting the resolution of the image, performing image scaling, etc. Texture compression focuses on processing the texture details of the image, and reduces the file size by removing some texture information that is not easily noticed by the human eye; thus, the data can be compressed more effectively to ensure a better compression effect.

[0036] Preferably, then, according to the processing logic, compression learning is performed on each thread compression port, that is, by processing and analyzing a large amount of data, continuously adjusting the parameters of the compression algorithm, and optimizing the compression process to find the most suitable compression method for processing data on this port; finally, the multiple thread compression ports of the compression learning are integrated in parallel to obtain the second compression processing node, and each thread compression port continues to work in parallel to jointly complete the compression task of the entire PDF file, realizing fast and high-quality compression of the PDF file, and improving the compression efficiency and quality.

[0037] Step S200: Scan the PDF file, perform file data interpretation, hierarchically archive and encapsulate the file data according to the data value and the cold-hot coefficient, determine multiple encapsulated blocks and perform decoupled compression to determine the compressed file, where the decoupled compression is a two-way compression of semantic-character for text data and geometric-texture for image data.

[0038] Preferably, scanning the PDF file means scanning and viewing the PDF file page by page and element by element to accurately locate all data elements in the file, such as text paragraphs, various images, complex graphics, and different tables, etc.; then perform file data interpretation, that is, deeply analyze the structure and encoding rules of the file to identify different types of data and extract key information. For example, for text data, determine its font, font size, color and other format information, and for image data, clarify its resolution, color mode and other attributes; then hierarchically archive and encapsulate the file data according to the data value and the cold-hot coefficient, including classifying high-value and high cold-hot coefficient data into high levels, such as the title of the file, important conclusions, etc., classifying medium-value and frequently used data into intermediate levels, and classifying low-value and low-frequency data into low levels; then encapsulate the data at the same level to form multiple independent encapsulated blocks.

[0039] Preferably, performing decoupled compression on the encapsulated blocks includes two-way compression of semantic-character for text data and geometric-texture for image data. Specifically, semantic compression uses natural language processing technology to analyze the grammatical structure and semantic relationship of the text, focusing on understanding and processing the meaning of the text to remove redundant information and retain the core semantics. For example, simplify some repetitive sentences, merge synonyms and near-synonyms, and remove modifiers that have little impact on the overall semantics; thereby reducing the number of characters in the text and achieving compression. Character compression is based on the optimization of character encoding to reduce data storage space. For example, use Huffman coding to assign different lengths of codes to each character according to the frequency of the character in the text. Characters with high frequencies use shorter codes, and characters with low frequencies use longer codes. After encoding the original longer text, the overall data volume is reduced. Semantic compression and character compression cooperate with each other, first streamline the text content at the semantic level, and then further optimize at the character encoding level to achieve two-way compression and improve the compression efficiency.

[0040] Preferably, geometric compression mainly processes the geometric features of an image, such as size and shape, including adjusting the resolution of the image, scaling the image, etc. For example, if the original resolution of the image is too high, the resolution can be appropriately reduced to reduce the number of pixels in the image, thereby reducing the size of the image file. Texture compression refers to analyzing the texture details of an image and removing fine texture information by quantifying and filtering the texture to reduce the storage space of texture data. For example, the wavelet transform is used to decompose the image into sub-bands of different frequencies, and then appropriate quantization processing is performed on the high-frequency sub-bands to remove unimportant information. Geometric compression and texture compression are combined to process the image to achieve effective compression of the image data. After the decoupled compression of each encapsulation block is completed, the compressed encapsulation blocks are combined to form the final compressed file.

[0041] Further, step S200 further includes step S210 of determining the content architecture system of the PDF file through file data interpretation according to the first preprocessing node; step S220 of performing a primary hierarchical filing based on data value for the content architecture system to determine the first filing result; step S230 of performing a secondary hierarchical filing based on the cold and hot coefficients to determine the second filing result; and step S240 of performing filing reorganization under weighted calculation on the first filing result and the second filing result to determine multiple filing layers and perform data encapsulation to determine multiple encapsulation blocks, where each encapsulation block corresponds to a redundancy ratio.

[0042] Preferably, the PDF file is processed by the first preprocessing node, that is, file data interpretation is performed, including identifying various elements in the file, such as text (different fonts, sizes, colors, paragraph formats, etc.), images (resolution, color mode, image type, etc.), tables, graphics, etc., and at the same time extracting relevant metadata information, so as to clearly understand the organization method and logical structure of the data in the file, and then constructing the content architecture system of the PDF file with this data. For example, determining the chapter hierarchy relationship of the text, the position of the image on the page and its correspondence with the text, etc.

[0043] Preferably, based on the content architecture system, the PDF file data is evaluated according to the importance of the data in the file, the degree of contribution to understanding the core content of the file, etc., to judge the value of each piece of data. For example, the title of the file, key conclusions, important data charts, etc. usually have relatively high data value, while the data value of auxiliary explanations and decorative elements is relatively low; then, according to the high and low data values, the data is divided into different levels, with high-value data in higher levels and low-value data in lower levels. Finally, the filing result based on the data value, that is, the first filing result, is obtained. On the basis of the first filing result, the hot and cold coefficients of the data are further evaluated. For the data in each level, it is further subdivided according to its hot and cold coefficients. For example, in the high-value data level, the hot data and cold data are divided into different sub-levels respectively, and then a more detailed filing structure, that is, the second filing result, is obtained.

[0044] Preferably, according to the degree of emphasis on data value and usage frequency, corresponding weights are assigned to different levels in the first filing result (hierarchical division based on data value) and the second filing result (hierarchical filing based on hot and cold coefficients). The two filing results are integrated and adjusted through weighted calculation to re-determine the hierarchical relationship of the data, forming multiple filing levels; then the data in the same filing level is encapsulated, that is, the relevant data is combined together to form an independent encapsulation block. Each encapsulation block contains data with similar characteristics (such as similar value and usage frequency), and each encapsulation block corresponds to a redundancy ratio. If there is a lot of duplicate or removable redundant information in the data within the encapsulation block, its redundancy ratio is high; otherwise, the redundancy ratio is low.

[0045] Further, step S240 further includes step S241 of determining a first encapsulation block, where the first encapsulation block is any one of the multiple encapsulation blocks; step S242 of determining a first redundancy coefficient for the first encapsulation block based on the filing level of the first filing result; step S243 of determining a second redundancy coefficient based on the filing level of the second filing result; and step S244 of determining a first redundancy ratio according to the first redundancy coefficient and the second redundancy coefficient.

[0046] Preferably, any one of the multiple encapsulation blocks is selected as the first encapsulation block, and the first redundancy coefficient is determined according to the archival layer of the first archival result to which it belongs. Specifically, since data values are different, their redundancy levels may also be different. High-value data (such as core conclusions and key data) is the core of the file and has little deletable content, usually with low redundancy. Low-value data (such as some auxiliary explanations and repeated examples) may have more redundant information. Therefore, based on the data value level reflected by the archival layer where the first encapsulation block is located, the redundancy level of the data within the encapsulation block is evaluated to determine the first redundancy coefficient. The larger the value, the higher the redundancy level of the encapsulation block based on the data value level.

[0047] Preferably, also for the first encapsulation block, the second redundancy coefficient is determined based on its position in the archival layer of the second archival result. Among them, there is also a certain relationship between the usage frequency of data and the redundancy level. Hot data (data that is frequently accessed) is carefully organized and has little redundancy; while cold data (data that is rarely accessed) has some unnecessary information or repeated content. Therefore, based on the archival layer of the first encapsulation block based on the hot-cold coefficient, the redundancy level of the data within the encapsulation block is evaluated to determine the second redundancy coefficient. The larger the value, the higher the redundancy level of the encapsulation block based on the data usage frequency level. Finally, according to the degree of emphasis on data value and usage frequency, different weights are assigned to the first redundancy coefficient and the second redundancy coefficient, and the first redundancy ratio is calculated using weighted average to comprehensively measure the redundancy level of the data within the first encapsulation block. The higher the redundancy ratio, the greater the compression space of the encapsulation block, and a more aggressive compression strategy can be adopted to reduce the file size.

[0048] Furthermore, step S200 further includes step S250, importing the multiple encapsulation blocks into the second compression processing node, activating the multi-threaded compression port by identifying the redundancy ratio, and importing the target compression port, where the target compression port includes at least one; step S260, the target compression port identifies the imported encapsulation blocks, performs data structure identification within the encapsulation blocks, and determines the first text data and the second image data; step S270, performs semantic-character decoupling on the first text data to determine the first compression round trip; step S280, performs geometric-texture decoupling on the second image data to determine the second compression round trip; step S290, performs concurrent compression processing according to the first compression round trip and the second compression round trip to determine the compressed encapsulation blocks; step S2100, integrates the compressed encapsulation blocks output by each target compression port as the compressed file.

[0049] Preferably, multiple encapsulated blocks are imported into the second compression processing node, and then the redundancy ratio corresponding to each encapsulated block is identified. According to the identified redundancy ratio, multi-threaded compression ports are activated. The multi-threaded compression ports are multiple parallel processing channels, and each port can independently compress data. After the multi-threaded compression ports are activated, the encapsulated blocks are imported into appropriate target compression ports according to the size of the redundancy ratio, data type, etc. And there is at least one target compression port, which can realize parallel processing of multiple encapsulated blocks and improve the compression efficiency. After each target compression port receives an encapsulated block, it identifies its data structure, distinguishes whether the data in the encapsulated block is text data or image data, and marks the text data in the encapsulated block as first text data and the image data as second image data, so as to adopt different compression strategies for different types of data.

[0050] Preferably, for the first text data, a semantic-character decoupling operation is performed. Specifically, semantic decoupling is to analyze the meaning of the text, remove redundant expressions, and extract core semantic information; character decoupling is to optimize the character encoding of the text, reduce the storage space occupied by characters, and then determine the first compression two-way process, that is, form a complete compression strategy for text data, including performing semantic processing first and then character processing. For the second image data, a geometry-texture decoupling operation is performed. Among them, geometry decoupling mainly processes geometric features such as the shape and size of the image, such as adjusting the resolution of the image, performing image scaling, etc.; texture decoupling focuses on the texture details of the image, removes texture information that is not easily perceptible to the human eye to reduce the file size, and then determines the second compression two-way process, that is, a complete compression strategy for image data, including geometry processing and texture processing. Based on the first compression two-way process (for text data) and the second compression two-way process (for image data), the target compression port simultaneously performs compression processing on the text and image data in the encapsulated block, that is, concurrent compression, compresses the data in the encapsulated block to obtain a compressed encapsulated block, and the data volume has been reduced compared with the original encapsulated block; finally, the compressed encapsulated blocks output by each target compression port are integrated to form a complete compressed file, and finally the efficient compression of the entire PDF file is realized.

[0051] Step S300, as the sink end receives the compressed file, the compression processing module is called and reversely activated, and reverse decompression based on the compression logic is performed to generate a pre-view file on the display interface, where the pre-decompression of high-value data and the progressive decompression of image data are used as decompression constraint conditions.

[0052] Preferably, the destination end refers to the end that receives data, which can be devices such as a user's computer, mobile phone, tablet, etc. The destination-end device receives the PDF compressed file sent from the source end through a network (such as the Internet, local area network, etc.), and then calls the compression processing module and performs reverse activation. Among them, the compression processing module is used to compress files at the source end and decompress files at the destination end. At the source end, the compression processing module performs forward activation, that is, first performs hierarchical archiving and encapsulation of data, and then performs steps such as multi-threaded adaptive loss compression to compress the file; at the destination end, the compression processing module is made to enter a working mode opposite to the compression process, and it processes the received data according to the reverse logic. Specifically, after the compression processing module is reversely activated, reverse operations are performed according to the compression logic determined by the source end. For example, if the source end performs semantic-character two-way compression on text data, during decompression, the character encoding must first be restored to the original state, and then the original text content is restored according to the semantic information; for image data, if geometric-texture two-way compression is performed during compression, during decompression, the texture information must first be restored, and then features such as the geometric shape and size of the image are restored. Finally, the decompressed data is converted into a format suitable for presentation on a display interface (such as the display screen of a computer, the screen of a mobile phone), thereby generating a pre-view file, that is, a quickly loaded preview version of the original PDF file, which enables users to see the general content of the file in a short time.

[0053] Preferably, with the pre-decompression of high-value data and the progressive decompression of image data as the decompression constraint conditions. Specifically, when decompressing the compressed file data, high-value data is preferentially pre-decompressed to enable users to quickly obtain key information. For example, in a PDF report, the title, core conclusions, important data charts, etc. of the file belong to high-value data, and they are preferentially decompressed so that they can be quickly displayed on the screen, allowing users to understand the core points of the file; since image data generally occupies a large storage space, if it is completely decompressed at one time, it may cause the file loading time to be too long, affecting the user experience. Therefore, progressive decompression is adopted, that is, the low-resolution version or basic contour information of the image is first decompressed, enabling users to quickly see the general appearance of the image. Then, when users perform operations such as magnifying the image, the higher-resolution and more detailed texture information of the image is gradually decompressed. Users can not only quickly browse the file but also obtain a satisfactory visual effect when they need to view a clear image, thereby effectively improving the decompression efficiency and the user's browsing experience.

[0054] Further, step S300 further includes step S310 of invoking the compression processing module to the physical data end; step S320 of temporarily determining the end type of the physical data end by identifying the file status of the PDF file; and step S330 of determining the activation direction of the compression processing module according to the end type, where the activation direction includes forward activation and reverse activation, and forward activation performs compression processing while reverse activation performs decompression processing.

[0055] Step S330 further includes step S331 of, if the file status is the compressed state, identifying the end type as the sink end for the processing method of decompression; and step S332 of, if the file status is the uncompressed state, identifying the end type as the source end for the processing method of compression.

[0056] Preferably, invoking the compression processing module to the physical data end means loading it from its original storage or deployment location (possibly a cloud server) into the device memory of the physical data end, enabling it to run on the device and process the PDF file. By identifying the file status of the PDF file, including the uncompressed original state, that is, the file exists in its originally created or stored form, which may be large in size and has not been compressed yet; and the compressed state, that is, the file has been compressed and its size has become smaller for easier storage and transmission. Then, according to the identified PDF file status, the end type of the physical data end is determined. Specifically, if the PDF file owned by the physical data end is in the uncompressed original state, the physical data end is temporarily determined as the source end (the end where information is generated or sent out) to compress the file for subsequent storage or transmission. On the contrary, if the physical data end receives a compressed PDF file, it is temporarily determined as the sink end (the end where information is received) to decompress the compressed file to view the file content.

[0057] Preferably, the activation direction of the compression processing module is determined according to the end type. Specifically, if it is detected that the PDF file stored or received by the physical data end has not been compressed, that is, the PDF file on the physical data end is in the original state, the file size remains the same as when it was created, and the data has not undergone any form of reduction or encoding optimization, the physical data end is determined as the source end, indicating that the uncompressed PDF file needs to be processed to make it easier to store and transmit. Then, the compression processing module is forward-activated to process the PDF file according to the preset compression logic, such as analyzing the characteristics and value of the data through data interpretation, then performing hierarchical archiving and encapsulation, and then executing decoupled compression (such as semantic-character two-way compression for text data and geometric-texture two-way compression for image data, etc.). Finally, the original PDF file is converted into a compressed file to reduce the file size and improve the storage and transmission efficiency.

[0058] Preferably, if it is detected that the PDF file stored or received by the physical data terminal has been compressed, that is, the file exists in a compressed format, such as the file size is significantly reduced compared to the original state, or the file format is a format processed by a specific compression algorithm, the physical data terminal is determined as the destination terminal, indicating that the received is a compressed PDF file, and the purpose is to restore it to a browsable original state. Then, the compression processing module is reversely activated. The compression processing module decompresses the compressed file according to the logic opposite to the compression process. By performing reverse operations based on the compression logic, such as restoring the semantic and character information of the compressed text data and restoring the geometric and texture features of the compressed image data, etc., finally, a browsable file is generated on the display interface.

[0059] Further, step S300 further includes step S340, performing pre-decompression deployment according to the data value gradient to determine the first decompression condition; step S350, using progressive decompression as the decompression method for image data to determine the second decompression condition; step S360, reversely activating the compression processing module, introducing the first decompression condition and the second decompression condition, performing decompression processing on the compressed file, and performing progressive restoration of the interface of the PDF file in the execution timing of the display interface to determine the pre-view file.

[0060] Preferably, according to factors such as the importance of the data and the key degree for understanding the file content, the data value is evaluated, and the data is divided into different levels to form a data value gradient. Based on the data value gradient, pre-decompression deployment is performed, including preferentially arranging pre-decompression for high-value data, such as determining the storage location of high-value data in the compressed file, allocating corresponding computing resources, etc., and then determining the first decompression condition, which mainly stipulates under what circumstances and in what way to decompress high-value data. For example, when the user requests to browse the file, first quickly decompress the encapsulation block where the high-value data is located, or stipulate the computing resources and time limit required for decompressing high-value data.

[0061] Preferably, progressive decompression means first quickly decompressing a low-resolution version or basic contour information of the image, so that the user can see the general appearance of the image in a short time. When the user performs operations such as zooming in on the image, then gradually decompress higher-resolution and more detailed texture information, and then determine the second decompression condition, including setting the specific steps of progressive decompression, such as the resolution ratio increased by each decompression, the time interval of decompression in different stages, and under what circumstances to trigger decompression of higher resolution. For example, it is stipulated to initially decompress a 1 / 4 resolution version of the image, and when the user's mouse hovers over the image, further decompress it to 1 / 2 resolution, and decompress it to the original resolution after clicking on the image.

[0062] Preferably, when decompressing a compressed file, the compression processing module for the compressed file is reversely activated, that is, the compression processing module is made to enter the decompression working mode, and the compressed file is processed according to the logic opposite to the compression process. During the decompression process, the first decompression condition (the pre-decompression condition for high-value data) and the second decompression condition (the progressive decompression condition for image data) are applied to the decompression process, including restoring the original state of the data according to the type of data (text, image, etc.) and the compression logic. At the same time, on the display interface, the interface of the PDF file is gradually restored in chronological order (time sequence). First, the decompressed high-value data, such as the title and core conclusions of the file, are displayed, and then as the image data is progressively decompressed, clearer images and other data are gradually displayed. In this way, the user can see the general content of the file in a short time and obtain a more complete and clear browsing experience with other operations; finally, a pre-viewed file is presented on the display interface.

[0063] The above specific embodiments do not limit the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present application shall be included within the protection scope of the present application. In some cases, the actions or steps recorded in the present application can be executed in a different order from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multi-tasking and parallel processing are also possible or may be advantageous.

Claims

1. A PDF file browsing method based on image compression and fast decompression technology, characterized in that: The method comprises: At the source end, the compression processing module is called and forward activated; Scanning a PDF file, interpreting the file data, hierarchically archiving and packaging the file data according to the data value and the hot / cold coefficient, determining a plurality of packaging blocks and performing decoupled compression to determine a compressed file, wherein the decoupled compression is a semantic-character two-way compression of text data and a geometric-texture two-way compression of image data; As the destination receives the compressed file, the compression processing module is called and reversely activated to perform reverse decompression based on compression logic, and a preview file is generated on the display interface, wherein the pre-decompression of high-value data and the progressive decompression of image data are used as decompression constraints.

2. The PDF file browsing method based on image compression and fast decompression technology as claimed in claim 1, characterized in that: Before positively activating the compression processing module, construct the compression processing module, including: The compression processing module is constructed by using data archiving encapsulation as the first pre-processing node and adaptive lossy compression based on multi-threading as the second compression processing node; Deploy the compression processing module on a cloud processor; At the physical data end, the compression processing module is called and activated from the cloud processor.

3. The PDF file browsing method based on image compression and fast decompression technology as claimed in claim 2, characterized in that: Constructing the compression processing module includes: Performing forward learning on the compression processing module to determine a first compression logic; Performing reverse learning on the compression processing module to determine a second decompression logic; Based on the file status, the compression processing module is activated in a directional manner based on the first compression logic and the second decompression logic.

4. The PDF file browsing method based on image compression and fast decompression technology as claimed in claim 3, characterized in that: The second compression processing node is an adaptive lossy compression based on multi-threading, including: Constructing a multi-threaded compression port according to a redundancy ratio-compression lossy coefficient, wherein the redundancy ratio is determined by balancing data value and a cold and hot coefficient; For the multi-threaded compression port, the parallel compression under format segmentation-two-way decoupling is used as the processing logic to perform compression learning on the compression port of each thread; The multi-threaded compression port after parallel ensemble learning serves as the second compression processing node.

5. The PDF file browsing method based on image compression and fast decompression technology as claimed in claim 1, characterized in that: Call the compression processing module, including: Calling the compression processing module to the physical data end; Temporarily determining the terminal type of the physical data terminal by identifying the file status of the PDF file; According to the terminal type, the activation direction of the compression processing module is determined, including forward activation and reverse activation. Forward activation performs compression processing, and reverse activation performs decompression processing.

6. The PDF file browsing method based on image compression and fast decompression technology as claimed in claim 5, characterized in that: If the file state is a compressed state, the end type is identified as a sink end, and decompression is used as a processing method; If the file status is a non-compressed status, the end type is identified as a source end, and compression is used as a processing method.

7. The PDF file browsing method based on image compression and fast decompression technology as claimed in claim 2, characterized in that: According to the data value and hot / cold coefficient, the file data is archived and packaged in a hierarchical manner, including: According to the first pre-processing node, the content architecture system of the PDF file is determined by interpreting the file data; For the content architecture system, perform a hierarchical archiving based on data value to determine a first archiving result; Perform secondary stratification archiving based on the cold and hot coefficients to determine the second archiving result; The first archiving result and the second archiving result are archived and reorganized under weighted calculation, a plurality of archiving layers are determined, data is encapsulated, and a plurality of encapsulation blocks are determined, wherein each encapsulation block corresponds to a redundancy ratio.

8. The PDF file browsing method based on image compression and fast decompression technology as claimed in claim 7, characterized in that: Each encapsulation block corresponds to a redundancy ratio, including: Determine a first encapsulation block, wherein the first encapsulation block is any one of the multiple encapsulation blocks; For the first encapsulation block, determine a first redundancy coefficient based on an archiving layer of the first archiving result; determining a second redundancy coefficient at an archiving layer based on the second archiving result; A first redundancy ratio is determined according to the first redundancy coefficient and the second redundancy coefficient.

9. The PDF file browsing method based on image compression and fast decompression technology as claimed in claim 4, characterized in that: Perform decoupling compression to determine the compressed files, including: Importing the plurality of encapsulated blocks into the second compression processing node, activating the multi-threaded compression port by identifying the redundancy ratio, and importing the target compression port, wherein the target compression port includes at least one; The target compression port identifies the imported encapsulation block, performs data structure identification in the encapsulation block, and determines the first text data and the second image data; Performing semantic-character decoupling on the first text data to determine a first compression round trip; performing geometry-texture decoupling on the second image data to determine a second compression round-trip; Performing concurrent compression processing according to the first compression round-trip and the second compression round-trip to determine a compression encapsulation block; The compressed encapsulation blocks output by each target compression port are integrated as the compressed file.

10. The PDF file browsing method based on image compression and fast decompression technology as claimed in claim 1, characterized in that: Perform reverse decompression based on compression logic and generate a pre-browsing file on the display interface, including: Perform pre-decompression deployment according to the data value gradient and determine the first decompression condition; Determine a second decompression condition by using progressive decompression as the image data decompression mode; The compression processing module is activated in reverse, the first decompression condition and the second decompression condition are introduced, the decompression processing is performed on the compressed file, the interface of the PDF file is gradually restored under the timing of the display interface execution, and the pre-browsed file is determined.

Citation Information

Patent Citations

  • Interaction-oriented editable art image generation system and method

    CN117036528A

  • Article analysis method and device

    CN118780277A

  • Bookmark directory generation method and device and electronic equipment

    CN118917279A

  • Compression method, system and device of multi-modal weight file and storage medium

    CN119493778A

  • Fontless structured document image representations for efficient rendering

    US5884014A