PDF file browsing method based on image compression and fast decompression technology

By interpreting and hierarchically archiving data at the source, utilizing multi-threaded adaptive lossy compression and cloud processors to perform two-way compression of text and images in PDF files, and performing reverse decompression at the destination, the problems of slow transmission and insufficient browsing fluency caused by the large amount of image data in PDF files are solved, achieving efficient data transmission and a fast browsing experience.

CN120123296BActive Publication Date: 2025-09-05BEIJING GUANGLIANDA YUNTU DREAM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510608656.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-09-05
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

In the existing technology, the large amount of image data in PDF files leads to slow transmission, and compression and decompression are difficult to balance efficiency and quality, resulting in insufficient data transmission efficiency and browsing fluency.

Method used

A method based on image compression and fast decompression technology is adopted to perform data interpretation and hierarchical archiving at the source end. Multi-threaded adaptive lossy compression and cloud processors are used to perform semantic-character two-way compression and geometric-texture two-way compression on text and image data respectively, and reverse decompression is performed at the destination end, with priority given to decompressing high-value data and progressively decompressing image data.

Benefits of technology

It improves data transmission efficiency and user browsing experience, achieves efficient compression and decompression, ensures rapid acquisition of key information, and improves data transmission and browsing fluency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123296B_ABST
    Figure CN120123296B_ABST
Patent Text Reader

Abstract

The present invention discloses a PDF file browsing method based on image compression and fast decompression technology, which relates to the technical field related to data processing. The method comprises: at the source end, calling a compression processing module and performing forward activation; scanning a PDF file, interpreting the file data, hierarchically archiving and packaging according to the data value and hot and cold coefficients, determining multiple packaging blocks and performing decoupling compression to determine a compressed file; calling a compression processing module and performing reverse activation, performing reverse decompression based on compression logic, generating a pre-browsing file on a display interface, and using pre-decompression of high-value data and progressive decompression of image data as decompression constraints. The method solves the technical problems existing in the prior art that the large amount of image data in PDF files leads to slow transmission, compression and decompression are difficult to balance efficiency and quality, resulting in insufficient data transmission efficiency and browsing fluency, thereby achieving the technical effect of improving data transmission efficiency and user browsing experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field related to data processing, and in particular to a PDF file browsing method based on image compression and fast decompression technology. Background Art

[0002] PDF files, due to their cross-platform nature, format stability, and content integrity, have become a widely used electronic document format in fields such as office, education, and publishing. However, PDF files typically contain large amounts of text and image data, presenting numerous challenges during data transmission and browsing. Traditional file compression technologies, such as ZIP and RAR, have limitations when handling PDF files. While they can reduce file size to a certain extent, they fail to fully exploit redundant information at the semantic and character levels for text data within PDF files, preventing deep, optimized compression. Furthermore, they cannot effectively distinguish between geometric structures and texture information for image data, making efficient compression difficult. Consequently, compression ratios and efficiencies fall short of meeting transmission and storage requirements. Furthermore, existing decompression technologies fail to differentiate data processing based on data value and user browsing needs when browsing PDF files. They require the entire file to be fully decompressed before content can be displayed. This not only consumes significant time and computing resources, but also hinders users from quickly accessing key information during the initial browsing process. This severely limits the browsing experience, especially on mobile devices or in poor network environments. This also impacts the data transmission efficiency and browsing responsiveness of PDF files.

[0003] At present, relevant technologies have technical problems such as the large amount of image data in PDF files leading to slow transmission, and the difficulty in balancing efficiency and quality during compression and decompression, resulting in insufficient data transmission efficiency and browsing fluency. Summary of the Invention

[0004] This application provides a PDF file browsing method based on image compression and fast decompression technology, which solves the technical problems in the prior art that the large amount of PDF file image data leads to slow transmission, and compression and decompression are difficult to balance efficiency and quality, resulting in insufficient data transmission efficiency and browsing fluency, thereby achieving the technical effect of improving data transmission efficiency and user browsing experience.

[0005] The present application provides a PDF file browsing method based on image compression and fast decompression technology, the method comprising: at a signal source end, calling a compression processing module and performing forward activation; scanning a PDF file, interpreting the file data, hierarchically archiving and encapsulating the file data according to data value and hot / cold coefficients, determining multiple encapsulation blocks and performing decoupled compression to determine a compressed file, wherein the decoupled compression is semantic-character two-way compression of text data and geometric-texture two-way compression of image data; as a signal sink end receives the compressed file, calling the compression processing module and performing reverse activation, performing reverse decompression based on compression logic, and generating a pre-browsed file on a display interface, wherein pre-decompression of high-value data and progressive decompression of image data are used as decompression constraints.

[0006] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology also performs the following processing: using data archiving encapsulation as the first pre-processing node and adaptive lossy compression based on multi-threading as the second compression processing node to construct the compression processing module; deploying the compression processing module on a cloud processor; and calling and activating the compression processing module from the cloud processor at the physical data end.

[0007] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology further performs the following processing: forward learning is performed on the compression processing module to determine the first compression logic; reverse learning is performed on the compression processing module to determine the second decompression logic; and based on the file status as a standard, directed activation of the compression processing module based on the first compression logic and the second decompression logic.

[0008] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology also performs the following processing: constructing a multi-threaded compression port based on the redundancy ratio-compression lossy coefficient, wherein the redundancy ratio is determined by balancing the data value and the hot and cold coefficient; for the multi-threaded compression port, compression learning is performed on each thread compression port using parallel compression under format segmentation-two-way decoupling as the processing logic; the multi-threaded compression port after parallel integrated learning serves as the second compression processing node.

[0009] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology also performs the following processing: calling the compression processing module to the physical data end; temporarily determining the end type of the physical data end by identifying the file status of the PDF file; and determining the activation direction of the compression processing module according to the end type, which includes forward activation and reverse activation, forward activation performs compression processing, and reverse activation performs decompression processing.

[0010] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology further performs the following processing: if the file status is a compressed state, the end type is identified as a destination end, and decompression is used as the processing method; if the file status is a non-compressed state, the end type is identified as a source end, and compression is used as the processing method.

[0011] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology also performs the following processing: according to the first pre-processing node, the content architecture system of the PDF file is determined by interpreting the file data; for the content architecture system, a hierarchical archiving is performed based on the data value to determine the first archiving result; a secondary hierarchical archiving is performed based on the hot and cold coefficients to determine the second archiving result; the first archiving result and the second archiving result are archived and reorganized under weighted calculation to determine multiple archiving layers and perform data encapsulation to determine multiple encapsulation blocks, wherein each encapsulation block corresponds to a redundancy ratio.

[0012] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology further performs the following processing: determining a first encapsulation block, wherein the first encapsulation block is any one of the multiple encapsulation blocks; determining a first redundancy coefficient for the first encapsulation block based on an archiving layer based on the first archiving result; determining a second redundancy coefficient based on an archiving layer based on the second archiving result; and determining a first redundancy ratio based on the first redundancy coefficient and the second redundancy coefficient.

[0013] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology further performs the following processing: importing the multiple encapsulated blocks into the second compression processing node, activating the multi-threaded compression port by identifying the redundancy ratio, and importing the target compression port, wherein the target compression port includes at least one; the target compression port identifies the imported encapsulated blocks, performs data structure identification within the encapsulated blocks, and determines the first text data and the second image data; performs semantic-character decoupling on the first text data to determine the first compression two-way pass; performs geometric-texture decoupling on the second image data to determine the second compression two-way pass; performs concurrent compression processing based on the first compression two-way pass and the second compression two-way pass to determine the compressed encapsulated blocks; and integrates the compressed encapsulated blocks output by each target compression port as the compressed file.

[0014] In a possible implementation, the PDF file browsing method based on image compression and fast decompression technology also performs the following processing: pre-decompression deployment is performed according to the data value gradient to determine the first decompression condition; progressive decompression is used as the image data decompression method to determine the second decompression condition; reverse activation of the compression processing module, introduction of the first decompression condition and the second decompression condition, decompression processing is performed on the compressed file, and the interface of the PDF file is gradually restored under the execution sequence of the display interface to determine the pre-browsed file.

[0015] The proposed PDF file browsing method based on image compression and fast decompression technology in this application is to call the compression processing module and perform forward activation at the source end; scan the PDF file, interpret the file data, perform hierarchical archiving and packaging based on the data value and hot and cold coefficients, determine multiple packaging blocks and perform decoupling compression to determine the compressed file; call the compression processing module and perform reverse activation, perform reverse decompression based on compression logic, generate a pre-browsed file on the display interface, and use the pre-decompression of high-value data and the progressive decompression of image data as decompression constraints. This solves the technical problems existing in the prior art of slow transmission due to the large amount of image data in PDF files, the difficulty in balancing efficiency and quality in compression and decompression, and the resulting lack of data transmission efficiency and browsing fluency, thereby achieving the technical effect of improving data transmission efficiency and user browsing experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention are briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in precise order. Instead, various steps may be processed in reverse order or simultaneously as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0017] Figure 1 A flowchart of a PDF file browsing method based on image compression and fast decompression technology provided in an embodiment of the present application.

[0018] Figure 2 This is a flow chart of constructing a compression processing module in a PDF file browsing method based on image compression and fast decompression technology provided in an embodiment of the present application. DETAILED DESCRIPTION

[0019] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.

[0020] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0021] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict, and the terms “first\second” involved are merely to distinguish similar objects and do not represent a specific ordering of the objects. The terms “including” and “having” and any variations are intended to cover non-exclusive inclusions. For example, a process, method, product, or server comprising a series of steps is not necessarily limited to those steps clearly listed, but may include other steps that are not clearly listed or inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application only.

[0022] The embodiment of the present application provides a PDF file browsing method based on image compression and fast decompression technology, such as Figure 1 As shown, the method includes:

[0023] Step S100: At the information source end, the compression processing module is called and forward activated.

[0024] Preferably, the source end, i.e., the source of information, refers to the end that generates or owns the original PDF file and needs to process it for transmission or storage. For example, a user creates a PDF file containing a large amount of images and text on his or her computer. This computer is the source end, or a server stores a PDF file to be sent to other devices. This server is considered the source end. Calling the compression processing module and performing positive activation means that at the source end, the compression processing module is started through specific instructions and interfaces to start working. For example, in PDF processing software, when the user clicks "Compress File", the corresponding code is executed to call the compression processing module and process the data transmission of the currently opened PDF file. The compression processing module is a functional module for compressing PDF file data and can perform compression processing on different data types (such as text data, image data, etc.) in the PDF file. For example, text data is processed using semantic analysis and character compression, and image data is processed through geometric feature extraction and texture compression. Positive activation indicates that it is consistent with the normal process direction of the compression processing, that is, starting from the original uncompressed PDF file data, the compression processing module executes the processing logic, and finally obtains the compressed file.

[0025] Furthermore, step S100 also includes step S110, using data archiving encapsulation as the first pre-processing node and adaptive loss compression based on multi-threading as the second compression processing node to construct the compression processing module; step S120, deploying the compression processing module on the cloud processor; step S130, at the physical data end, calling and activating the compression processing module from the cloud processor.

[0026] Preferably, for complex PDF file data, which contains multiple types such as text, images, and charts, and has different data values ​​and hot / cold coefficients (i.e., data heat or usage frequency), data archiving and packaging is used as the first pre-processing node. Data archiving and packaging refers to interpreting the PDF file data, analyzing its data characteristics and attributes, and dividing the file data into layers according to data value and hot / cold coefficients. The data of different layers are then packaged separately to form multiple packaged blocks. Specifically, the data in the PDF file is first comprehensively parsed to identify different types of data, such as text, images, graphics, tables, etc., and relevant metadata, such as font information, image resolution, color mode, etc., are extracted. At the same time, the structure and organization of the data, such as the paragraph structure of the text and the position of the image on the page, are analyzed. Then, based on the importance and frequency of use of the data, its value and hot / cold coefficient are evaluated by analyzing the data's semantics, position in the document, and association with other data. Important text content, such as titles and key conclusions, is assigned a higher value and hot coefficient, while auxiliary images or less important text are assigned a relatively lower value and hot coefficient.

[0027] Preferably, the data is divided into levels according to its value and hot / cold coefficient, for example, into three levels: high, medium and low. The data with high value and high hot coefficient are classified into the high level, such as core text content, important images, etc. The data with medium value and frequency of use are classified into the middle level, and the data with low value and low frequency of use are classified into the low level, such as some decorative images or less critical annotations. Then the data at the same level are encapsulated to form independent encapsulation blocks. For text data, they are encapsulated according to paragraphs or chapters. For image data, they are grouped and encapsulated according to their position or logical relationship in the page. Preprocessing is performed during the encapsulation process, such as character encoding conversion for text, format conversion for images or resolution adjustment, etc., which helps to improve compression efficiency and quality.

[0028] Preferably, adaptive loss compression based on multi-threading is used as the second compression processing node. During the compression process, multi-threading can be used to perform compression operations on different encapsulated blocks in parallel, thereby significantly improving the compression speed. Adaptive loss compression is an intelligent compression strategy that can automatically adjust the degree of compression according to the type and importance of the data. For data with higher quality requirements, such as important text content, the loss is minimized during compression to ensure the integrity and accuracy of the data; for data with relatively low quality requirements, such as fine textures of images, the compression loss is appropriately increased in exchange for a larger compression ratio without affecting the overall visual effect. Specifically, the archived and encapsulated data is divided into multiple data blocks according to the data type, data quality, data size or encapsulation block level, and then one or more threads are assigned to each data block for processing according to the complexity of the data block and the expected processing time; then the data importance is evaluated, and the corresponding loss threshold is determined for data of different types and importance. For example, for high-resolution images, the resolution is allowed to be appropriately reduced without affecting the visual effect. For text data, the proportion of character replacement or deletion is limited; finally, according to the data type and loss threshold, a suitable compression algorithm is selected. For example, for text, Huffman coding, run-length coding, etc. can be used. For images, JPEG, PNG and other compression formats can be used, and the compression parameters are adjusted according to the loss threshold; and multiple threads compress their assigned data blocks at the same time, and each thread selects a suitable compression algorithm and parameters for compression according to a pre-established adaptive loss strategy.

[0029] Preferably, the first pre-processing node and the second compression processing node are integrated to construct a compression processing module, and the module is deployed on a cloud processor, wherein the cloud processor is a remote computing resource based on cloud computing technology, that is, the entire compression processing function is migrated to the cloud. Users do not need to install complex compression software and computing hardware on local devices, but only need to connect to the cloud through the network to perform compression processing. The cloud processor can dynamically adjust computing resources according to actual usage needs. When a large number of users request compression services at the same time, the cloud can automatically allocate more computing resources to ensure efficient operation of the service.

[0030] Preferably, the compression processing module is then called and activated at the physical data end, wherein the physical data end refers to the device that actually stores or uses the PDF file data, such as a personal computer, mobile device, etc. The user sends a request to the cloud processor through the network to call and activate the compression processing module deployed on the cloud processor. Specifically, when the user needs to compress a PDF file, the application at the physical data end uploads the data of the file to the cloud processor and sends a call instruction. After receiving the instruction and data, the cloud processor activates the compression processing module and performs data archiving and packaging according to the pre-configuration, and then performs multi-threaded adaptive loss compression, and finally returns the compressed file to the physical data end.

[0031] Further, such as Figure 2 As shown, step S110 also includes step S111, performing forward learning on the compression processing module to determine the first compression logic; step S112, performing reverse learning on the compression processing module to determine the second decompression logic; step S113, performing directional activation of the compression processing module based on the first compression logic and the second decompression logic based on the file status.

[0032] Preferably, the compression processing module is forward-learned, i.e., it processes the original uncompressed PDF file data, including evaluating the characteristics of data from a large number of actual processing cases, the distribution of different types of data, the correlation between data, etc., learning how to compress the data more efficiently, and then summarizing a set of optimal compression processes, i.e., the first compression logic, which includes compression strategies for different types of data (such as text and images), such as which encoding method to use for semantic-character two-way compression of text data, how to perform geometric-texture two-way compression of image data, and the specific operation steps and parameter settings for each link such as data archiving and packaging, multi-threaded adaptive lossy compression, etc. The compression processing module is reverse-learned, i.e., it analyzes and processes the compressed data, efficiently and accurately restores the original data, and determines the second decompression logic, which specifies the order and method in which the compressed data should be processed during decompression. For example, if semantic analysis and character encoding are first performed on text data during compression, the reverse steps need to be followed during decompression, first decoding the encoded characters and then restoring them based on semantic information. It also includes progressive decompression of image data.

[0033] Preferably, the file status is used as a standard, and the compression processing module is activated in a targeted manner according to the first compression logic and the second decompression logic. The file status refers to the current stage of the PDF file, which is mainly divided into the to-be-compressed state and the decompressed browsing state. The to-be-compressed state indicates that the file is original and has not been compressed and needs to be compressed for storage or transmission; the decompressed browsing state indicates that the file is already a compressed file and the user needs to decompress it and browse it on the display interface. According to the different states of the file, the compression processing module activates the corresponding logic in a targeted manner. That is, when the file is in the to-be-compressed state, the first compression logic is activated, and the compression processing module compresses the file according to the optimal compression process determined by forward learning, converting the original PDF file into a compressed file; when the file is in the decompressed browsing state, the second decompression logic is activated, and the module decompresses the compressed file according to the decompression order determined by reverse learning, and generates a pre-browsed file on the display interface to meet the user's browsing needs. In this way, the compression processing module can complete the compression and decompression tasks of the PDF file more intelligently and efficiently.

[0034] Furthermore, step S110 also includes step S114, constructing a multi-threaded compression port based on the redundancy ratio-compression lossy coefficient, wherein the redundancy ratio is determined by balancing the data value and the hot and cold coefficient; step S115, for the multi-threaded compression port, using parallel compression under format splitting-two-way decoupling as the processing logic, and performing compression learning on each thread compression port; step S116, the multi-threaded compression port after parallel integration learning is used as the second compression processing node.

[0035] Preferably, the redundancy ratio is determined by balancing data value and hot / cold coefficients, wherein data value reflects the importance of the data in the entire PDF file, such as key text content and core images, and the hot / cold coefficient reflects the frequency of data use, with frequently accessed data being hot data and cold data being cold data. By combining data value and hot / cold coefficients, the degree of redundancy in the data can be assessed. The compression loss coefficient indicates the amount of data or the degree of quality loss allowed during the compression process. Different types of data can have different compression loss coefficients. For data with high quality requirements, such as important text, the coefficient should be set smaller. For data with relatively low quality requirements, such as partial details of an image, the coefficient can be appropriately increased to achieve a higher compression ratio. Then, based on the redundancy ratio and the set compression loss coefficient, a multi-threaded compression port is constructed, i.e., multiple parallel compression channels are constructed, each of which can independently perform data compression processing. By allocating data to different ports for processing based on data type, data redundancy ratio, etc., the computing power of the multi-core processor is fully utilized to improve data compression efficiency.

[0036] Preferably, the PDF file contains multiple types of data formats, such as text, images, graphics, tables, etc. Data in different formats have different characteristics and compression requirements. Format segmentation refers to segmenting different data formats (such as text, images, tables, etc.) in the PDF file, and then compressing them separately, that is, performing parallel compression under two-way decoupling, including semantic-character two-way compression and geometric-texture two-way compression. Among them, semantic compression is processed from the meaning level of the text, such as removing redundant expressions, merging similar semantic information, etc. Character compression is to optimize the encoding of the characters themselves, such as using Huffman coding to reduce the storage space of characters; geometric compression mainly processes geometric features such as the shape and size of the image, such as adjusting the resolution of the image, scaling the image, etc. Texture compression focuses on processing the texture details of the image, reducing the file size by removing some texture information that is not easily perceived by the human eye; thereby compressing data more effectively to ensure better compression effect.

[0037] Preferably, compression learning is then performed on each thread compression port according to the processing logic, that is, by processing and analyzing a large amount of data, the parameters of the compression algorithm are continuously adjusted and the compression process is optimized to find the compression method that is most suitable for the port to process the data; finally, the multiple thread compression ports of the compression learning are integrated in parallel to obtain a second compression processing node, and the various thread compression ports continue to work in parallel to collaboratively complete the compression task of the entire PDF file, thereby achieving fast and high-quality compression of the PDF file and improving compression efficiency and quality.

[0038] In step S200, a PDF file is scanned and the file data is interpreted. According to the data value and the hot / cold coefficient, the file data is hierarchically archived and packaged, a plurality of package blocks are determined, and decoupling compression is performed to determine a compressed file. The decoupling compression is a semantic-character two-way compression of text data and a geometric-texture two-way compression of image data.

[0039] Preferably, scanning a PDF file means scanning and viewing the PDF file in detail page by page and element by element to accurately locate all data elements in the file, such as text paragraphs, various images, complex graphics, and different tables; then interpreting the file data, that is, conducting an in-depth analysis of the file's structure and encoding rules, so as to identify different types of data and extract key information. For example, for text data, determine its font, font size, color and other format information; for image data, clarify its resolution, color mode and other attributes; then perform hierarchical archiving and packaging of file data according to data value and hot and cold coefficients, including classifying high-value, high-hot coefficient data into a high level, such as the title of the file, important conclusions, etc., classifying medium-value and frequently used data into a middle level, and classifying low-value, low-use-frequency data into a low level; then encapsulate the data at the same level to form multiple independent encapsulation blocks.

[0040] Preferably, decoupled compression is performed on the encapsulated blocks, including semantic-character two-pass compression of text data and geometric-texture two-pass compression of image data. Specifically, semantic compression uses natural language processing technology to analyze the grammatical structure and semantic relationships of the text, focusing on understanding and processing the meaning of the text to remove redundant information and retain the core semantics. For example, some repeated sentences are simplified, synonyms and near-synonyms are merged, and modifiers that have little impact on the overall semantics are removed; thereby reducing the number of characters in the text, thereby achieving compression. Character compression is based on the optimization of character encoding to reduce data storage space. For example, using Huffman coding, each character is assigned a code of different lengths based on the frequency of its occurrence in the text. Frequently occurring characters use shorter codes, while less frequently occurring characters use longer codes. After encoding, the original longer text has a reduced overall data volume. Semantic compression and character compression work together to first streamline the text content at the semantic level and then further optimize it at the character encoding level, achieving two-pass compression and improving compression efficiency.

[0041] Preferably, geometric compression mainly processes the geometric features of the image, such as size and shape, including adjusting the resolution of the image, scaling the image, etc. For example, if the original resolution of the image is too high, the resolution can be appropriately lowered to reduce the number of pixels in the image, thereby reducing the size of the image file. Texture compression refers to analyzing the texture details of the image, removing subtle texture information by quantizing and filtering the texture, and reducing the storage space of texture data. For example, wavelet transform is used to decompose the image into sub-bands of different frequencies, and then the high-frequency sub-bands are appropriately quantized to remove unimportant information. Geometric compression and texture compression are combined to process the image to achieve effective compression of the image data. After completing the decoupling compression of each encapsulation block, the compressed encapsulation blocks are combined to form the final compressed file.

[0042] Furthermore, step S200 also includes step S210, determining the content architecture system of the PDF file by interpreting the file data according to the first pre-processing node; step S220, performing a hierarchical archiving based on the data value for the content architecture system to determine the first archiving result; step S230, performing a secondary hierarchical archiving based on the hot and cold coefficients to determine the second archiving result; step S240, archiving and reorganizing the first archiving result and the second archiving result under weighted calculation to determine multiple archiving layers and perform data encapsulation to determine multiple encapsulation blocks, wherein each encapsulation block corresponds to a redundancy ratio.

[0043] Preferably, the PDF file is processed by the first pre-processing node, that is, the file data is interpreted, including identifying various elements in the file, such as text (different fonts, font sizes, colors, paragraph formats, etc.), images (resolution, color mode, image type, etc.), tables, graphics, etc., and extracting relevant metadata information, so as to clearly understand the organization and logical structure of the data in the file, and then use these data to construct the content architecture system of the PDF file, for example, to determine the chapter hierarchy of the text, the position of the image in the page and the corresponding relationship with the text, etc.

[0044] Preferably, based on the content architecture system, the PDF file data is evaluated according to the importance of the data in the file, the degree of contribution to understanding the core content of the file, etc., to determine the value of each piece of data. For example, the title of the file, key conclusions, important data charts, etc. usually have higher data value, while the data value of auxiliary explanations and decorative elements is relatively low. Then, according to the level of data value, the data is divided into different levels, with high-value data at higher levels and low-value data at lower levels. Finally, the archiving result based on data value is obtained, that is, the first archiving result. Based on the first archiving result, the hot and cold coefficients of the data are further evaluated. For each level, the data is further subdivided according to its hot and cold coefficients. For example, in the high-value data level, hot data and cold data are divided into different sub-levels, thereby obtaining a more detailed archiving structure, that is, the second archiving result.

[0045] Preferably, according to the degree of emphasis on data value and usage frequency, corresponding weights are assigned to different levels in the first archiving result (hierarchical division based on data value) and the second archiving result (hierarchical archiving based on hot and cold coefficients), and the two archiving results are integrated and adjusted through weighted calculation to redefine the hierarchical relationship of the data and form multiple archiving layers; the data in the same archiving layer is then encapsulated, that is, related data is combined together to form independent encapsulation blocks, each encapsulation block contains data with similar characteristics (such as similar value and usage frequency), and each encapsulation block corresponds to a redundancy ratio. If the data in the encapsulation block contains more repeated or removable redundant information, its redundancy ratio is higher; otherwise, the redundancy ratio is lower.

[0046] Furthermore, step S240 also includes step S241, determining a first encapsulation block, wherein the first encapsulation block is any one of the multiple encapsulation blocks; step S242, determining a first redundancy coefficient for the first encapsulation block based on the archiving layer of the first archiving result; step S243, determining a second redundancy coefficient based on the archiving layer of the second archiving result; step S244, determining a first redundancy ratio based on the first redundancy coefficient and the second redundancy coefficient.

[0047] Preferably, any one of the multiple encapsulation blocks is selected as the first encapsulation block, and the first redundancy coefficient is determined according to the archiving layer of the first archiving result to which it belongs. Specifically, the redundancy degree of data may be different depending on its value. High-value data (such as core conclusions, key data) is the core of the file, with almost no part that can be deleted, and usually has low redundancy. Low-value data (such as some auxiliary instructions, repeated examples) may have more redundant information. Therefore, according to the data value reflected by the archiving layer where the first encapsulation block is located, the redundancy degree of the data in the encapsulation block is evaluated to determine the first redundancy coefficient. The larger the value, the higher the redundancy degree of the encapsulation block based on the data value level.

[0048] Preferably, a second redundancy coefficient is also determined for the first encapsulated block based on its position in the archive layer based on the second archiving result. The frequency of data use is also correlated with the degree of redundancy. Hot data (frequently accessed data) is carefully organized and lacks redundancy, while cold data (rarely accessed data) contains unnecessary information or duplicate content. Therefore, based on the archive layer based on the hot-cold coefficient where the first encapsulated block is located, the redundancy of the data within the encapsulated block is evaluated to determine the second redundancy coefficient. A larger value indicates a higher degree of redundancy for the encapsulated block based on the frequency of data use. Finally, based on the emphasis on data value and frequency of use, different weights are assigned to the first and second redundancy coefficients, and a weighted average is used to calculate a first redundancy ratio, which comprehensively measures the redundancy of the data within the first encapsulated block. A higher redundancy ratio indicates that the encapsulated block has greater room for compression, allowing for a more aggressive compression strategy to reduce file size.

[0049] Furthermore, step S200 also includes step S250, importing the multiple encapsulation blocks into the second compression processing node, activating the multi-threaded compression port by identifying the redundancy ratio, and importing the target compression port, wherein the target compression port includes at least one; step S260, the target compression port identifies the imported encapsulation blocks, performs data structure identification in the encapsulation blocks, and determines the first text data and the second image data; step S270, performs semantic-character decoupling on the first text data, and determines the first compression two-way; step S280, performs geometric-texture decoupling on the second image data, and determines the second compression two-way; step S290, performs concurrent compression processing according to the first compression two-way and the second compression two-way, and determines the compressed encapsulation block; step S2100, integrates the compressed encapsulation blocks output by each target compression port as the compressed file.

[0050] Preferably, multiple encapsulated blocks are imported into a second compression processing node, and then the redundancy ratio corresponding to each encapsulated block is identified. Based on the identified redundancy ratio, a multi-threaded compression port is activated, wherein the multi-threaded compression port is a plurality of parallel processing channels, each of which can independently compress data. After the multi-threaded compression port is activated, the encapsulated block is imported into a suitable target compression port based on the redundancy ratio size, data type, etc., and there is at least one target compression port, which can realize parallel processing of multiple encapsulated blocks and improve compression efficiency. After receiving the encapsulated block, each target compression port identifies its data structure, distinguishes whether the data in the encapsulated block is text data or image data, and marks the text data in the encapsulated block as first text data and the image data as second image data, so that different compression strategies can be adopted for different types of data.

[0051] Preferably, for the first text data, a semantic-character decoupling operation is performed. Specifically, semantic decoupling is to analyze the meaning of the text, remove redundant expressions, and extract core semantic information; character decoupling is to optimize the character encoding of the text to reduce the storage space occupied by the characters, thereby determining the first compression two-way process, that is, forming a complete compression strategy for the text data, including first performing semantic processing and then performing character processing. For the second image data, a geometric-texture decoupling operation is performed, wherein geometric decoupling mainly processes geometric features such as the shape and size of the image, such as adjusting the image resolution and performing image scaling; texture decoupling focuses on the texture details of the image, removing texture information that is not easily perceived by the human eye to reduce the file size, thereby determining the second compression two-way process, that is, forming a complete compression strategy for the image data, including geometric processing and texture processing. Based on the first compression round-trip (for text data) and the second compression round-trip (for image data), the target compression port compresses the text and image data in the encapsulated block simultaneously, that is, concurrent compression, compressing the data in the encapsulated block to obtain a compressed encapsulated block. The data volume is reduced compared to the original encapsulated block; finally, the compressed encapsulated blocks output by each target compression port are integrated to form a complete compressed file, ultimately achieving efficient compression of the entire PDF file.

[0052] In step S300, as the destination receives the compressed file, the compression processing module is called and reversely activated, reverse decompression based on the compression logic is performed, and a pre-browsing file is generated on the display interface, wherein the pre-decompression of high-value data and the progressive decompression of image data are used as decompression constraints.

[0053] Preferably, the destination refers to the end that receives data, which can be a user's computer, mobile phone, tablet computer and other devices. The destination device receives the PDF compressed file sent from the source end through a network (such as the Internet, local area network, etc.), and then calls the compression processing module and performs reverse activation. The compression processing module is used to compress files at the source end and to decompress files at the destination end. At the source end, the compression processing module performs forward activation, that is, first performs hierarchical archiving and packaging of data, and then performs multi-threaded adaptive loss compression and other steps to compress the file; at the destination end, the compression processing module enters a working mode opposite to the compression process, so that it processes the received data according to the reverse logic. Specifically, after deactivating the compression processing module, the reverse operation is performed based on the compression logic determined by the source. For example, if the source performs semantic-character two-pass compression on text data, decompression must first restore the character encoding to its original state, and then restore the original text content based on the semantic information. For image data, if geometry-texture two-pass compression was performed during compression, decompression must first restore the texture information, and then restore the image's geometric shape and size. Finally, the decompressed data is converted into a format suitable for presentation on a display interface (such as a computer monitor or mobile phone screen), generating a preview file, a fast-loading preview version of the original PDF file, allowing users to quickly see the general content of the file.

[0054] Preferably, the pre-decompression of high-value data and the progressive decompression of image data are used as decompression constraints. Specifically, when decompressing compressed file data, high-value data is pre-decompressed first so that users can quickly obtain key information. For example, in a PDF report, the title, core conclusions, important data charts, etc. of the file are high-value data, and they are decompressed first so that they can be quickly displayed on the screen, allowing users to understand the core points of the file. Since image data generally occupies a large storage space, if it is completely decompressed at one time, it may cause the file loading time to be too long, affecting the user experience. Therefore, progressive decompression is adopted, that is, the low-resolution version of the image or basic outline information is first decompressed so that the user can quickly see the general appearance of the image. Then, when the user performs operations such as zooming in on the image, the higher resolution and finer texture information of the image are gradually decompressed. The user can quickly browse the file and obtain a satisfactory visual effect when a clear image is needed, thereby effectively improving the decompression efficiency and the user's browsing experience.

[0055] Furthermore, step S300 also includes step S310, calling the compression processing module to the physical data end; step S320, temporarily determining the end type of the physical data end by identifying the file status of the PDF file; step S330, determining the activation direction of the compression processing module according to the end type, which includes forward activation and reverse activation, forward activation performs compression processing, and reverse activation performs decompression processing.

[0056] Step S330 further includes step S331, if the file status is compressed, the end type is identified as the destination end, and decompression is used as the processing method; step S332, if the file status is uncompressed, the end type is identified as the source end, and compression is used as the processing method.

[0057] Preferably, the compression processing module is called to the physical data terminal, i.e., loaded from the original storage or deployment location (possibly a cloud server) into the device memory of the physical data terminal, so that it can run on the device and process the PDF file. The file state of the PDF file is identified, including an uncompressed original state, i.e., the file exists in its originally created or stored form, which may be large and has not yet been compressed; and a compressed state, i.e., the file has been compressed and has a smaller size for easier storage and transmission. The terminal type of the physical data terminal is then determined based on the identified PDF file state. Specifically, if the PDF file held by the physical data terminal is in an uncompressed original state, the physical data terminal is temporarily determined as a source terminal (the terminal that generates or sends information) and is used to compress the file for subsequent storage or transmission. Conversely, if the physical data terminal receives a compressed PDF file, it is temporarily determined as a destination terminal (the terminal that receives information) and is used to decompress the compressed file to browse the file content.

[0058] Preferably, the activation direction of the compression processing module is determined according to the terminal type. Specifically, if it is detected that the PDF file stored or received by the physical data terminal has not been compressed, that is, the PDF file on the physical data terminal is in the original state, the file size remains the size when it was created, and the data has not undergone any form of reduction or encoding optimization, the physical data terminal is determined to be the source terminal, indicating that the uncompressed PDF file needs to be processed to facilitate storage and transmission. The compression processing module is then positively activated, and the PDF file is processed according to the preset compression logic, such as data interpretation and analysis of data characteristics and value, and then hierarchical archiving and packaging are performed, and then decoupling compression is performed (such as semantic-character two-way compression of text data and geometric-texture two-way compression of image data, etc.), and finally the original PDF file is converted into a compressed file to reduce file size and improve storage and transmission efficiency.

[0059] Preferably, if it is detected that the PDF file stored or received by the physical data end has been compressed, that is, the file exists in a compressed format, such as the file size is significantly reduced compared to the original state, or the file format is a format processed by a specific compression algorithm, the physical data end is determined to be the destination end, indicating that the received file is a compressed PDF file, and the purpose is to restore it to its original browsable state, so the compression processing module is reversely activated, and the compression processing module decompresses the compressed file according to the logic opposite to the compression process, and finally generates a browsable file on the display interface by performing reverse operations based on the compression logic, such as restoring the semantics and character information of the compressed text data, and restoring the geometric and texture features of the compressed image data.

[0060] Furthermore, step S300 also includes step S340, performing pre-decompression deployment according to the data value gradient to determine the first decompression condition; step S350, using progressive decompression as the image data decompression method to determine the second decompression condition; step S360, reversely activating the compression processing module, introducing the first decompression condition and the second decompression condition, performing decompression processing on the compressed file, and gradually restoring the interface of the PDF file under the execution sequence of the display interface to determine the pre-browsed file.

[0061] Preferably, the data value is evaluated based on factors such as the importance of the data and the criticality to understanding the file content, and the data is divided into different levels to form a data value gradient. Based on the data value gradient, pre-decompression deployment is performed, including giving priority to pre-decompression of high-value data, such as determining the storage location of high-value data in the compressed file, allocating corresponding computing resources, etc., and then determining the first decompression condition, which mainly stipulates under what circumstances and in what manner the high-value data is decompressed. For example, when a user requests to browse a file, the encapsulation block where the high-value data is located is first quickly decompressed, or the computing resources and time limits required for decompression of high-value data are stipulated.

[0062] Preferably, progressive decompression means first quickly decompressing a low-resolution version of the image or basic outline information, so that the user can see the approximate appearance of the image in a short time. When the user zooms in on the image, etc., it gradually decompresses higher resolution and finer texture information, and then determines the second decompression condition, including setting specific steps for progressive decompression, such as the resolution ratio increased each time decompression, the time interval between decompression stages, and the circumstances under which higher resolution decompression is triggered. For example, it is stipulated that the initial decompression is a 1 / 4 resolution version of the image, and when the user hovers the mouse over the image, it is further decompressed to 1 / 2 resolution, and decompressed to the original resolution after clicking the image.

[0063] Preferably, when a compressed file needs to be decompressed, the compression processing module used to compress the file is reversely activated, that is, the compression processing module is put into a decompression working mode, and the compressed file is processed according to the logic opposite to the compression process. During the decompression process, the first decompression condition (pre-decompression condition for high-value data) and the second decompression condition (progressive decompression condition for image data) are applied to the decompression process, including restoring the original state of the data according to the type of data (text, image, etc.) and the compression logic, and gradually restoring the interface of the PDF file in chronological order (time sequence) on the display interface. First, the decompressed high-value data, such as the title and core conclusion of the file, is displayed, and then as the image data is progressively decompressed, clearer images and other data are gradually displayed. The user can see the approximate content of the file in a short time and obtain a more complete and clear browsing experience with other operations; finally, the pre-browsed file is presented on the display interface.

[0064] The above specific embodiments do not constitute a limitation to the scope of protection of this application. It should be understood by those skilled in the art that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of this application should be included in the scope of protection of this application. In some cases, the actions or steps recorded in this application can be performed in an order different from that in the embodiments and can still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A PDF file browsing method based on image compression and fast decompression technology, characterized in that: The method comprises: At the source end, the compression processing module is called and forward activated; Scan PDF files and interpret the file data. Based on the data value and hot / cold coefficient (where the hot / cold coefficient reflects the frequency of data use, with frequently accessed data being hot data and cold data being cold data), the file data is hierarchically archived and packaged. Multiple packaged blocks are determined and decoupled compression is performed to determine a compressed file. Decoupled compression includes semantic-character two-pass compression of text data and geometry-texture two-pass compression of image data. As the destination receives the compressed file, the compression processing module is called and reversely activated to perform reverse decompression based on the compression logic, and a pre-browsing file is generated on the display interface, wherein the pre-decompression of high-value data and the progressive decompression of image data are used as decompression constraints.

2. The PDF file browsing method based on image compression and fast decompression technology according to claim 1, characterized in that: Before forward activation of the compression processing module, construct the compression processing module, including: The compression processing module is constructed by using data archiving and packaging as a first pre-processing node and adaptive lossy compression based on multi-threading as a second compression processing node; Deploy the compression processing module on a cloud processor; At the physical data end, the compression processing module is called and activated from the cloud processor.

3. The PDF file browsing method based on image compression and fast decompression technology according to claim 2, characterized in that: Constructing the compression processing module includes: Performing forward learning on the compression processing module to determine a first compression logic; Performing reverse learning on the compression processing module to determine a second decompression logic; Based on the file status, the compression processing module is activated in a directed manner based on the first compression logic and the second decompression logic.

4. The PDF file browsing method based on image compression and fast decompression technology according to claim 3, characterized in that: The second compression processing node is an adaptive lossy compression based on multi-threading, including: Constructing a multi-threaded compression port based on a redundancy ratio-compression lossy coefficient, wherein the redundancy ratio is determined by balancing data value and hot and cold coefficients; For the multi-threaded compression port, compression learning is performed on each thread compression port using parallel compression under format splitting-two-way decoupling as the processing logic; The multi-threaded compression port after parallel ensemble learning serves as the second compression processing node.

5. The PDF file browsing method based on image compression and fast decompression technology according to claim 1, characterized in that: Call the compression processing module, including: Calling the compression processing module to the physical data end; Temporarily determining the terminal type of the physical data terminal by identifying the file status of the PDF file; The activation direction of the compression processing module is determined according to the terminal type, including forward activation and reverse activation. Forward activation performs compression processing, and reverse activation performs decompression processing.

6. The PDF file browsing method based on image compression and fast decompression technology according to claim 5, characterized in that: If the file state is compressed, the end type is identified as a sink end, and decompression is used as a processing method; If the file status is a non-compressed state, the end type is identified as a source end, and compression is used as a processing method.

7. The PDF file browsing method based on image compression and fast decompression technology as claimed in claim 2, characterized in that: According to the data value and hot / cold coefficient, the file data is archived and packaged in a hierarchical manner, including: Determining the content structure of the PDF file by interpreting the file data according to the first pre-processing node; For the content architecture system, perform a hierarchical archiving based on data value to determine a first archiving result; Perform secondary stratification archiving based on the cold and hot coefficients to determine the second archiving result; The first archiving result and the second archiving result are archiving reorganized under weighted calculation to determine multiple archiving layers and perform data encapsulation to determine multiple encapsulation blocks, wherein each encapsulation block corresponds to a redundancy ratio.

8. The PDF file browsing method based on image compression and fast decompression technology according to claim 7, characterized in that: Each encapsulation block corresponds to a redundancy ratio, including: determining a first encapsulation block, wherein the first encapsulation block is any one of the plurality of encapsulation blocks; determining, for the first encapsulation block, a first redundancy coefficient using an archiving layer based on the first archiving result; determining a second redundancy coefficient based on the archiving layer of the second archiving result; A first redundancy ratio is determined according to the first redundancy coefficient and the second redundancy coefficient.

9. The PDF file browsing method based on image compression and fast decompression technology as claimed in claim 4, characterized in that: Perform decoupling compression to determine the compressed files, including: Importing the plurality of encapsulated blocks into the second compression processing node, activating the multi-threaded compression port by identifying a redundancy ratio, and importing the multi-threaded compression port into a target compression port, wherein the target compression port includes at least one; The target compression port identifies the imported encapsulation block, performs data structure identification in the encapsulation block, and determines the first text data and the second image data; performing semantic-character decoupling on the first text data to determine a first compression round trip; performing geometry-texture decoupling on the second image data to determine a second compression round-pass; performing concurrent compression processing based on the first compression round-trip and the second compression round-trip to determine a compression encapsulation block; The compressed encapsulation blocks output by each target compression port are integrated as the compressed file.

10. The PDF file browsing method based on image compression and fast decompression technology according to claim 1, characterized in that: Perform reverse decompression based on compression logic and generate a preview file on the display interface, including: Perform pre-decompression deployment based on the data value gradient and determine the first decompression condition; Determining a second decompression condition by using progressive decompression as the image data decompression mode; The compression processing module is activated in reverse, the first decompression condition and the second decompression condition are introduced, the compressed file is decompressed, the interface of the PDF file is gradually restored under the timing of the display interface execution, and the pre-browsed file is determined.

Citation Information

Patent Citations

  • Interaction-oriented editable art image generation system and method

    CN117036528A

  • Article analysis method and device

    CN118780277A