Loading optimization method for PDF (Portable Document Format) instant preview
By dividing PDF files into loading units by pages and capturing user tag operations, generating a collection of tag pages to store them in an independent tag file library, solving the problem that the files after tagging in PDF instant preview are larger and the loading time is increased, and fast loading and efficient storage are achieved.
Patent Information
- Application Number
- CN202510714678.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, PDF instant preview adopts full parsing and lazy loading strategies to embed user tag content into the original document, resulting in the overall larger file after tagging, increasing loading time, affecting cache consistency and loading efficiency.
Divide PDF files into loading units by pages, capture user tag operations, generate a set of tag pages and store them in an independent tag file library. When the user closes and opens them again, use a combined loading process to read the tag page from the tag file library to overwrite the corresponding page of the PDF file for preview.
It significantly improves the loading speed of the tagged page, reduces storage resource redundancy, enhances the efficiency of multiple terminals, and optimizes PDF preview performance.
Smart Images

Figure CN120234487A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of graphic data reading, and particularly to a method for optimizing the loading of PDF instant preview. Background Art
[0002] PDF instant preview usually adopts the strategies of full parsing and lazy loading. Full parsing means loading the entire PDF file into memory at one time, and then parsing and rendering the content of each page; lazy loading is a strategy of loading on demand, that is, only loading the page content that the user currently needs to view, and dynamically loading other pages when the user scrolls to them. When the user performs operations such as color marking or annotation on the document, these marked contents are usually directly embedded in the original PDF file, resulting in a sharp increase in the file size and a significant reduction in the loading speed. Especially when there are many marked contents and the file is large, it affects cache consistency and loading efficiency.
[0003] In summary, in the prior art, there are technical problems that due to the fact that PDF instant preview mostly adopts the strategies of full parsing and lazy loading, and embeds the user's marked content into the original document, the overall size of the marked file becomes larger, the loading time increases, and further affects cache consistency and loading efficiency. Summary of the Invention
[0004] The purpose of this application is to provide a method for optimizing the loading of PDF instant preview, so as to solve the technical problems in the prior art that due to the fact that PDF instant preview mostly adopts the strategies of full parsing and lazy loading, and embeds the user's marked content into the original document, the overall size of the marked file becomes larger, the loading time increases, and further affects cache consistency and loading efficiency.
[0005] In view of the above problems, this application provides a method for optimizing the loading of PDF instant preview. Among them, the method for optimizing the loading of PDF instant preview includes: dividing the PDF file for instant preview into loading units by page; capturing the marking operations of the user previewing the PDF file to obtain a set of marked pages; storing the set of marked pages in an independent marked file library, and the independent marked file library is grouped and managed according to the user's ID and the ID of the PDF file; when the user closes the PDF file, if the user opens the same PDF file again, start a combined loading process, and the combined loading process includes reading the marked pages from the independent marked file library according to the user's ID, covering the corresponding pages of the PDF file according to the marked pages, and outputting a combined PDF file for the user to preview.
[0006] Optionally, the marking operations include at least one of color highlighting, graffiti annotation, text annotation, underline, strikethrough, or graphic marking.
[0007] Optionally, capture the marking operations of the user previewing the PDF file, extract the marked text fields; identify the text fields of the PDF file based on the marked text fields, and obtain the set of the same text fields in the PDF file; if the user activates the same-field marking instruction, when the user marks any text field, synchronously mark the set of the same text fields.
[0008] Optionally, build a field matching model based on text hash comparison; call the field matching model to perform synonymous recognition on the PDF file with the marked text fields to generate a set of synonymous fields; if the user activates the same-field marking instruction, when the user marks any text field, synchronously mark the set of the same text fields and the set of synonymous fields.
[0009] Optionally, record the first marking operation of the user previewing the PDF file at the previous timestamp, and the first marking data corresponding to the first marking operation; record the second marking operation of the user previewing the PDF file at the next timestamp, and the second marking data corresponding to the second marking operation; if the first marking data and the second marking data are repeated, use the second marking data to overwrite the first marking data for preview display.
[0010] Optionally, capture the marking operations of the user previewing the PDF file according to the timestamp ID, and store the version marking data according to the timestamp ID.
[0011] Optionally, capture the marking operations of the user previewing the PDF file, and extract the marking data, where the marking data includes page number, marking color, marking area position, marking timestamp, and marking type; convert the marking data into structured data and store it in an independent marking file library, and the independent marking file library is grouped and managed according to the ID of the user, the ID of the PDF file, and the ID of the loading unit.
[0012] Optionally, store the set of marked pages in the independent marking file library; where the independent marking file library is a hierarchical storage structure, used to create an index directory according to the ID of the PDF file, and store it according to the ID of the user under each directory.
[0013] Optionally, the combined loading process includes reading the marked pages and the ID of the loading unit from the independent marking file library according to the ID of the user, and reading the marked layer data; overlaying the marked layer data in a vector manner at the corresponding page position of the PDF file, and outputting a combined PDF file for the user to preview.
[0014] Optionally, a set of marked pages and a set of unmarked pages are obtained; when outputting a combined PDF file, the loading priority of the set of marked pages is higher than that of the set of unmarked pages.
[0015] The technical solutions provided in this application have at least the following beneficial effects: By dividing the instant preview PDF file into loading units page by page; capturing the marking operations of the user previewing the PDF file to obtain a set of marked pages; storing the set of marked pages in an independent marked file library, and the independent marked file library is grouped and managed according to the user ID and the PDF file ID; when the user closes the PDF file, if the user opens the same PDF file again, a combined loading process is started, and the combined loading process includes reading the marked pages from the independent marked file library according to the user ID, covering the corresponding pages of the PDF file according to the marked pages, and outputting a combined PDF file for the user to preview. That is to say, by capturing the marking operations of the user, automatically detecting the pages that have changed, extracting the marked data, storing the marked data separately in an independent marked file library, separating it from the original PDF document, only including the pages with marking changes and their marked content. When the user opens the same PDF file again, the combined loading method is used to optimize the PDF preview performance, significantly improving the loading speed of the marked pages, reducing the redundancy of storage resources, and enhancing the multi-terminal reuse efficiency.
[0016] The above description is only an overview of the technical solutions of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specifically gives the specific implementation manners of this application. It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of this application, nor is it used to limit the scope of this application. Other features of this application will become easily understood through the following specification. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0018] Figure 1 It is a schematic flowchart of the loading optimization method for instant preview of PDF in this application.
[0019] Figure 2This is a flowchart showing the process of capturing user marking operations in the PDF instant preview loading optimization method of the present application. Specific embodiments
[0020] By providing a PDF instant preview loading optimization method, the present application solves the technical problems in the prior art. Since PDF instant previews mostly adopt full - scale parsing and lazy - loading strategies, embedding user - marked content into the original document results in an increase in the overall size of the marked file, an increase in loading time, and further affects cache consistency and loading efficiency. By capturing the user's marking operations, automatically detecting the pages where changes occur, extracting the marked data, and storing the marked data separately in an independent marked file library, separated from the original PDF document, only including the pages where marking changes occur and their marked content. When the user opens the same PDF file again, a combined loading method is used to optimize the PDF preview performance, significantly improving the loading speed of the marked pages, reducing the redundancy of storage resources, and enhancing the multi - terminal reuse efficiency.
[0021] Next, the technical solutions in the present application will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described here. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application. Additionally, it should be noted that for the sake of description, only parts related to the present application are shown in the accompanying drawings, not all of them.
[0022] Embodiment. Please refer to the attached Figure 1 The present application provides a PDF instant preview loading optimization method. Among them, the PDF instant preview loading optimization method specifically includes the following steps: S100: Divide the PDF file for instant preview into loading units by page.
[0023] Specifically, instant preview means that when a user views a PDF file, they can quickly see the content of the document. Usually, it is based on the pre - rendering or fast - loading mechanism of the file, without the need to fully download or parse the entire document, and can display the file content in a short time.
[0024] During instant preview, large files or complex documents usually need to be processed. In this case, the loading and display of the file will become very slow. Especially when the network is unstable or the device performance is low, the loading delay will affect the user experience. Regarding each page of the PDF file as a separate loading unit, when loading, the original PDF pages are loaded according to the priority of each page, that is, the corresponding page content is dynamically loaded according to the user's page - turning operation, which is suitable for large PDF files that require quick preview.
[0025] S200: Capture the marking operations of the user previewing the PDF file to obtain a marked page set.
[0026] The marking operations include at least one of color highlighting, scribbling annotation, text circling, underlining, strikethrough, or graphic marking.
[0027] Specifically, when the user views the PDF file, the marking operations performed by the user on the document are monitored and recorded in real time, including color highlighting, scribbling annotation, text circling, underlining, strikethrough, or graphic marking, etc., to obtain a marked page set. The marked page set includes all the page sets where the user has performed marking operations. Whenever the user marks a page (such as highlighting, circling, etc.), that page is regarded as a marked page and added to the marked page set for subsequent processing and loading.
[0028] Marking operations are various modification or annotation operations performed by the user on certain areas or elements in the PDF document. For example, color highlighting is that the user uses color to highlight the text content in the document, scribbling annotation is that the user adds annotation text or graphics on the page in a freehand drawing way, text circling is that the user circles specific text or paragraphs in the document, underlining is that the user underlines the text, strikethrough is that the user adds a strikethrough to the text, and graphic marking is that the user draws shapes or graphics (such as rectangles, circles, etc.) in the document.
[0029] Implement an event listening mechanism to capture in real time the marking operations performed by the user during the preview, including any visibility changes of the PDF file by the user, such as adding highlights to the text, drawing a circle on a certain area, or inserting a freehand-drawn graphic, etc. Whenever the user performs a marking operation, marking data is generated according to the type and position of the mark. For example, when the user draws a rectangle on the document, record information such as the coordinates, dimensions, and color of this rectangle and store it in an independent mark file library. Each time the user performs a marking operation, check whether the operation has changed the page content. If so, add the page to the marked page set and generate new marking data for it. In this way, the marked page set is continuously updated and always contains all the pages marked by the user.
[0030] Marked data is usually stored in a structured format, such as JSON or XML files, which record detailed information such as the type, location, and color of the marking operations on each page, facilitating subsequent combined loading and display. When the user performs marking operations, these operations are recorded and associated with a specific page, implemented through the built-in application of the PDF, which converts the user's operations into a data format that can be stored. Monitor the user's operations on the PDF file, such as mouse clicks, selecting certain text areas, applying color highlights, or scribbling. When the user performs a marking operation, capture the specific content of this operation (such as location, type of mark, color, etc.) and determine the text area of the mark. For example, assume a PDF document has 5 pages, the user highlights in red on the first page, adds a text annotation on the third page, and draws a graphic on the fifth page. These marking operations are captured and stored as a set of marked pages, which includes the type, location, and page information involved in each marking operation. The set of marked pages includes the following data: First page: Red highlight, location (100, 200) to (150, 250); Third page: Text annotation, "This is an important point", location (150, 300); Fifth page: Graphic mark, circle, center point (200, 400), radius 50 pixels.
[0031] By capturing and recording each marking operation of the user on the document in real time, ensure that the marked data of each page is accurately stored and managed. By recording the marking operations, easily trace back the user's modification history of the document.
[0032] Furthermore, as shown in the appendix Figure 2 it is shown that the present application S200 includes: S210: Capture the marking operation of the user previewing the PDF file and extract the marked text fields; S220: Identify the text fields of the PDF file based on the marked text fields and obtain the set of the same text fields in the PDF file; S230: If the user activates the same-field marking instruction, when the user marks any text field, synchronously mark the set of the same text fields.
[0033] Specifically, extract the marked text field of the user's marking operation during preview, that is, the selected text area or field. For example, the user may select a certain piece of text and mark it with highlighting, underlining, etc. This piece of text becomes a marked text field. Analyze the area marked by the user, extract the text field from it, and identify the text in this area through the PDF text extraction algorithm. The extracted text field can be a certain paragraph of text, a word, or a sentence. The PDF text extraction algorithm is a computer program specifically used to extract text content from PDF documents, which can analyze the structure of PDF files and extract the text content therein for further processing or analysis. Copy the text from the original content of the PDF document, or if the text exists in the form of an image (such as in a scanned document), optical character recognition (OCR) technology is required to identify the text in the image.
[0034] Extract the actual text content from the marked text field of the user, perform text recognition, and search for the same text field throughout the PDF document. Through simple string matching, retrieve the same text field throughout the PDF file to obtain a set of the same text fields in the PDF file. Traverse the entire PDF document, check the text content page by page. After retrieving the same text fields on all pages, including the same text fields on different pages, integrate these text fields to obtain a set of the same text fields.
[0035] If the user starts the same-field marking instruction, automatically synchronously mark all text fields in the document that are the same as the text content marked by the user. Specifically, after the user starts the same-field marking instruction, according to the marking operation performed by the user, traverse the set of the same text fields and perform synchronous marking. All areas with the same content as this text field will be subjected to the same marking operation, and the user does not need to manually mark each identical text field.
[0036] By identifying and synchronously marking all the same text fields, the user does not need to manually mark each area where the same text appears, significantly improving the marking efficiency and convenience. The user only needs to mark the text once, and the system will intelligently complete the remaining marking work, making the document marking operation smoother and more intelligent, and improving the user experience.
[0037] Furthermore, the present application further includes the following steps: S240: Construct a field matching model based on text hash comparison; S250: Invoke the field matching model to perform synonymous recognition on the PDF file with the marked text field to generate a set of synonymous fields; S260: If the user starts the same-field marking instruction, when the user marks any text field, synchronously mark the set of the same text fields and the set of synonymous fields.
[0038] Specifically, a field matching model is constructed based on text hash comparison. Text hash comparison is a process of converting text data into a fixed-length value (hash value) through a hash function. The hash value is used to represent the uniqueness of the text content, and when two text contents are the same or similar, their hash values should also be the same or similar. The field matching model refers to an algorithm or method for finding similar or identical parts among multiple text fields (such as paragraphs or table contents in a PDF document), and can identify similar texts based on direct content matching, similarity calculation, or semantic analysis.
[0039] Obtain a large amount of text data, perform data preprocessing on the text data, and perform unified formatting processing on the text, such as converting all letters to lowercase to avoid different hash values due to inconsistent case. When necessary, perform stemming (such as removing the plural form, tense changes, etc. of English words) or synonym normalization processing. Divide the text data into independent text fields according to a certain logic (such as by paragraph, by line, by table cell, etc.). For each text field, remove irrelevant parts such as punctuation marks, spaces, and special characters, and retain valid lexical information. Perform word segmentation on the text data (especially for languages such as Chinese), so that the content of each field becomes a sequence of words.
[0040] Decompose each text field into words, phrases, or other basic semantic units (such as n-gram) as features. Each vocabulary is assigned a weight, usually using term frequency (TF) or TF-IDF (term frequency-inverse document frequency) as the weight. More important vocabulary will have a greater impact on the final hash value. Term frequency is the frequency of a word appearing in a document, indicating the number of times the word appears in the document. The basic idea of term frequency is that words with a high frequency of occurrence may have a higher importance in the current document. TF-IDF is a weighting method that combines term frequency (TF) and inverse document frequency (IDF), mainly used to measure the importance of a word in a certain document. Inverse document frequency (IDF) measures the prevalence of a word in the entire document collection. The smaller the IDF value, the more frequently the word appears in the corpus and the lower the information content.
[0041] Each vocabulary corresponds to a high-dimensional vector representation. Calculate a fixed-length vector for each vocabulary, and this vector will be mapped to an integer value (hash value) through certain rules (such as a hash function). For example, the hash value of each vocabulary may be a floating value between 0 and 1, used to represent the position and weight of the vocabulary in the text.
[0042] Hash each feature (such as a word) in the text to obtain the hash value of each feature. The length of the hash value is usually a fixed number of bits (such as 64 bits or 128 bits). Perform a weighted aggregation on each feature hash value to obtain a final hash value. The weights during aggregation are usually determined by the importance of the feature in the text (such as word frequency). Combine the hash values of all features into a vector according to certain rules, that is, synthesize a final hash value by means of weighted summation.
[0043] For each text field, store its hash value in a hash table. When field matching is required, directly compare the hash values of the fields. If the hash values of two fields are the same, it is considered that these two fields are the same and suitable for tag synchronization. For fields with different but similar hash values, calculate the Hamming distance of the hash values to evaluate their similarity. The Hamming distance refers to the number of different bits in two binary hash values. The smaller the value, the higher the similarity of the fields. Through the comparison of hash values, if it is found that the hash values of multiple fields are very close (that is, the Hamming distance is less than a preset threshold), it is considered that they are synonymous fields and should be grouped into one category. Adjust the threshold of text hash comparison according to actual needs. A smaller threshold will increase the matching range of similar fields but may also increase the risk of false matching. Through multiple experimental data verifications, determine the most appropriate threshold to ensure the accuracy and efficiency of the model. After the model is constructed, all field matching results will be stored in the form of a table or a database. Each field corresponds to a hash value and is marked with the set of synonymous fields it belongs to.
[0044] Extract all text fields from the PDF file, including individual words, phrases, or sentences in the document. Process each extracted text field through a text hashing algorithm to convert it into a hash value. This hash value is a binary representation of a fixed length, which can effectively represent the content of the field while avoiding directly storing long texts, improving storage and calculation efficiency. That is, perform text parsing on the PDF file to extract all text fields. Call the field matching model to convert each extracted text field into a unique hash value. By performing weighted calculations on the features of the text content, ensure that words with high frequencies of occurrence can have a greater impact on the final hash value. By performing weighted calculations on the features of the text content, ensure that words with high frequencies of occurrence can have a greater impact on the final hash value. According to the field matching model, calculate the similarity between the text content in the PDF file and the marked text fields. Fields with higher similarity (that is, fields with close hash values) are considered semantically similar or identical fields. For example, according to the Hamming distance between hash values, take the text fields corresponding to the smaller distance as text fields with similar semantics. For these similar fields, a set of synonymous fields can be generated and associated with the original marked fields.
[0045] The set of synonymous fields is identified by comparing the hash values of multiple text fields in the PDF document and belongs to a set of texts that are synonymous or semantically similar. These text fields may be syntactically different but represent the same concept or content in practical applications. For example, temperature, air temperature, and weather temperature may be multiple fields in the set of synonymous fields.
[0046] When the user performs a marking operation on the PDF document, through the above-mentioned field matching model, other text fields identical to the marked field are automatically identified and the marking operation is synchronously performed. After the user starts the same-field marking instruction, the text field currently marked by the user is checked, and the hash value of this text field is identified. By comparing the hash values of the fields, text fields identical to the currently marked text field (i.e., fields with the same or similar hash values) are automatically searched for in the PDF file. If the system detects that this field belongs to the set of synonymous fields, these fields are automatically marked.
[0047] Through the identification and synchronous marking of synonymous fields, the workload of manual repeated marking by the user is reduced. The synonymous fields are automatically identified and synchronously marked, reducing errors and omissions that may occur during manual marking and ensuring the accuracy and consistency of marking.
[0048] Furthermore, the present application further includes the following steps: S270: Record the first marking operation of the user previewing the PDF file at the previous timestamp, and the first marking data corresponding to the first marking operation; S280: Record the second marking operation of the user previewing the PDF file at the next timestamp, and the second marking data corresponding to the second marking operation; S290: If the first marking data and the second marking data are repeated, use the second marking data to overwrite the first marking data for preview display.
[0049] Specifically, a timestamp is recorded each time the user makes a mark to accurately record the moment when the event occurs. When the user performs the first marking operation on the PDF file, this timestamp, the first marking operation, and the first marking data are recorded. The first marking operation refers to the marking operation performed by the user this time, including at least one of color highlighting, graffiti annotation, text annotation, underlining, strikethrough, or graphic marking. The first marking data is the detailed information marked by the user during this marking, including the type, location, color, text content, etc. of the mark, describing the specific content marked by the user in the PDF.
[0050] Similarly, record the second marking operation of the user previewing the DF file at the next timestamp, and the second marking data corresponding to the second marking operation. Here, the first and second do not refer to the number of times the user marks, but only indicate that the content marked by the user later can overwrite the content marked by the user before.
[0051] Check whether there are duplicate or conflicting situations between the first marked data and the second marked data. By comparing the first marked data and the second marked data, determine whether there is overlapping or duplicate marked content. For example, judge the type of the mark (highlight and strikethrough), the position (whether it covers the same text paragraph), and the content of the mark (whether it is the same text). If the second marked data is the same as the first marked data (such as the positions of the marks completely coincide and the content of the marks is consistent), then these two marked data are considered duplicates. If the second marked data is not completely the same as the first marked data (such as different mark types or different mark contents), it is not regarded as a duplicate mark.
[0052] If duplicate marked data is found, perform an overwrite operation. The second marked data will overwrite the first marked data and be previewed and displayed in the document. This means that the second mark will be the final display result and the first mark will be removed. For example, assume that the strikethrough of the second mark covers the highlighted area of the first mark. If it is detected that these two marked data are exactly the same in terms of position and content, then the second mark (strikethrough) will overwrite the first mark (highlight) and finally present the strikethrough effect.
[0053] After completing the processing of the marked data, apply the final marked data to the PDF document for preview display. The preview page will show the results of all valid marking operations, including the processed marks (such as the situation where strikethrough covers highlight). By automatically judging and processing duplicate marking operations, users do not need to manually revoke or adjust the previous marks, effectively avoiding redundant marked data, reducing meaningless duplicate marks in the document, and ensuring the clarity and conciseness of the marked content.
[0054] Furthermore, the present application further includes the following steps: Capture the marking operations of the user previewing the PDF file according to the timestamp ID, and store the version marked data according to the timestamp ID.
[0055] Specifically, when the user marks the PDF file, each marking operation will be captured and the time when it occurs will be recorded. A timestamp ID will be assigned to each marking operation, usually a unique identifier generated according to the current time. The timestamp ID can adopt a time format accurate to milliseconds to ensure that each marking operation has a unique ID. Each time the user makes a mark, the detailed information of the mark will be stored together with the timestamp ID to ensure that each marking operation can correspond to a specific time point.
[0056] As the user performs marking operations multiple times, multiple versions of the marking data are generated. Each version is associated with a specific timestamp ID. Version management of the marking data is carried out through the timestamp ID to ensure that the marking data at each time point can be accurately recorded and accessed. When the user previews the document, the corresponding marked content is displayed according to the current version of the marking data and the timestamp ID.
[0057] Through the timestamp ID, each marking operation can be accurately traced, ensuring clear and accurate version management of the marking data. This allows users to more easily view and compare the marked content at different time points, avoiding marking conflicts and incorrect operations in the traditional method, and improving the user experience.
[0058] S300: Store the set of marked pages in an independent marking file library, and the independent marking file library is grouped and managed according to the ID of the user and the ID of the PDF file.
[0059] Furthermore, step S300 of the present application includes: S310: Capture the marking operation of the user previewing the PDF file, and extract the marking data, where the marking data includes the page number, marking color, marking area location, marking timestamp, and marking type; S320: Convert the marking data into structured data and store it in an independent marking file library, and the independent marking file library is grouped and managed according to the ID of the user, the ID of the PDF file, and the ID of the loading unit.
[0060] Specifically, capture all relevant information generated when the user previews the PDF file for marking, including the page number, the color of the marking, the location of the marking area, the timestamp of the marking (i.e., the time when the marking operation occurs), and the marking type (such as highlighting, underlining, scribbling, strikethrough, etc.). Once the marking data is captured, it is converted into structured data for easy storage and management. The marking data will be converted according to a certain structure, and common storage formats include JSON, XML, or tabular database formats.
[0061] Structured data refers to data organized according to a predetermined format and rules, usually having clear data fields for easy storage, retrieval, and processing. Compared with unstructured data (such as text files, images), structured data is more standardized and is usually stored in tabular, database, or JSON / XML formats. For example, when the text color is red, the structured data is mark_color:#FF0000, and when the marking type is highlighting, the structured data is mark_type:highlight.
[0062] The structured data is stored in an independent markup file library and grouped and managed according to the user ID, PDF file ID, and loading unit ID. The markup data for each loading unit (usually a single page or a part of a page in a PDF) is grouped and managed to ensure that the markup information can be quickly loaded according to the page number during loading and preview. An index is established based on the user ID, PDF file ID, and loading unit ID to speed up the query and retrieval of the markup data. For example, when a user opens a certain PDF file again, the markup data for that page is quickly retrieved and loaded according to the user ID, PDF file ID, and loading unit ID.
[0063] When the user continues to mark the same file, only the newly added markup data is stored instead of storing the entire file's markup data repeatedly, effectively avoiding data redundancy and improving storage efficiency. Through the storage of structured data, the markup data becomes easy to query, update, and manage. Different markup operations are clearly recorded in an independent markup library, preventing the continuous expansion of the original PDF file.
[0064] Furthermore, the present application further includes the following steps: S330: Store the set of marked pages into the independent markup file library; S340: Among them, the independent markup file library has a hierarchical storage structure for creating an index directory according to the ID of the PDF file and storing according to the ID of the user under each directory.
[0065] Specifically, the set of marked pages is the set of pages marked by the user when previewing a PDF file. Whenever the user marks a certain area of a page in a PDF file (such as highlighting, annotating, etc.), that page will be added to the set of marked pages. This set includes all the page information marked by the user, and the markup data for each page includes color, markup type, markup location, timestamp, etc.
[0066] After the set of marked pages is captured and organized, it is stored in the independent markup file library. To manage and query the markup data more efficiently, the independent markup file library adopts a hierarchical storage structure. First, a corresponding directory is created in the independent markup file library according to the ID of each PDF file. The ID of each PDF file is unique, ensuring that each directory in the file library corresponds to a specific PDF file.
[0067] Under the directory of each PDF file, subdirectories are created according to the IDs of different users. For example, a PDF file with the ID "pdf_001" will correspond to a directory named "pdf_001". If a user with the ID "user_001" has marked in the PDF file, a subdirectory named "user_001" is created under the "pdf_001" directory, and the marking data of this user in this PDF file is stored in this subdirectory. For multiple users in the same PDF file, their marking data will be stored in different subfolders under the same folder, and each user has an independent subfolder to store their marking data.
[0068] The marking data can be stored in the user subdirectory in the form of structured data files (such as JSON, XML, etc.). The file contains page marking data and corresponding marking attributes (such as marking type, marking color, marking position, timestamp, etc.). An index is established for each PDF file and user in the independent marking file library, so that when querying the marking data of a certain PDF file and user, the corresponding directory and subdirectory can be quickly located. By storing the marked page set in the independent marking file library and adopting a hierarchical storage structure, the marking data is efficiently managed and stored, making the query, loading, and updating of the marking data more efficient. At the same time, redundant storage is reduced, and the isolation and security of the data are ensured.
[0069] S400: When the user closes the PDF file and then opens the same PDF file again, a combined loading process is started. The combined loading process includes reading the marked pages from the independent marking file library according to the user's ID, covering the corresponding pages of the PDF file based on the marked pages, and outputting a combined PDF file for the user to preview.
[0070] Furthermore, S400 of this application includes: S410: The combined loading process includes reading the marked pages and the ID of the loading unit from the independent marking file library according to the user's ID, and reading the marked layer data; S420: Overlaying the marked layer data in vector form at the corresponding page position of the PDF file, and outputting a combined PDF file for the user to preview.
[0071] Specifically, when the user closes the PDF file and reopens it, the combined loading process is initiated. The previously marked data of the user is reloaded and combined with the original PDF file, ensuring that all previous marking operations are retained and correctly displayed when the user reopens the file. Specifically, the original PDF pages are loaded according to priority, and the marked data files corresponding to the pages are read in parallel, and the two are synthesized at the rendering layer: the original PDF page serves as the background image, and the marking information is superimposed as an overlay.
[0072] The core goal of the combined loading process is to ensure that the previous markings on the file (such as highlighting, scribbling, etc.) can be correctly displayed in the file. When the user reopens the PDF file, the independent marking file library is first queried, and the relevant marking data is read according to the user ID and the PDF file ID to obtain the set of pages marked by all users in the PDF file. Each page will contain relevant marking operation records. If the file is large, the marked content is further accurately located according to the loading unit ID. The loading unit ID is used to identify a specific page or a part of a page in the PDF file. In a long document, each page or each segment can be regarded as an independent loading unit.
[0073] The marked layer data associated with each marked page and the loading unit ID is read, including the marking type, marking position, marking color, timestamp, etc. The marked layer data includes the specific content and position of all marking operations (such as color marking, highlighting, scribbling, etc.), and is usually stored as vector layer information (such as paths, coordinates, colors, marking types, etc.). When displaying the marked content, vector graphics are used instead of bitmaps. Vector graphics can remain clear at different resolutions and can be easily scaled and superimposed. When previewing the PDF, the markings are usually superimposed on the page as a vector layer instead of directly modifying the original content.
[0074] The content of the original PDF page is displayed as the background image, and the marked layer is superimposed as an overlay on the corresponding page area. Vector graphics can very precisely control the position, shape, color, and transparency of the markings. According to the coordinates and positions of the marked data, ensure that the marked content is accurately aligned with elements such as text and graphics in the PDF page. For example, if the user highlights a certain paragraph of text on the page, according to the coordinate information of the marking, the highlighting is accurately displayed at the corresponding text position.
[0075] The combination of the marked layer and the original PDF page is output to generate a combined PDF file, including the original PDF page and the marked storage layer. The combined PDF file will be presented to the user for preview. When previewing, the user can see the content marked last time and can continue to operate on the document (such as adding new markings or modifying existing markings).
[0076] Since the marked data is stored separately from the original PDF file, the combined loading process avoids regenerating the marked data every time the file is opened. By reading the independent marked layer data and overlaying it on the PDF pages, the system can display the file content more quickly, significantly reducing the loading time. The marked layer data only stores the marking information and does not need to repeatedly store the PDF pages themselves, saving storage space, avoiding redundant storage of duplicate data, and improving storage efficiency. Through the independent marked file library and hierarchical storage method, it supports the reuse of marked data across devices and platforms. The content marked by the user on one device can be quickly loaded and displayed on another device, thus enhancing the application experience in a multi-terminal environment.
[0077] Furthermore, the present application further includes the following steps: S430: Obtain the marked page set and the unmarked page set; S440: When outputting the combined PDF file, the loading priority of the marked page set is higher than that of the unmarked page set.
[0078] Specifically, the PDF file is divided into a marked page set and an unmarked page set. The marked page set is the set of pages on which the user has performed marking operations (such as highlighting, scribbling, annotating, etc.) in the PDF file; the unmarked page set is the set of pages on which there are no user marking operations in the PDF file. Based on the page content in the PDF file and the user's marking operations, it is identified which pages have been marked by the user and which pages have not been marked.
[0079] When outputting the combined PDF file, the loading order is determined according to the different loading priorities of the marked page set and the unmarked page set. To enhance the user experience, generally, the loading priority of the marked page set should be higher than that of the unmarked page set. That is to say, when the user opens the PDF file, the pages in the marked page set are first loaded and rendered, including the user's marked content, with a higher loading priority, ensuring that the user can quickly see their previous marking operations, reducing the waiting time for the user, and allowing the user to see their marks earlier.
[0080] After the marked pages are loaded, the unmarked page set is then loaded. Since there are no user marks on these pages, they can usually be loaded later, and the delayed loading will not affect the user's mark viewing experience. The purpose of this step is to optimize the allocation of resources and avoid unnecessary waiting.
[0081] When the loading of all pages is completed, the final combined PDF file is output. The combined PDF file contains the original PDF file and the user's marked layer, and all the marked content will be overlaid at the corresponding page positions. When the file is generated, the user can see the marked pages and other unmarked content during preview.
[0082] By setting the loading priority of the marked pages to high, users can see the marked content they have made earlier without waiting for the entire PDF file to finish loading. The response speed of the file is significantly improved through the priority loading method. Especially for PDF files containing a large number of pages, users do not have to wait for the unmarked pages to finish loading, avoiding redundant waiting during loading.
[0083] In summary, the loading optimization method for PDF instant preview provided by this application has the following beneficial effects: The PDF file for instant preview is divided into loading units page by page; the marking operations of the user previewing the PDF file are captured to obtain a set of marked pages; the set of marked pages is stored in an independent marked file library, and the independent marked file library is grouped and managed according to the user ID and the PDF file ID; when the user closes the PDF file and then opens the same PDF file again, a combined loading process is started. The combined loading process includes reading the marked pages from the independent marked file library according to the user ID, covering the corresponding pages of the PDF file according to the marked pages, and outputting a combined PDF file for the user to preview. That is to say, by capturing the marking operations of the user, automatically detecting the pages that have changed, and extracting the marked data, the marked data is separately stored in an independent marked file library, separated from the original PDF document, only containing the pages with marking changes and their marked content. When the user opens the same PDF file again, the combined loading method is used to optimize the PDF preview performance, significantly improving the loading speed of the marked pages, reducing the redundancy of storage resources, and enhancing the multi-terminal reuse efficiency.
[0084] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0085] Obviously, for those skilled in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A method for optimizing the loading of instant PDF preview, characterized in that, Including: Dividing the instant preview PDF file into loading units page by page; Capturing the marking operations of the user previewing the PDF file to obtain a set of marked pages; Storing the set of marked pages in an independent marking file library, and the independent marking file library is grouped and managed according to the user ID and the PDF file ID; When the user closes the PDF file, if the user opens the same PDF file again, start a combined loading process, and the combined loading process includes reading the marked pages from the independent marking file library according to the user ID, covering the corresponding pages of the PDF file according to the marked pages, and outputting a combined PDF file for the user to preview.
2. The loading optimization method for PDF instant preview according to claim 1, characterized in that The marking operations include at least one of color highlighting, graffiti annotation, text annotation, underlining, strikethrough, or graphic marking.
3. The loading optimization method for PDF instant preview according to claim 1, characterized in that Constructing an independent marking database, including: Capturing the marking operations of the user previewing the PDF file and extracting marking data, where the marking data includes page number, marking color, marking area position, marking timestamp, and marking type; Converting the marking data into structured data and storing it in the independent marking database, and the independent marking file library is grouped and managed according to the user ID, the PDF file ID, and the loading unit ID.
4. The loading optimization method for PDF instant preview according to claim 3, wherein, The combined loading process further includes: The combined loading process includes reading the marked pages and the loading unit ID from the independent marking database according to the user ID, and reading the marking layer data; Overlaying the PDF file corresponding page position in a vector manner according to the marking layer data, and outputting a combined PDF file for the user to preview.
5. The method for optimizing the loading of instant PDF preview according to claim 1, wherein Outputting a combined PDF file, including: Obtaining a set of marked pages and a set of unmarked pages; When outputting the combined PDF file, the loading priority of the set of marked pages is higher than that of the set of unmarked pages.
6. The loading optimization method for PDF instant preview according to claim 1, wherein Capturing the marking operations of the user previewing the PDF file further includes: Capturing the marking operations of the user previewing the PDF file and extracting the marked text fields; Identifying the text fields of the PDF file based on the marked text fields to obtain a set of the same text fields in the PDF file; If the user starts the same-field marking instruction, when the user marks any text field, synchronously mark the set of the same text fields.
7. The method for optimizing the loading of instant PDF preview according to claim 6, characterized in that, After extracting the marked text fields, it further includes: Constructing a field matching model based on text hash comparison; Invoking the field matching model to perform synonymous recognition on the marked text fields in the PDF file to generate a set of synonymous fields; If the user starts the same-field marking instruction, when the user marks any text field, synchronously mark the set of the same text fields and the set of synonymous fields.
8. The loading optimization method for PDF instant preview according to claim 1, wherein Storing the set of marked pages in the independent marking file library; Among them, the independent marking file library is a hierarchical storage structure, used to create an index directory according to the PDF file ID, and store according to the user ID under each directory.
9. The method for optimizing the loading of instant PDF preview according to claim 1, wherein, Capturing the marking operations of the user previewing the PDF file includes: Record the first marking operation of the user previewing the PDF file at the previous timestamp, and the first marking data corresponding to the first marking operation; Record the second marking operation of the user previewing the PDF file at the next timestamp, and the second marking data corresponding to the second marking operation; If the first marking data is repeated with the second marking data, use the second marking data to overwrite the first marking data for preview display.
10. The loading optimization method for PDF instant preview according to claim 9, characterized in that, Capture the marking operation of the user previewing the PDF file according to the timestamp ID, and store the version marking data according to the timestamp ID.
Citation Information
Patent Citations
Digital book interaction sharing system and realization method therefor
CN106599219A
Efficient browsing method and device applied to pdf document
CN115048339A
Internet-of-data reliable log record management, construction and assembly method and log traceability method
CN118689858A
System and method for mobile document preview
US20110173188A1