Processing method and device for previewing PDF (Portable Document Format) file at front end, storage medium and electronic equipment
By splitting, tagging, and replacing front-end PDF files to generate hierarchical independent files, and performing progressive rendering based on the user's position, the problems of long request times and slow rendering in front-end PDF previews are solved, achieving efficient data requests and fast rendering.
Patent Information
- Application Number
- CN202511038487.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies suffer from long request times, slow rendering, and functional loss when previewing PDF files on the front end. This is especially true when the browser rendering engine directly opens the PDF file or when each page of the PDF is converted into an image for previewing, resulting in a poor user experience.
By preprocessing PDF files with preview functionality on the front end, including splitting, tagging, and replacing, multiple independent files with a hierarchical structure are generated for preview. Then, progressive rendering is performed based on the user's access position to acquire and render the target data step by step.
It improves the efficiency and accuracy of fragmented requests, reduces the amount of data requests, increases rendering speed, and retains the functionality of the original PDF file, thus enhancing the user experience.
Smart Images

Figure CN120929688A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, applicable to the fields of financial technology and / or healthcare, and particularly to a method and apparatus for processing PDF files for front-end preview, a storage medium, and an electronic device. Background Technology
[0002] In front-end business software, there is a frequent need to preview PDF files. For example, in smart healthcare scenarios, doctors can preview a PDF of a diagnosis they have written on the front end; in fintech scenarios, users can preview PDF contracts through the front-end software. Integrating PDF preview functionality into the front-end software allows users to quickly view PDF file content without leaving the current page or software interface, saving time spent opening dedicated PDF reader software and waiting for files to load, making the entire operation process smoother.
[0003] Currently, previewing PDF files uses a browser and its built-in rendering engine. However, this method requires the browser to directly open the PDF file, resulting in long request times and slow rendering. Another approach is to convert each page of the PDF into an image and preview it by loading the image. However, this method suffers from the loss of PDF's inherent capabilities, such as the loss of copyable text and hyperlinks. Furthermore, the request time remains significant even after conversion, leading to issues like long white screen times and fragmented loading for users. Summary of the Invention
[0004] In view of this, the present invention provides a method and apparatus for processing PDF files for front-end preview, a storage medium, and an electronic device, the main purpose of which is to solve the problems of long request time, slow rendering, and loss of function when previewing PDF files on the front end.
[0005] According to one aspect of the present invention, a method for processing PDF files for front-end preview is provided, comprising:
[0006] The PDF file with a preview function is preprocessed to obtain multiple independent files with a hierarchical structure for preview; the file preprocessing includes splitting, tagging and replacement.
[0007] Receive a user's preview request for a PDF file, and determine the target data from multiple independent files to be previewed based on the user's PDF access location;
[0008] Based on the hierarchical structure of the target data, progressive rendering is performed on the content related to the user's PDF access location to obtain a rendered PDF preview result.
[0009] Furthermore, the preprocessing of the PDF file with preview functionality at the front end yields multiple independent files to be previewed, each with a hierarchical structure, including:
[0010] A PDF file with a front-end preview function is split into multiple independent files; the number of independent files is the same as the number of pages in the PDF file.
[0011] Each independent file is tagged based on its height position in the PDF file to obtain tagged independent files;
[0012] By replacing high-frequency strings in each of the independent files with placeholders, multiple independent files with a hierarchical structure are obtained for preview.
[0013] Furthermore, the process of splitting PDF files with preview functionality on the front end includes:
[0014] Obtain the PDF file to be split from the front end, and create a folder corresponding to the PDF file to be split;
[0015] Obtain the total number of pages in the PDF file to be split, and split the PDF file according to the number of pages to obtain multiple independent files;
[0016] The split, multiple independent files are saved to the corresponding folders so that the target data can be determined from the folders based on the user's PDF access location.
[0017] Further, the tagging process based on the height position of each independent file in the PDF file to obtain tagged independent files includes:
[0018] Get the height of a single page in a PDF file;
[0019] The height position corresponding to each independent file is determined based on the page number of each independent file in the PDF file;
[0020] Each corresponding independent file is tagged based on its height position to obtain tagged independent files.
[0021] Furthermore, the step of replacing high-frequency strings in each of the independent files with placeholders at each level to obtain multiple independent files to be previewed with a hierarchical structure includes:
[0022] The independent file is subjected to a first high-frequency word extraction process using a word frequency analysis model to obtain a first high-frequency string set.
[0023] Set a first placeholder set corresponding to the first high-frequency string set, and use the first placeholder set to replace the first high-frequency string in the independent file to obtain a first replaced independent file;
[0024] Associate the first replaced independent file with the first set of placeholders to obtain the first replaced full file;
[0025] Based on the first high-frequency string set, a word frequency analysis model is used to perform a second high-frequency word extraction process on the independent file to obtain a second high-frequency string subset contained in the first high-frequency string set;
[0026] Set a second set of placeholders corresponding to the second high-frequency string subset, and use the second set of placeholders to replace the second high-frequency string in the first replaced full file;
[0027] When the length of the extracted high-frequency string is less than the set threshold, the high-frequency string extraction operation is stopped, and the hierarchical independent file to be previewed is obtained.
[0028] Furthermore, determining the target data from multiple independent files to be previewed based on the user-accessed PDF location includes:
[0029] The location of the PDF accessed by the user is determined based on the position of the preview page window on the PDF file when the user scrolls through the preview.
[0030] Based on the tags on each of the individual files to be previewed, target data matching the location of the PDF accessed by the user is determined from each of the individual files to be previewed corresponding to the PDF file to be previewed; the target data contains at least one target individual file to be previewed.
[0031] Furthermore, the progressive rendering process of content related to the user's PDF access location based on the hierarchical structure of the target data includes:
[0032] The placeholder set of the target data is obtained sequentially in the reverse order of high-frequency word extraction and then rendered.
[0033] When the first set of placeholders and the associated first replaced independent file are obtained, it is determined that the full information of the target data has been extracted, and the rendering of the target data is completed.
[0034] According to another aspect of the present invention, a processing apparatus for front-end previewing PDF files is provided, comprising:
[0035] The file preprocessing module is used to preprocess PDF files with preview functionality on the front end to obtain multiple independent files to be previewed with a hierarchical structure; the file preprocessing includes splitting, tagging, and replacement processing.
[0036] The target determination module is used to receive a user's preview request for a PDF file and determine target data from multiple independent files to be previewed based on the user's PDF access location;
[0037] The progressive rendering module is used to progressively render the content related to the user's PDF access location based on the hierarchical structure of the target data, and obtain the rendered PDF preview result.
[0038] Furthermore, the file preprocessing module includes:
[0039] The splitting unit is used to split a PDF file with a front-end preview function into multiple independent files; the number of independent files is the same as the number of pages in the PDF file.
[0040] The marking unit is used to mark each of the independent documents based on the height position of the independent document in the PDF file, so as to obtain the marked independent documents;
[0041] The replacement unit is used to replace high-frequency strings in each of the independent files by using placeholders to obtain multiple independent files to be previewed with a hierarchical structure.
[0042] Furthermore, the splitting unit is also used for:
[0043] Obtain the PDF file to be split from the front end, and create a folder corresponding to the PDF file to be split;
[0044] Obtain the total number of pages in the PDF file to be split, and split the PDF file according to the number of pages to obtain multiple independent files;
[0045] The split, multiple independent files are saved to the corresponding folders so that the target data can be determined from the folders based on the user's PDF access location.
[0046] Furthermore, the marking unit is also used for:
[0047] Get the height of a single page in a PDF file;
[0048] The height position corresponding to each independent file is determined based on the page number of each independent file in the PDF file;
[0049] Each corresponding independent file is tagged based on its height position to obtain tagged independent files.
[0050] Furthermore, the replacement unit is also used for:
[0051] The independent file is subjected to a first high-frequency word extraction process using a word frequency analysis model to obtain a first high-frequency string set.
[0052] Set a first placeholder set corresponding to the first high-frequency string set, and use the first placeholder set to replace the first high-frequency string in the independent file to obtain a first replaced independent file;
[0053] Associate the first replaced independent file with the first set of placeholders to obtain the first replaced full file;
[0054] Based on the first high-frequency string set, a word frequency analysis model is used to perform a second high-frequency word extraction process on the independent file to obtain a second high-frequency string subset contained in the first high-frequency string set;
[0055] Set a second set of placeholders corresponding to the second high-frequency string subset, and use the second set of placeholders to replace the second high-frequency string in the first replaced full file;
[0056] When the length of the extracted high-frequency string is less than the set threshold, the high-frequency string extraction operation is stopped, and the hierarchical independent file to be previewed is obtained.
[0057] Furthermore, the target determination module is also used for:
[0058] The location of the PDF accessed by the user is determined based on the position of the preview page window on the PDF file when the user scrolls through the preview.
[0059] Based on the tags on each of the individual files to be previewed, target data matching the location of the PDF accessed by the user is determined from each of the individual files to be previewed corresponding to the PDF file to be previewed; the target data contains at least one target individual file to be previewed.
[0060] Furthermore, the progressive rendering module is also used for:
[0061] The placeholder set of the target data is obtained sequentially in the reverse order of high-frequency word extraction and then rendered.
[0062] When the first set of placeholders and the associated first replaced independent file are obtained, it is determined that the full information of the target data has been extracted, and the rendering of the target data is completed.
[0063] According to another aspect of the present invention, a storage medium is provided, wherein at least one executable instruction is stored therein, the executable instruction causing a processor to perform operations corresponding to the above-described front-end preview PDF file processing method.
[0064] According to another aspect of the present invention, an electronic device is provided, including a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus;
[0065] The memory is used to store at least one executable instruction, which causes the processor to perform operations corresponding to the above-described front-end preview PDF file processing method.
[0066] By employing the above-described technical solutions, the technical solutions provided by the embodiments of the present invention have at least the following advantages:
[0067] This invention provides a method, apparatus, storage medium, and electronic device for processing PDF files for front-end preview. Compared with existing technologies, this invention preprocesses PDF files with front-end preview functionality to obtain multiple independent files to be previewed with a hierarchical structure, laying a data foundation for segmentation requests and improving the efficiency and accuracy of segmentation requests. This invention receives user preview requests for PDF files and determines target data from the multiple independent files to be previewed based on the user's PDF access location, achieving precise segmentation requests based on the user's preview location, reducing the amount of data requested in the current request, and improving the efficiency of data requests. Furthermore, this invention uses the hierarchical structure of the target data to perform progressive rendering processing on the content related to the user's PDF access location, improving rendering speed without changing the functionality of the original PDF file.
[0068] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0069] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0070] Figure 1 A flowchart illustrating a method for processing PDF files for front-end preview provided by an embodiment of the present invention is shown;
[0071] Figure 2 This diagram illustrates a file preprocessing flow provided in an embodiment of the present invention.
[0072] Figure 3 This diagram illustrates the process of determining target data according to an embodiment of the present invention.
[0073] Figure 4 This diagram illustrates the flow chart of the progressive rendering process provided in an embodiment of the present invention.
[0074] Figure 5 This diagram illustrates the structure of a front-end PDF file preview processing device provided in an embodiment of the present invention.
[0075] Figure 6 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention is shown. Detailed Implementation
[0076] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0077] This invention provides a method for previewing PDF files on a front-end, primarily executed in front-end business software to meet the previewing needs of front-end users for PDF format files. The front-end business software encompasses various business scenarios; for example, in a smart healthcare scenario, doctors can preview a diagnosis report as a PDF after writing it on the front-end; in a fintech scenario, users can preview contract documents in PDF format through the front-end software, etc. This invention does not impose specific limitations. Figure 1 As shown, the method includes:
[0078] 101. Perform file preprocessing on PDF files with front-end preview functionality to obtain multiple independent files with a hierarchical structure to be previewed;
[0079] In this embodiment of the invention, the current execution end obtains a PDF file with preview function on the front end. For example, in the smart medical scenario, doctors will issue various forms during the diagnosis and treatment process. Some forms need to be previewed and confirmed by the doctor before being uploaded or submitted to ensure the accuracy and completeness of the information. The following are some common forms that require doctors to preview and confirm: (1) Medical records. Medical records are the core records of patients' medical information. It is necessary to ensure that the text description is accurate, the logic is clear, and the information is complete. Doctors need to preview and confirm whether the contents of the medical records accurately reflect the patient's condition and diagnosis and treatment process to avoid subsequent diagnosis and treatment problems due to record errors or omissions. (2) Examination request forms (such as blood test, imaging examination (X-ray, CT, MRI, etc.), ultrasound examination, etc., including examination items, examination purpose, patient basic information, etc.). Examination request forms need to clarify whether the examination items are consistent with the patient's condition and diagnostic needs. Doctors need to preview and confirm whether the contents of the request form are accurate to avoid unnecessary examinations or omission of key examination items. (3) Prescription forms (including drug name, dosage, usage, course of treatment, etc.). Prescription forms are directly related to the patient's medication safety and treatment effect. Doctors need to preview documents to confirm whether the chosen medication is correct, the dosage is appropriate, and whether there are any interactions with other drugs, to ensure medication safety. In addition to the documents listed above that require doctor preview, this also includes surgical application forms, referral forms, discharge summaries, etc., which are not specifically limited in this embodiment. Furthermore, in some financial business systems, a document preview function before uploading is necessary to ensure that the document content meets business requirements. For example, in some financial data analysis platforms, users need to preview the content of uploaded documents to confirm that the data structure and format are correct before uploading; similarly, contract terms may be previewed before uploading, etc., which are not specifically limited in this embodiment.
[0080] In this embodiment of the invention, the current execution end also needs to perform file preprocessing on the acquired PDF file, including splitting, tagging, and replacement. Splitting involves dividing the PDF file to be previewed into multiple independent files, allowing subsequent data requests to be made based on these independent files. Tagging involves adding identifiers to the split independent files, enabling differentiation between them and facilitating subsequent fragment requests. Replacement involves replacing parts of the original PDF file with placeholders according to the length of high-frequency strings, resulting in multiple independent files to be previewed with a hierarchical structure. This facilitates progressive rendering, reduces rendering granularity, and further improves the user experience.
[0081] 102. Receive a user's preview request for a PDF file, and determine the target data from multiple independent files to be previewed based on the user's PDF access location;
[0082] In this embodiment of the invention, the current execution terminal receives a user's preview request for a PDF file. For example, when a user clicks the "Preview" button on a front-end business page, a preview request for the PDF file is sent based on this operation. This embodiment of the invention does not impose specific limitations. After receiving the user's preview request, the current execution terminal first opens a preview page window, then detects the position where the user stops scrolling through the preview page window, and determines this position as the user's current access position to the PDF file. Then, based on the user's PDF access position, target data is determined from multiple independent files to be previewed. The target data is not the complete data of the PDF file to be previewed, but rather the data corresponding to the PDF access position. For example, if the user is detected scrolling through the preview page window and stops at page 1 of the PDF file, the data on page 1 is determined as the target data from the independent files to be previewed; if the user is detected scrolling through the preview page window and stops at pages 3 and 4 of the PDF file, the data on pages 3 and 4 are determined as the target data from the independent files to be previewed, etc. This embodiment of the invention does not impose specific limitations.
[0083] 103. Based on the hierarchical structure of the target data, progressively render the content related to the user's PDF access location to obtain the rendered PDF preview result.
[0084] In this embodiment of the invention, the current execution end performs progressive rendering processing on the content related to the user's PDF access location based on the hierarchical structure of the target data, obtaining a rendered PDF preview result. The progressive rendering processing represents the process of rendering level by level according to the hierarchy of high-frequency strings extracted from the target data's hierarchical structure. Because each level involves less data and has lower data complexity during rendering, the rendering pressure on the browser is reduced, and the rendering speed is improved without losing the original PDF's functionality.
[0085] Furthermore, as a refinement and extension of the specific implementation methods described above, in order to lay a solid data foundation for fragmentation requests and improve the efficiency and accuracy of fragmentation requests, another method for processing PDF files for front-end preview is provided, such as... Figure 2 As shown, the steps involve preprocessing a PDF file with a preview function on the front end to obtain multiple independent files to be previewed with a hierarchical structure, including:
[0086] 201. Split the PDF file with preview function on the front end to obtain multiple independent files;
[0087] In this embodiment of the invention, the current execution end splits the PDF file with preview function on the front end. The specific splitting process is as follows:
[0088] First, obtain the PDF file to be split from the front end and create a folder corresponding to the PDF file to be split. For example, if the PDF file to be split is "XXX Discharge Summary", then create a folder named "XXX Discharge Summary" corresponding to "XXX Discharge Summary". The naming method of the folder can be set as needed, and this embodiment of the invention does not impose specific limitations.
[0089] Next, the total number of pages in the PDF file to be split is obtained, and the PDF file is split according to the number of pages to obtain multiple independent files. For example, if the PDF file "XXX Discharge Summary" has 40 pages, it is split into 40 independent files according to the number of pages, with the content of each independent file corresponding to the content of each page. Specifically, 40 new PDF documents can be created, and the content of each page can be copied and pasted into each of the 40 new PDF documents according to the page number, thus obtaining 40 independent files. This embodiment of the invention does not impose specific limitations. Since the split is performed by page number, the number of independent files after splitting is the same as the number of pages in the PDF file.
[0090] Finally, the multiple independent files are saved to the corresponding folders so that the target data can be determined from the folders based on the user's PDF access location. For example, the 40 independent files split from the above "XXX Discharge Summary" are saved to the corresponding "XXX Discharge Summary" folder. Each time a preview request for "XXX Discharge Summary" is received, the target data is determined from the "XXX Discharge Summary" folder. This embodiment of the invention does not impose specific limitations.
[0091] 202. Tag each of the independent files based on their height position in the PDF file to obtain the tagged independent files;
[0092] In this embodiment of the invention, the specific process of the current execution end performing the tagging process is as follows:
[0093] First, obtain the height of each page in the PDF file;
[0094] Next, the height position corresponding to each independent document is determined based on the page number of each independent document in the PDF file; for example, if the single page height of "XXX Discharge Summary" is 50, then the height position corresponding to the independent document on page 1 is [0, 50]; the height position corresponding to the independent document on page 2 is [50, 100]; the height position corresponding to the independent document on page 3 is [100, 150], and so on, the height position corresponding to the independent document on page 40 is [1950, 2000], etc., and the embodiments of the present invention do not make specific limitations.
[0095] Finally, each corresponding independent file is tagged based on the height position to obtain the tagged independent files. For example, for the 40 independent files obtained after splitting the "XXX Discharge Summary" above, the independent file on page 1 is tagged with height position [0,50]; the independent file on page 2 is tagged with height position [50,100]; the independent file on page 3 is tagged with height position [100,150]; and so on, the independent file on page 40 is tagged with height position [1950,2000], etc. This embodiment of the invention does not make specific limitations.
[0096] 203. Replace the high-frequency strings in each of the independent files with placeholders to obtain multiple independent files with a hierarchical structure to be previewed.
[0097] In this embodiment of the invention, the specific process of the current execution end performing progressive replacement processing is as follows:
[0098] S1. The independent file is subjected to a first high-frequency word extraction process using a word frequency analysis model to obtain a first high-frequency string set; for example, a first high-frequency word extraction process is performed on an independent file containing the string "aaabbbcccddqqaaabbbcccddxxaasssaaabbbcccddqqaaabbbcccddooppooppaaabbbccc ddqqaaabbbcccdd" to obtain a first high-frequency string set of "aaabbbcccddqqaaabbbcccdd", etc. This embodiment of the invention does not make specific limitations.
[0099] S2. Set a first placeholder set corresponding to the first high-frequency string set. For example, set the first placeholder set to json{1: aaabbbcccddqqaaabbbcccdd}. This embodiment of the invention does not make specific limitations. Then, use the first placeholder set to replace the first high-frequency string in the independent file to obtain the first replaced independent file. For example, use the first placeholder set json{1: aaabbbcccddqqaaabbbcccdd} to replace the string "aaabbbcccddqqaaabbbcccddxxaasssaaabbbcccddqqaaabbbcccddooppooppaaabbbcccddqqaaabbbcccdd" to obtain the string "{1}xxaasss{1}ooppoopp{1}" in the first replaced independent file. This embodiment of the invention does not make specific limitations.
[0100] S3. Associate the first replaced independent file with the first placeholder set to obtain the first replaced full file; For the first replaced independent file obtained in S2 above containing the string "{1}xxaasss{1}ooppoopp{1}", associate it with the first placeholder set json{1:aaabbbcccddqqaaabbbcccdd} to obtain the first replaced full file. This embodiment of the invention does not make specific limitations.
[0101] S4. Based on the first high-frequency string set, a word frequency analysis model is used to perform a second high-frequency word extraction process on the independent file to obtain a second high-frequency string subset contained in the first high-frequency string set; for example, the second high-frequency word extraction process is performed on the first high-frequency string set {“aaabbbcccddqqaaabbbcccdd”} extracted in step S1 to obtain the second high-frequency string subset contained in the high-frequency string set as {“aaabbbcccdd”}, etc. The embodiments of the present invention do not make specific limitations.
[0102] It should be noted that the length of the high-frequency string extracted in the second high-frequency word extraction process is shorter than that extracted in the first high-frequency word extraction process, and so on. After each high-frequency string extraction process, the length of the extracted high-frequency string will decrease.
[0103] S5. Set a second placeholder set corresponding to the second high-frequency string subset. For example, set the corresponding second placeholder set to json{2:aaabbbcccdd} for the second high-frequency string subset extracted in step S4. This embodiment of the invention does not make specific limitations. Then, use the second placeholder set json{2:aaabbbcccdd} to replace the second high-frequency string "aaabbbcccdd" in the first replaced full file.
[0104] S6. When the length of the extracted high-frequency string is less than the set threshold, the high-frequency string extraction operation is stopped, and the independent file to be previewed with a hierarchical structure is obtained. For example, if the set threshold is that the string length is greater than or equal to 10, then the high-frequency string extraction operation is no longer performed on the second high-frequency string "aaabbbcccdd", and the independent file to be previewed with two hierarchical structures is obtained. If the set threshold is that the string length is greater than or equal to 3, then a third high-frequency word extraction process can be performed on the second high-frequency string "aaa", "bbb", "ccc"}, and the third placeholder set is set to json{3:aaa, 4:bbb, 5:ccc}. Finally, the second high-frequency string "aaabbbcccdd" is replaced with the third placeholder set json{3:aaa, 4:bbb, 5:ccc}, and the independent file to be previewed with three hierarchical structures is obtained. This embodiment of the invention does not make specific limitations.
[0105] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, in order to make precise segmentation requests for the user's preview position, another method for processing front-end preview PDF files is provided, such as... Figure 3 As shown, the steps for determining target data from multiple independent files to be previewed based on the location of the PDF accessed by the user include:
[0106] 301. Determine the location of the PDF file accessed by the user based on the position of the preview page window on the PDF file when the user scrolls through the preview.
[0107] In this embodiment of the invention, the current execution terminal determines the PDF position accessed by the user based on the position where the preview page window hovers over the PDF file during the user's scroll preview operation. For example, for a PDF file with a single page height of 50, if it is detected that the user hovers over the first page of the PDF file during the scroll preview operation in the preview page window, then the PDF position accessed by the user is determined to be a height position of [0, 50]; if it is detected that the user hovers over the third or fourth page of the PDF file during the scroll preview operation in the preview page window, then the PDF position accessed by the user is determined to be a height position of [100, 150] and [150, 200], etc. This embodiment of the invention does not impose specific limitations.
[0108] 302. Based on the markers on each of the individual files to be previewed, determine the target data that matches the location of the PDF accessed by the user from each of the individual files to be previewed corresponding to the PDF file to be previewed;
[0109] In this embodiment of the invention, the current execution terminal determines the target data matching the location of the PDF accessed by the user from each of the independent files to be previewed, based on the markers on each of the independent files to be previewed corresponding to the PDF file to be previewed. For example, when the determined location of the PDF accessed by the user is a height position of [0, 50], the independent file to be previewed marked with [0, 50] is determined as the matching target data; when the determined location of the PDF accessed by the user is a height position of [100, 150] and [150, 200], the two independent files to be previewed marked with [100, 150] and [150, 200] are determined as the matching target data, respectively. Therefore, the target data contains at least one target independent file to be previewed. It should be noted that when the PDF preview has a zoom function, the number of target independent files to be previewed contained in the target data changes with the size of the file zoom.
[0110] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, in order to improve rendering speed without changing the functionality of the original PDF file, another method for front-end previewing PDF files is provided, such as... Figure 4 As shown, the steps involve progressively rendering content related to the user's PDF access location based on the hierarchical structure of the target data, including:
[0111] 401. Obtain the placeholder set of the target data in the reverse order of high-frequency word extraction and perform rendering operations;
[0112] In this embodiment of the invention, the current execution end sequentially obtains the placeholder set of the target data and performs rendering operations in the reverse order of high-frequency word extraction. For example, regarding the order of high-frequency word extraction in steps S1 to S6 above, the third placeholder set json{3: aaa, 4: bbb, 5: ccc} is obtained first and the first rendering operation is performed; then the second placeholder set json{2: aaabbbcccdd} is obtained and the second rendering operation is performed; then the first placeholder set json{1: aaabbbcccddqqaaabbbcccdd} is obtained and the third rendering operation is performed, etc. This embodiment of the invention does not make specific limitations.
[0113] 402. When the first set of placeholders and the associated first replaced independent file are obtained, it is determined that the full information of the target data has been extracted, and the rendering of the target data is completed.
[0114] In this embodiment of the invention, when the first placeholder set and the associated first replaced independent file are obtained at the current execution end, it is determined that the full information of the target data has been extracted, and the rendering of the target data is completed. That is, when the first placeholder set json{1:aaabbbcccddqqaaabbbcccdd} and the associated first replaced independent file “{1}xxaasss{1}ooppoopp{1}” are obtained, it is determined that the full information of the target data has been extracted. This embodiment of the invention does not make specific limitations. This embodiment of the invention performs progressive rendering based on the hierarchical structure of the target data. Each level has a small amount of data and low data complexity during rendering, thus reducing the rendering pressure on the browser and improving the rendering speed without losing the functionality of the original PDF.
[0115] This invention provides a method for processing PDF files for front-end preview. Compared with existing technologies, this invention preprocesses PDF files with front-end preview functionality to obtain multiple independent files to be previewed with a hierarchical structure, laying a data foundation for segmentation requests and improving the efficiency and accuracy of segmentation requests. This invention receives user preview requests for PDF files and determines target data from the multiple independent files to be previewed based on the user's PDF access location, achieving precise segmentation requests based on the user's preview location, reducing the amount of data requested in the current request, and improving the efficiency of data requests. Furthermore, this invention performs progressive rendering processing on the content related to the user's PDF access location through the hierarchical structure of the target data, improving rendering speed without changing the functionality of the original PDF file.
[0116] As a response to the above Figure 1 The implementation of the method shown in this embodiment of the invention provides a front-end PDF file preview processing device, such as... Figure 5 As shown, the device includes:
[0117] The file preprocessing module 51 is used to preprocess PDF files with preview functionality on the front end to obtain multiple independent files to be previewed with a hierarchical structure; the file preprocessing includes splitting, tagging and replacement processing.
[0118] The target determination module 52 is used to receive a user's preview request for a PDF file and determine target data from multiple independent files to be previewed based on the user's PDF access location.
[0119] The progressive rendering module 53 is used to progressively render the content related to the user's PDF access location based on the hierarchical structure of the target data, and obtain the rendered PDF preview result.
[0120] Furthermore, the file preprocessing module 51 includes:
[0121] The splitting unit is used to split a PDF file with a front-end preview function into multiple independent files; the number of independent files is the same as the number of pages in the PDF file.
[0122] The marking unit is used to mark each of the independent documents based on the height position of the independent document in the PDF file, so as to obtain the marked independent documents;
[0123] The replacement unit is used to replace high-frequency strings in each of the independent files by using placeholders to obtain multiple independent files to be previewed with a hierarchical structure.
[0124] Furthermore, the splitting unit is also used for:
[0125] Obtain the PDF file to be split from the front end, and create a folder corresponding to the PDF file to be split;
[0126] Obtain the total number of pages in the PDF file to be split, and split the PDF file according to the number of pages to obtain multiple independent files;
[0127] The split, multiple independent files are saved to the corresponding folders so that the target data can be determined from the folders based on the user's PDF access location.
[0128] Furthermore, the marking unit is also used for:
[0129] Get the height of a single page in a PDF file;
[0130] The height position corresponding to each independent file is determined based on the page number of each independent file in the PDF file;
[0131] Each corresponding independent file is tagged based on its height position to obtain tagged independent files.
[0132] Furthermore, the replacement unit is also used for:
[0133] The independent file is subjected to a first high-frequency word extraction process using a word frequency analysis model to obtain a first high-frequency string set.
[0134] Set a first placeholder set corresponding to the first high-frequency string set, and use the first placeholder set to replace the first high-frequency string in the independent file to obtain a first replaced independent file;
[0135] The first replaced full file is associated with the first placeholder set to obtain the first replaced full file;
[0136] Based on the first high-frequency string set, a word frequency analysis model is used to perform a second high-frequency word extraction process on the independent file to obtain a second high-frequency string subset contained in the first high-frequency string set;
[0137] Set a second set of placeholders corresponding to the second high-frequency string subset, and use the second set of placeholders to replace the second high-frequency string in the first replaced full file;
[0138] When the length of the extracted high-frequency string is less than the set threshold, the high-frequency string extraction operation is stopped, and the hierarchical independent file to be previewed is obtained.
[0139] Furthermore, the target determination module 52 is also used for:
[0140] The location of the PDF accessed by the user is determined based on the position of the preview page window on the PDF file when the user scrolls through the preview.
[0141] Based on the tags on each of the individual files to be previewed, target data matching the location of the PDF accessed by the user is determined from each of the individual files to be previewed corresponding to the PDF file to be previewed; the target data contains at least one target individual file to be previewed.
[0142] Furthermore, the progressive rendering module 53 is also used for:
[0143] The placeholder set of the target data is obtained sequentially in the reverse order of high-frequency word extraction and then rendered.
[0144] When the first set of placeholders and the associated first replaced independent file are obtained, it is determined that the full information of the target data has been extracted, and the rendering of the target data is completed.
[0145] This invention provides a processing device for front-end previewing PDF files. Compared with existing technologies, this invention preprocesses PDF files with front-end preview functionality to obtain multiple independent files with a hierarchical structure for preview, laying a data foundation for segmentation requests and improving the efficiency and accuracy of segmentation requests. This invention receives user preview requests for PDF files and determines target data from the multiple independent files based on the user's PDF access location, achieving precise segmentation requests based on the user's preview location, reducing the amount of data requested in the current request, and improving the efficiency of data requests. Furthermore, this invention performs progressive rendering processing on the content related to the user's PDF access location through the hierarchical structure of the target data, improving rendering speed without changing the functionality of the original PDF file.
[0146] According to one embodiment of the present invention, a storage medium is provided, the storage medium storing at least one executable instruction, the computer-executable instruction being able to execute the front-end preview PDF file processing method in any of the above method embodiments.
[0147] Figure 6 The diagram illustrates the structure of an electronic device according to an embodiment of the present invention. The specific embodiments of the present invention do not limit the specific implementation of the electronic device.
[0148] like Figure 6 As shown, the electronic device may include: a processor 602, a communications interface 604, a memory 606, and a communications bus 608.
[0149] The processor 602, communication interface 604, and memory 606 communicate with each other via communication bus 608.
[0150] Communication interface 604 is used to communicate with other network elements such as clients or other servers.
[0151] The processor 602 is used to execute program 610, which can specifically perform the relevant steps of the above-mentioned front-end preview PDF file processing method.
[0152] Specifically, program 610 may include program code that includes computer operation instructions.
[0153] Processor 602 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The electronic device may include one or more processors of the same type, such as one or more CPUs; or it may include processors of different types, such as one or more CPUs and one or more ASICs.
[0154] Memory 606 is used to store program 610. Memory 606 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0155] Specifically, program 610 can be used to cause processor 602 to perform the following operations:
[0156] The PDF file with a preview function is preprocessed to obtain multiple independent files with a hierarchical structure for preview; the file preprocessing includes splitting, tagging and replacement.
[0157] Receive a user's preview request for a PDF file, and determine the target data from multiple independent files to be previewed based on the user's PDF access location;
[0158] Based on the hierarchical structure of the target data, progressive rendering is performed on the content related to the user's PDF access location to obtain a rendered PDF preview result.
[0159] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0160] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for processing PDF files for front-end preview, characterized in that, include: Preprocess PDF files with preview functionality on the front end to obtain multiple independent files with a hierarchical structure for preview; The file preprocessing includes splitting, tagging, and replacement. Receive a user's preview request for a PDF file, and determine the target data from multiple independent files to be previewed based on the user's PDF access location; Based on the hierarchical structure of the target data, progressive rendering is performed on the content related to the user's PDF access location to obtain a rendered PDF preview result.
2. The method according to claim 1, characterized in that, The process of preprocessing PDF files with preview functionality on the front end results in multiple independent files with a hierarchical structure to be previewed, including: A PDF file with a front-end preview function is split into multiple independent files; the number of independent files is the same as the number of pages in the PDF file. Each independent file is tagged based on its height position in the PDF file to obtain tagged independent files; By replacing high-frequency strings in each of the independent files with placeholders, multiple independent files with a hierarchical structure are obtained for preview.
3. The method according to claim 2, characterized in that, The process of splitting PDF files with preview functionality on the front end includes: Obtain the PDF file to be split from the front end, and create a folder corresponding to the PDF file to be split; Obtain the total number of pages in the PDF file to be split, and split the PDF file according to the number of pages to obtain multiple independent files; The split, multiple independent files are saved to the corresponding folders so that the target data can be determined from the folders based on the user's PDF access location.
4. The method according to claim 2, characterized in that, The process of tagging each independent file based on its height position in the PDF file to obtain tagged independent files includes: Get the height of a single page in a PDF file; The height position corresponding to each independent file is determined based on the page number of each independent file in the PDF file; Each corresponding independent file is tagged based on its height position to obtain tagged independent files.
5. The method according to claim 2, characterized in that, The process involves replacing high-frequency strings in each of the independent files with placeholders at each level to obtain multiple independent files to be previewed with a hierarchical structure, including: The independent file is subjected to a first high-frequency word extraction process using a word frequency analysis model to obtain a first high-frequency string set. Set a first placeholder set corresponding to the first high-frequency string set, and use the first placeholder set to replace the first high-frequency string in the independent file to obtain a first replaced independent file; Associate the first replaced independent file with the first set of placeholders to obtain the first replaced full file; Based on the first high-frequency string set, a word frequency analysis model is used to perform a second high-frequency word extraction process on the independent file to obtain a second high-frequency string subset contained in the first high-frequency string set; Set a second set of placeholders corresponding to the second high-frequency string subset, and use the second set of placeholders to replace the second high-frequency string in the first replaced full file; When the length of the extracted high-frequency string is less than the set threshold, the high-frequency string extraction operation is stopped, and the hierarchical independent file to be previewed is obtained.
6. The method according to claim 1, characterized in that, The determination of target data from multiple independent files to be previewed based on the PDF location accessed by the user includes: The location of the PDF accessed by the user is determined based on the position of the preview page window on the PDF file when the user scrolls through the preview. Based on the tags on each of the individual files to be previewed, target data matching the location of the PDF accessed by the user is determined from each of the individual files to be previewed corresponding to the PDF file to be previewed; the target data contains at least one target individual file to be previewed.
7. The method according to any one of claims 1 to 6, characterized in that, The progressive rendering of content related to the user's PDF access location based on the hierarchical structure of the target data includes: The placeholder set of the target data is obtained sequentially in the reverse order of high-frequency word extraction and then rendered. When the first set of placeholders and the associated first replaced independent file are obtained, it is determined that the full information of the target data has been extracted, and the rendering of the target data is completed.
8. A processing device for front-end previewing PDF files, characterized in that, include: The file preprocessing module is used to preprocess PDF files with preview functionality on the front end to obtain multiple independent files with a hierarchical structure for preview. The file preprocessing includes splitting, tagging, and replacement. The target determination module is used to receive a user's preview request for a PDF file and determine target data from multiple independent files to be previewed based on the user's PDF access location; The progressive rendering module is used to progressively render the content related to the user's PDF access location based on the hierarchical structure of the target data, and obtain the rendered PDF preview result.
9. A storage medium storing at least one executable instruction that performs an operation corresponding to the front-end preview PDF file processing method as described in any one of claims 1-7.
10. An electronic device, comprising a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the front-end preview PDF file processing method as described in any one of claims 1-7.