A method for repairing content truncation during the conversion of web pages to PDF

By obtaining web page DOM structure tree data and browser screenshot technology, analyzing the relationship between DOM elements, the problem of element truncation during web page conversion to PDF is solved, and the automated processing of static and dynamic web pages and complete PDF generation is realized.

CN119830865BActive Publication Date: 2025-07-01ANHUI HIGH QUALITY MINING TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510307667.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-01
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

In the process of converting web pages to PDF, there are problems of truncating text, tables, pictures or seals. Especially when the web page height is greater than 32,767 pixels, the image cannot be completely cut, resulting in truncating elements in PDF files.

Method used

By obtaining the web page DOM structure tree data, using browser screenshot technology to take full screenshots, combining the screenshots to form large images, analyzing the relationship of DOM elements, forming truncation processing rules, and truncating the picture within the A4 size, finally forming a complete PDF.

Benefits of technology

It realizes automated processing of static and dynamic web pages, supports custom element repair, and flexibly handles multiple tasks to ensure that elements in PDF files are not truncated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830865B_ABST
    Figure CN119830865B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for repairing content truncation during the conversion of a web page to a PDF. The method includes obtaining a basic data set of the DOM structure tree for the web page to be processed; using browser screenshot technology to perform full-screen screenshot processing on the web page to be processed, obtaining multiple web page screenshot images; merging all the web page screenshot images in order from top to bottom to form a large image, and determining the height of the large image based on the sum of the heights of all the web page screenshot images; performing DOM element definition analysis on the large image through the generated DOM structure tree data set to form a first truncation processing rule; truncating the large image within the A4 size according to the first truncation processing rule to form a plurality of segmented image data sets; performing PDF conversion on the generated image data sets to form a complete PDF. The present invention realizes the repair of the element truncation phenomenon during the conversion of a web page to a PDF based on the combination of html structured data and image processing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a method for repairing content truncation during the conversion of web pages to PDF. Background Art

[0002] The function of converting web pages to PDF is currently quite common. Generally, users will achieve PDF conversion through manual printing; since printing can only be processed manually, batch conversion of web pages to PDF can only be achieved through software. Whether it is manual processing or software processing, the converted PDF usually has the following problems:

[0003] 1. Software processing: When the web page is very long (generally with a height greater than 32767 pixels), it is impossible to capture a complete picture through front-end code or functions provided by the browser.

[0004] 2. Software processing: Since the standard PDF A4 size is basically inconsistent with the web page size, after converting the web page to PDF, there will be problems such as truncation of text, tables, seals, pictures, or other elements.

[0005] 3. Manual processing: Although print preview will automatically process text, pictures, tables, etc. to avoid truncation; however, for some custom signatures, seals, etc., the truncation problem still cannot be avoided.

[0006] 4. Traditional archiving PDF technologies on the market have very poor support for dynamic web pages (the page first loads the structure and then loads the data), and there are problems such as incomplete data loading and element truncation. Summary of the Invention

[0007] Aiming at the problems of truncation of text, tables, pictures, or seals that occur during the conversion of web pages to PDF in the above-mentioned existing technologies, the present invention provides a method for repairing content truncation during the conversion of web pages to PDF, effectively solving the problem of element truncation.

[0008] The present invention provides a method for repairing content truncation during the conversion of web pages to PDF, including the following steps:

[0009] Step 1, for the web page to be processed, obtain the basic data set of the web page DOM structure tree;

[0010] Step 2, for the web page to be processed, use browser screenshot technology to perform full-screen screenshot processing on the web page to be processed, obtaining multiple web page screenshot pictures;

[0011] Step 3, merge all the web page screenshot pictures in order from top to bottom to form a large picture, and the height of the large picture is determined based on the sum of the heights of all the web page screenshot pictures;

[0012] Step 4, perform width normalization on the large image in Step 3, and correct the width to the width of the standard A4 size;

[0013] Step 5, perform DOM element definition analysis on the large image with width normalization in Step 4 through the DOM structure tree data set generated in Step 1, analyze the relationship between the position of each row of pixels in the large image and the position of the DOM element in the web page, and form the first truncation processing rule;

[0014] Step 6, truncate the image generated in Step 4 within the A4 size according to the first truncation processing rule to form a data set of multiple segmented images;

[0015] Step 7, perform PDF conversion on the image data set generated in Step 6 to form a complete PDF.

[0016] In some embodiments, Step 1 includes:

[0017] For the web page to be processed, obtain the current element, and analyze whether the current element belongs to a preset truncation information type, where the preset truncation information types include images, tables, and seals;

[0018] If the current element belongs to a preset truncation information type, extract the type, coordinates, and size parameters of the element through a processing function.

[0019] In some embodiments, Step 2 includes:

[0020] Use the browser screenshot technology to take screenshots of the web page opened by the browser. If it is a multi-screen page, multiple images are generated;

[0021] When the height of the web page is greater than the maximum height of a single static screenshot by the browser screenshot technology, perform scrolling screenshot processing on the web page to generate an image array data.

[0022] In some embodiments, Step 3 includes:

[0023] Pre-create a large image with the original size, where the width of the large image with the original size is greater than the width of the web page, and the height of the large image with the original size is greater than the height of the web page in Step 1;

[0024] Import the screenshot images in Step 2 into the large image one by one.

[0025] In some embodiments, Step 4 includes:

[0026] Determine the coordinate data of the page blank range according to the data generated in Step 1;

[0027] Edge-cut the large image generated in step 3 according to the page blank range coordinate data, and then magnify the page of the large image so that its width is the same as the A4 width.

[0028] In some embodiments, step 5 includes:

[0029] Step 51, scan and analyze the pixels of the width-normalized large image generated in step 4 line by line in reverse order from the bottom up;

[0030] Step 52, for each row of pixels, analyze the relationship between the position of the current row of pixels and the position of the elements belonging to the preset truncation information type; if a part or all of the pixels in the current row of pixels are within the coordinate range of the element, or a part or all of the pixels in the current row of pixels are at the edge of the coordinate range of the element, then mark the current row as non-truncatable to form the first truncation processing rule data.

[0031] In some embodiments, step 5 further includes: for the width-normalized large image in step 4, determine the second truncation processing rule through the background color continuity determination rule;

[0032] Step 6 further includes: truncate the width-normalized large image generated in step 4 within the A4 size according to the second truncation processing rule.

[0033] In some embodiments, for the width-normalized large image in step 4, determining the second truncation processing rule through the background color continuity determination rule includes:

[0034] Step 53, classify the pixels of the current pixel row by color, and calculate the total number of color types a and the total number of pixels of each color type T;

[0035] Step 54, obtain the distance standard deviation d of the positions of the pixels of the same color type;

[0036] Step 55, if the total number of color types a is less than the first preset value, and the distance standard deviation of the positions of the pixels of the color with the largest color proportion is less than the second preset value, then determine that the current row is included in the non-truncation rule, denoted as the second truncation processing rule.

[0037] In some embodiments, in step 6, truncating the width-normalized large image generated in step 4 within the A4 size according to the first truncation processing rule and the second truncation processing rule includes:

[0038] Step 61, for the width-normalized large image generated in step 4, start counting from the first pixel row, and when reaching the pixel row with the A4 size height, determine whether the current pixel row belongs to the first truncation processing rule and the second truncation processing rule; if not, enter step 62, if so, enter step 63;

[0039] Step 62: If not, truncate at the current pixel row to form a picture with the height of A4 size, and use the next pixel row of the current pixel row as the first row to start counting, then execute Step 61;

[0040] Step 63: If it belongs, go back up to the pixel row that does not belong to the first truncation processing rule and the second truncation processing rule, denoted as pixel row L1, truncate at the position of pixel row L1, and use the next pixel row of pixel L1 as the first row to start counting, then execute Step 61;

[0041] Step 64: For the truncated picture, if the height is less than the height of A4 size, add a blank picture filled with blank elements at the bottom so that the height of the truncated picture after adding the blank picture reaches the height of A4 size.

[0042] In some embodiments, in Step 7, it includes: for all the pictures with A4 size generated in Step 6, convert them into single pages of PDF with A4 standard size one by one through a PDF conversion tool, and finally merge them into a whole PDF.

[0043] A method for repairing content truncation during web page to PDF conversion according to the present invention has the following beneficial effects:

[0044] 1. The types of HTML elements to be processed and repaired for truncation can be customized, and the method is flexible.

[0045] 2. Support the repair of element truncation data.

[0046] 3. Support not only static web pages but also dynamic web pages (data content is changing) for conversion and repair.

[0047] 4. This method is implemented by software and can automatically process multiple tasks simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a schematic flowchart of a method for repairing content truncation during web page to PDF conversion in an embodiment of the present application;

[0049] Figure 2 is a schematic flowchart of a method for truncating the pictures generated in Step 4 within the A4 size in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0051] Embodiment 1

[0052] See Figure 1 andFigure 2 , an embodiment of the present application provides a method for repairing content truncation during web page to PDF conversion, including the following steps:

[0053] Step 1: For the web page to be processed, obtain the basic data set of the web page DOM structure tree;

[0054] Specifically, this step 1 includes the following steps:

[0055] Step 11: For the web page to be processed, obtain the current element and analyze whether the current element belongs to a preset truncation information type. Specifically, the preset truncation information type can be determined customarily according to the scenario. For example, the preset truncation information type includes pictures, tables, seals, etc.;

[0056] Step 12: If the current element belongs to the preset truncation information type, extract the type, coordinates, and size parameters of the element through the processing function getHtmlCoordinate(e). It should be noted that the coordinate parameter of the element is the coordinate data relative to the top of the web page, and the obtained coordinate data is the horizontal and vertical coordinate values relative to the upper left corner of the web page.

[0057] In the embodiment of the present application, the processing function getHtmlCoordinate(e) is a custom function.

[0058] This function first obtains the specified element information by calling the system function document.getElementsByTagName(“F”) (where F can be element types such as table, img, div, etc.) (but not all elements are required, and in one implementation, element information such as table, img, div, etc. can be filtered), and returns the result element;

[0059] Then, it obtains the element rectangle size and coordinate information by calling the system function element.getBoundingClientRect function. For elements with borders, it also obtains the border data through the system function window.getComputedStyle(element).getPropertyValue("border-width"), and sequentially obtains the coordinate, length, width, and border size data of the element.

[0060] Step 2: For the web page to be processed, use the browser screenshot technology to perform full-screen screenshot processing on the web page to be processed, and obtain multiple web page screenshot pictures;

[0061] In this step 2, it specifically includes the following steps:

[0062] Step 21, use the browser screenshot technology to take screenshots of the web pages opened by the browser. If it is a multi-screen page, multiple pictures will be generated.

[0063] Step 22, when the height of the web page is greater than the maximum height of a single static screenshot by the browser screenshot technology, perform a scrolling screenshot process on the web page to generate picture array data.

[0064] In one implementation, the above Step 2 can be implemented based on the following method:

[0065] Step 23, use selenium webdriver to take screenshots of the web pages opened by the browser. If it is a multi-screen page, multiple pictures will be generated.

[0066] In one implementation, the above Step 2 can also be implemented based on the following method:

[0067] Step 24, when the height of the web page is greater than 32767 pixels, perform a scrolling screenshot process on the web page to generate picture array data.

[0068] Step 3, merge all the web page screenshot pictures in order from top to bottom to form a large picture. The height of the large picture is determined based on the sum of the heights of all the web page screenshot pictures;

[0069] In this Step 3, it specifically includes the following steps:

[0070] Step 31, pre-create a large picture with the original size. The width of the large picture with the original size is greater than the width of the web page. For example, it is 2 times the size of A4 at 72 dpi, that is: 1190 pixels; the height of the large picture with the original size is greater than the height of the web page in Step 1;

[0071] Step 32, import the screenshot pictures in Step 2 into the large picture one by one.

[0072] Step 4, perform width normalization processing on the large picture in Step 3 to correct the width to the width of the standard A4 size;

[0073] In this Step 4, it specifically includes:

[0074] Step 41, determine the coordinate data of the page blank range according to the data generated in Step 1;

[0075] Step 42, perform edge cutting on the large picture generated in Step 3 according to the coordinate data of the page blank range, and then perform magnification processing on the page of the large picture to make its width consistent with the A4 width, so that the data in the picture will be clearer.

[0076] Step 5: Analyze the DOM element definitions of the large width-normalized image in Step 4 using the DOM structure tree dataset generated in Step 1, analyze the relationship between the position of each row of pixels in the large image and the position of the DOM element in the web page, and form the first truncation processing rule.

[0077] In this Step 5, it specifically includes:

[0078] Step 51: Scan and analyze the pixels of the large width-normalized image generated in Step 4 row by row in reverse order from the bottom up.

[0079] Step 52: For each row of pixels, analyze the relationship between the position of the current row of pixels and the position of the element belonging to the preset truncation information type; if a part or all of the pixels in the current row of pixels are within the coordinate range of the element, or a part or all of the pixels in the current row of pixels are at the edge of the coordinate range of the element, then mark the current row as non-truncatable to form the first truncation processing rule data.

[0080] Specifically, for the large width-normalized image generated in Step 4, based on the coordinate positions of each DOM element in the web page, locate the position of the corresponding DOM element on the large width-normalized image. In the embodiments of the present application, considering one case, if the large width-normalized image generated in Step 4 is not scaled compared to the display content of the web page in Step 1, at least not scaled in the height direction, then on the premise of knowing the coordinate position P of a certain DOM element (such as an image, table, seal, etc.) on the web page, it is correspondingly possible to determine the position of the DOM element on the large width-normalized image at the same position P, and further determine whether it can be truncated at the same coordinate position P on the large width-normalized image. At the same time, considering another case, after processing the web page screenshot image based on the foregoing Steps 2, 3, and 4, the large width-normalized image generated in Step 4 may deviate in size and position from the display content of the web page in Step 1. In this case, it is necessary to determine the relatively accurate pixel row area of elements such as images, tables, and seals on the large width-normalized image. Therefore, in the embodiments of the present application, before the above Step 51, it includes: Step 50: Determine whether the display content at the same position of the large width-normalized image generated in Step 4 and the web page in Step 1 is the same. It can be understood that in the case where the display content at the same position of the two is the same, based on the coordinate positions of the DOM elements on the web page, the pixel rows occupied by the elements corresponding to the large image can be determined. This Step 50 includes:

[0081] Step 501: According to the DOM structure tree dataset generated in Step 1, determine the coordinate positions (the coordinate values are relative to the top of the web page) of each element of the preset truncation information type in the web page.

[0082] Step 502: Select at least three elements of preset truncation information types in the web page as elements to be referenced, and obtain the coordinate positions of the elements to be referenced in the web page. Preferably, the at least three elements of preset truncation information types adopt different truncation information types;

[0083] Step 503: Based on the coordinate positions of the elements to be referenced, determine the same coordinate positions (coordinate values relative to the top of the large image) of the width-normalized large image generated in Step 4 as the preliminary pixel regions of the elements to be referenced;

[0084] Step 504: Perform pixel matching based on the preliminary pixel regions of the elements to be referenced and the image of the coordinate position regions of the elements to be referenced in the web page, and determine whether the same coordinate positions in the web page and the width-normalized large image are of the same content based on the matching results: If the features of a preset number of pixels scattered in the preliminary pixel regions of the elements to be referenced are consistent with the features of a preset number of pixels in the image of the coordinate position regions of the elements to be referenced in the web page (the features include the position of the pixel relative to the upper left corner of the region image, the pixel grayscale value, and the distribution of the grayscale values of the neighboring pixels of the pixel), it is determined that the content at the same positions in the width-normalized large image generated in Step 4 and the web page is consistent;

[0085] Furthermore, execute Step 51 to perform a reverse scan and analysis of the pixels of the width-normalized large image generated in Step 4 line by line from the bottom up;

[0086] Furthermore, execute Step 52 to analyze the relationship between the position of the current row of pixels in the large image and the coordinate positions of the elements belonging to the preset truncation information types on the web page for each row of pixels. If a part or all of the pixels in the current row of pixels are within the coordinate range of the element, or a part or all of the pixels in the current row of pixels are at the edge of the coordinate range of the element, mark the current row as non-truncatable to form the first truncation processing rule data. For example, if the coordinate position of the element belonging to the preset truncation information type on the web page is [(200, 14600), (400, 14800)] (the upper left corner coordinate and the lower right corner coordinate of the element on the web page), then the element corresponds to occupying the pixel rows between the 14600th row and the 14800th row in the large image (including the 14600th row and the 14800th row). If the currently scanned pixel row is between the 14600th row and the 14800th row (including the 14600th row and the 14800th row), it is determined that the currently scanned pixel row is non-truncatable.

[0087] In the above step 504, in one implementation, the pixel matching is determined based on whether the pixel values of the pixels to be matched in the two regional images are the same and whether the gray-scale distribution characteristics of the neighboring pixels of the pixels to be matched in the two regional images are the same. The pixels to be matched are a preset number of pixels at random positions in the regional image. Further, in order to determine whether to further reduce or enlarge the width-normalized large image generated in step 4 later, another pixel matching method is provided. In another implementation, the pixel matching includes: using at least one pixel point in the preliminary pixel region of the element to be referenced as the pixel point to be matched, and matching a matching pixel point within the regional image of the coordinate position of the element to be referenced in the web page. The pixel characteristics of the matching pixel point and the pixel point to be matched (the pixel characteristics can be pixel gray-scale values) are the same, and the pixel characteristic distribution of the neighboring pixel points of the matching pixel point and the pixel characteristic (the pixel characteristics can be pixel gray-scale values) distribution of the neighboring pixel points of the pixel point to be matched are the same; determining the matching result based on the difference in the positions of the pixel point to be matched and the corresponding matching pixel point relative to the upper left corner of the regional image: if the positions of the pixel point to be matched and the corresponding matching pixel point relative to the upper left corner of the regional image are the same, it is determined that the width-normalized large image generated in step 4 and the content at the same position on the web page are the same; otherwise, based on the position difference, it is determined whether to further reduce or enlarge the width-normalized large image generated in step 4 to eliminate the difference (it should be noted that the above "the pixel characteristic distribution of the neighboring pixel points of the matching pixel point and the pixel characteristic (the pixel characteristics can be pixel gray-scale values) distribution of the neighboring pixel points of the pixel point to be matched are the same" indicates that compared with other pixel points in the regional image of the coordinate position of the element to be referenced in the web page, the pixel characteristic distribution of the neighboring pixel points of the matching pixel point in the regional image of the coordinate position of the element to be referenced in the web page and the pixel characteristic distribution of the neighboring pixel points of the pixel point to be matched have the greatest similarity).

[0088] The above step 504 further includes: if the characteristics of a preset number of pixels with scattered positions in the two regional images cannot be matched, it is determined that the content at the same position on the width-normalized large image generated in step 4 and the web page is inconsistent; then, based on the matching result, it is determined whether to further reduce or enlarge the width-normalized large image generated in step 4. For example, if the pixel characteristics and the gray-scale value distribution of the neighboring pixels of the pixels are the same in the two regional images but the positions are different, it is determined whether to further reduce or enlarge the width-normalized large image generated in step 4 according to the position difference.

[0089] Step 6: Truncate the width-normalized large image generated in step 4 within the A4 size according to the first truncation processing rule to form a plurality of segmented picture data sets;

[0090] In this step 6, it specifically includes:

[0091] Step 61: For the large width-normalized image generated in Step 4, start counting from the first pixel row. When reaching the pixel row with the height of A4 size, determine whether the current pixel row belongs to the first truncation processing rule. If it does not belong, go to Step 62; if it belongs, go to Step 63;

[0092] Step 62: If it does not belong, perform truncation at the current pixel row to form an image with the height of A4 size, and use the next pixel row of the current pixel row as the first row to start counting, and execute Step 61;

[0093] Step 63: If it belongs, go back up to the pixel row that does not belong to the first truncation processing rule, denoted as pixel row L1. Perform truncation at the position of pixel row L1, and use the next pixel row of pixel L1 as the first row to start counting, and execute Step 61;

[0094] Step 64: For the truncated image, if the height is less than the height of A4 size, add a blank image filled with blank elements at the bottom so that the height of the truncated image after adding the blank image reaches the height of A4 size.

[0095] Step 7: Perform PDF conversion on the image dataset generated in Step 6 to form a complete PDF.

[0096] In this Step 7, it includes: for all A4-sized images generated in Step 6, convert them one by one into A4-standard-sized PDF single pages through a PDF conversion tool, and finally merge them into a whole PDF.

[0097] Embodiment 2

[0098] In another implementation, considering that the data processed in Step 5 is usually quite accurate, but since the dataset generated in Step 1 is floating-point, there is a slight error between it and the pixel coordinate set (integer data) in Step 5 (there is an error in floating-point to integer conversion); therefore, a repair step needs to be added to further improve Step 5 of the above Embodiment 1 and further make an adaptive improvement to Step 6.

[0099] See Figure 1 and Figure 2 , this application embodiment provides a method for repairing content truncation during web page to PDF conversion, including the following steps:

[0100] Sequentially execute Steps 1 to 4 as described in Embodiment 1 above. The definitions of Steps 1 to 4 in this application embodiment are the same as those in Steps 1 to 4 described in Embodiment 1 above;

[0101] Step 5 includes: analyzing the DOM element definitions of the images in Step 4 using the DOM structure tree dataset generated in Step 1 to form a first truncation processing rule; the method for obtaining the first truncation processing rule includes:

[0102] Step 51: Analyze the pixels of the image generated in Step 4 by scanning line by line in reverse order from the bottom upwards.

[0103] Step 52: For each row of pixels, analyze the relationship between the position of the current row of pixels and the position of the elements belonging to the preset truncation information type; if a part or all of the pixels in the current row of pixels are within the coordinate range of the element, or a part or all of the pixels in the current row of pixels are at the edge of the coordinate range of the element, then mark the current row as non-truncatable to form the first truncation processing rule data.

[0104] In this Step 5, it also includes determining a second truncation processing rule for the image in Step 4 through the background color continuity determination rule. The method for obtaining this second truncation processing rule includes:

[0105] Step 53: Classify the pixels of the current pixel row by color, and calculate the total number of color types a and the total number of pixels of each color type T.

[0106] Step 54: Obtain the distance standard deviation d of the positions of the pixels of the same color type.

[0107] Step 55: If the total number of color types a is less than the first preset value (for example, this first preset value is 5), and the distance standard deviation of the positions of the pixels of the color with the largest color proportion is less than the second preset value (for example, this second preset value is 50), then determine that the current row is included in the non-truncation rule, denoted as the second truncation processing rule.

[0108] It should be noted that before classifying the pixels of the current pixel row by color in Step 53, it also includes retrieving the background color data for exclusion, and then classifying the pixels of the current pixel row by color.

[0109] Furthermore, in Step 6, it includes: truncating the image generated in Step 4 within the A4 size according to the first truncation processing rule and the second truncation processing rule. This Step 6 specifically includes:

[0110] Step 61: For the image generated in Step 4, start counting from the first pixel row. When reaching the pixel row at the height of the A4 size, determine whether the current pixel row belongs to the first truncation processing rule and the second truncation processing rule; if not, enter Step 62, if so, enter Step 63;

[0111] Step 62: If not, truncate at the current pixel row to form an image with the height of A4 size, and use the next pixel row below the current pixel row as the first row to start counting, then execute Step 61;

[0112] Step 63: If so, move backward to the pixel row that does not belong to the first truncation processing rule and the second truncation processing rule, denoted as pixel row L1, truncate at the position where pixel row L1 is located, and use the next pixel row below pixel L1 as the first row to start counting, then execute Step 61;

[0113] Step 64: For the truncated image, if the height is less than the height of A4 size, add a blank image filled with blank elements at the bottom so that the height of the truncated image after adding the blank image reaches the height of A4 size.

[0114] Furthermore, execute Step 7 to perform PDF conversion on the image dataset generated in Step 6 to form a complete PDF. The specific limitation of this Step 7 is the same as that of Step 7 in the above Embodiment 1.

[0115] The present invention is not limited to the above specific embodiments. Various transformations made by those of ordinary skill in the art starting from the above concepts without creative labor fall within the protection scope of the present invention.

Claims

1. A method for repairing content truncation during web page to PDF conversion, characterized in that: The steps include: Step 1: Obtain a basic data set of the web page DOM structure tree for the web page to be processed; Step 2: For the web page to be processed, a browser screenshot technology is used to perform full-screen screenshot processing on the web page to be processed, and multiple web page screenshot images are obtained; Step 3: merge all webpage screenshots in order from top to bottom to form a large image, and the height of the large image is determined based on the height and height of all webpage screenshots; Step 4, standardize the width of the large image in step 3 and correct the width to the width of the standard A4 size; Step 5, performing DOM element definition analysis on the width-normalized large image of step 4 through the DOM structure tree data set generated in step 1, analyzing the relationship between the position of each row of pixels in the large image and the position of the DOM element in the web page, and forming a first truncation processing rule, including: step 51, scanning and analyzing pixels of the width-normalized large image generated in step 4 row by row from the bottom to the top in reverse order; step 52, for each row of pixels, analyzing the relationship between the position of the current row of pixels and the position of the element belonging to the preset truncation information type; if a part or all of the pixels in the current row of pixels are within the coordinate range of the element, or a part or all of the pixels in the current row of pixels are at the edge of the coordinate range of the element, then marking the current row as non-truncation, forming the first truncation processing rule data; Step 6, truncating the image generated in step 4 within the A4 size according to the first truncation processing rule to form multiple segmented image data sets; Step 7: Convert the image dataset generated in step 6 into a PDF file to form a complete PDF file.

2. The method for repairing content truncation during webpage to PDF conversion according to claim 1, characterized in that: The step 1 comprises: For the web page to be processed, the current element is obtained, and whether the current element belongs to a preset truncation information type is analyzed, and the preset truncation information type includes a picture, a table, and a seal; If the current element belongs to the preset truncation information type, the type, coordinates, and size parameters of the element are extracted through a processing function.

3. The method for repairing content truncation during webpage to PDF conversion according to claim 1, characterized in that: The step 2 includes: Use browser screenshot technology to take screenshots of web pages opened by the browser. If it is a multi-screen page, multiple pictures will be generated; When the height of a web page is greater than the maximum height of a single static screenshot of the browser screenshot technology, a scrolling screenshot is performed on the web page to generate image array data.

4. The method for repairing content truncation during webpage to PDF conversion according to claim 1, characterized in that: The step 3 comprises: Pre-create a large image of original size, wherein the width of the large image of original size is greater than the width of the webpage, and the height of the large image of original size is greater than the height of the webpage in step 1; Import the screenshots in step 2 into the large image one by one.

5. The method for repairing content truncation during webpage to PDF conversion according to claim 1, characterized in that: The step 4 comprises: According to the data generated in step 1, determine the coordinate data of the blank range of the page; The large image generated in step 3 is edge-cut according to the coordinate data of the blank range of the page, and then the page of the large image is enlarged to make its width consistent with the width of A4.

6. The method for repairing content truncation during webpage to PDF conversion according to claim 1, characterized in that: The step 5 further includes: for the width-normalized large image in step 4, determining a second truncation processing rule by using a background color continuity determination rule; The step 6 further includes: truncating the width-standardized large image generated in step 4 within the A4 size according to a second truncation processing rule.

7. The method for repairing content truncation during webpage to PDF conversion according to claim 6, characterized in that: The second truncation processing rule is determined for the width-normalized large image in step 4 by using the background color continuity determination rule, including: Step 53, color classification of the pixels in the current pixel row, and calculation of the total number of color categories a and the total number of color pixels in each category T; Step 54, obtaining the distance standard deviation d of the positions of the same category color pixels; Step 55, if the total number of color categories a is less than the first preset value, and the distance standard deviation of the pixel position of the color with the largest color proportion is less than the second preset value, determine that the current row is included in the non-truncation rule, which is recorded as the second truncation processing rule.

8. The method for repairing content truncation during webpage to PDF conversion according to claim 6, characterized in that: In step 6, the width-standardized large image generated in step 4 is truncated within the A4 size according to the first truncation processing rule and the second truncation processing rule, including: Step 61, for the width-normalized large image generated in step 4, count from the first pixel row, and when reaching the pixel row of A4 size height, determine whether the current pixel row belongs to the first truncation processing rule or the second truncation processing rule; if not, proceed to step 62; if yes, proceed to step 63; Step 62, if it does not belong to, then cut off at the current pixel row to form an A4 size image, and take the next pixel row of the current pixel row as the first row, start counting, and execute step 61; Step 63, if yes, then go back to the pixel row that does not belong to the first truncation processing rule and the second truncation processing rule, record it as pixel row L1, truncate at the position of pixel row L1, and start counting with the next pixel row of pixel L1 as the first row, and execute step 61; Step 64, if the height of the truncated image is less than the height of the A4 size, a blank image filled with blank elements is added at the bottom so that the height of the truncated image after adding the blank image reaches the height of the A4 size.

9. The method for repairing content truncation during webpage to PDF conversion according to claim 1, characterized in that: The step 7 includes: for all the A4-sized pictures generated in step 6, converting them one by one into single PDF pages of A4 standard size through a PDF conversion tool, and finally merging them into a whole PDF.

Citation Information

Patent Citations

  • Method and device for generating PDF (Portable Document Format) file of webpage and electronic equipment

    CN118260501A

  • PDF text generation method and device and storage medium

    CN118276798A