PDF File Generation Method, Device, Equipment and Storage Medium Based on Website Language

By obtaining the hypertext markup language file of the page to be downloaded, converting it into a canvas file, identifying and adjusting the text size, the problem that traditional PDF generation methods do not consider the font size, and achieving a better user reading experience and browsing experience.

CN113627126BActive Publication Date: 2025-06-13SHENZHEN PING AN MEDICAL HEALTH TECHNOLOGY SERVICES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110910220.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-09
Publication Date
2025-06-13
Estimated Expiration
2041-08-09

AI Technical Summary

Technical Problem

Traditional PDF file generation methods do not take into account the font size in the web page, resulting in the generated PDF file that may affect user reading.

Method used

By detecting the download instructions for the page to be downloaded uploaded by the user, obtain the hypertext markup language file and convert it into a canvas file through the canvas plug-in. Then, collect the pixel color in the screenshot, use the trivalue method to identify the text size, set the scaling multiple, and sort the text through the PDF plug-in and generate a PDF file.

Benefits of technology

It realizes the adjustment of the font size in PDF files, improves the fluency of users' reading, and improves the experience of users' browsing by retaining the color of the page to be downloaded.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113627126B_ABST
    Figure CN113627126B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence technology, and provides a method, device, equipment and storage medium for generating a PDF file based on website language. The method includes: obtaining a hypertext markup language file of the page to be downloaded through the website language, converting the hypertext markup language file into a canvas file through a canvas plug-in, then scaling and sorting the text in the canvas file through a PDF plug-in to generate a PDF file, and sending the PDF file to the user. Thus, the adjustment of the font size in the converted PDF file is realized, making the user's reading of the PDF file smoother. In addition, the colors in the page to be downloaded can also be retained in the form of a canvas, further improving the user browsing experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly relates to a method, device, equipment and storage medium for generating PDF files based on website language. Background Art

[0002] PDF is a document in Portable Document Format, which is a file format developed by Adobe Systems for file exchange in a manner independent of applications, operating systems, and hardware. Due to its characteristics such as being non-editable, digital watermarks can be added, and the display effects are consistent under different platforms, it has been widely used.

[0003] Currently, websites and systems are all implemented based on code. Traditional generation methods can generate PDF files by implementing simple plain text or simple tables. However, the PDF files generated by this method do not consider the font size on the web page, resulting in problems that the generated PDF files may affect user reading. Summary of the Invention

[0004] The main object of the present invention is to provide a method, device, equipment and storage medium for generating PDF files based on website language, aiming to solve the problem that the PDF files generated by traditional generation methods do not consider the font size on the web page, resulting in problems that the generated PDF files may affect user reading.

[0005] The present invention provides a method for generating PDF files based on website language, including:

[0006] Detecting whether there is a download instruction for the page to be downloaded uploaded by the user;

[0007] Obtaining the HyperText Markup Language file of the page to be downloaded through website language based on the download instruction;

[0008] Converting the HyperText Markup Language file into a canvas file through a canvas plugin;

[0009] Obtaining a first screenshot of the page to be downloaded;

[0010] Collecting the values of the R color channel, G color channel, and B color channel in the RGB color model of the pixel points in the first screenshot to obtain the RGB color of each pixel point;

[0011] Setting the RGB color of each pixel point in the screenshot to (0, 0, 0), (255, 255, 255), or (P, P, P) according to a preset binarization method, where P is a preset value greater than 0 and less than 255, so as to obtain a temporary picture composed of three colors;

[0012] Calculate the areas occupied by the three colors in the temporary picture, and use a preset text segmentation method for the areas occupied by the two colors with smaller areas to obtain separated single characters;

[0013] Identify the sizes of the individual characters through the canvas;

[0014] Based on the sizes of the individual characters, set scaling factors for the individual characters according to the correspondence table between text size and magnification factor;

[0015] Perform scaling processing on the individual characters based on the scaling factors of the individual characters;

[0016] Sort the text in the canvas file through a PDF plugin to generate a PDF file;

[0017] Send the PDF file to the user.

[0018] Further, the step of converting the hypertext markup language file into a canvas file through a canvas plugin includes:

[0019] Provide a configuration window for the user through variable style options in the canvas plugin;

[0020] Obtain the style parameters set by the user in the configuration window;

[0021] Convert the hypertext markup language file into a canvas file according to the style parameters.

[0022] Further, after the step of converting the hypertext markup language file into a canvas file through a canvas plugin, it further includes:

[0023] Obtain the text size in the page to be downloaded according to the hypertext markup language file;

[0024] Determine whether the text size is less than a preset size value;

[0025] If it is less than the preset size value, magnify the content in the canvas file.

[0026] Further, the step of sorting the text in the canvas file through a PDF plugin to generate a PDF file includes:

[0027] Obtain a second screenshot of the page to be downloaded;

[0028] Analyze the paragraph format in the second screenshot;

[0029] Sort the text in the canvas file based on the paragraph format to generate the PDF file.

[0030] Further, the step of sorting the text in the canvas file through the PDF plug-in to generate a PDF file includes:

[0031] Sort the text in the canvas file through the PDF plug-in and generate a preview interface in real time;

[0032] Send the preview interface to the user;

[0033] Determine whether the confirmation instruction of the user is received;

[0034] If the confirmation instruction of the user is received, generate the PDF file based on the current sorting.

[0035] Further, the step of sorting the text in the canvas file through the PDF plug-in to generate a PDF file includes:

[0036] Preliminarily sort the text in the canvas file according to the preset template in the PDF plug-in to obtain a first temporary PDF file;

[0037] Determine whether the layout content on the last page in the first temporary PDF file is less than the preset ratio in the page;

[0038] If it is less than the preset ratio, reselect the template in the PDF plug-in for rearrangement until the obtained PDF file reaches the preset ratio.

[0039] Further, the step of sorting the text in the canvas file through the PDF plug-in to generate a PDF file includes:

[0040] Preliminarily sort the text in the canvas file according to the preset template in the PDF plug-in to obtain a second temporary PDF file;

[0041] Extract the reference color tone in the second temporary PDF file through a preset extraction algorithm; wherein, the preset extraction algorithm is any one of octree, median cut, K-means, fuzzy, C-means algorithms;

[0042] Set the theme color corresponding to the reference color tone by using a preset theme color setting model; wherein, the theme color setting model is used to characterize the correlation relationship between the reference color tone and the theme color of the image.

[0043] The present invention also provides a PDF file generation device based on website language, including:

[0044] A detection module, configured to detect whether there is a download instruction for the page to be downloaded uploaded by the user;

[0045] A first acquisition module, configured to obtain a HyperText Markup Language (HTML) file of the page to be downloaded based on the download instruction through the website language;

[0046] A conversion module, configured to convert the HTML file into a canvas file through a canvas plugin;

[0047] A second acquisition module, configured to obtain a first screenshot of the page to be downloaded;

[0048] An acquisition module, configured to acquire the values of the R color channel, the G color channel, and the B color channel in the RGB color model of the pixel points in the first screenshot, so as to obtain the RGB colors of each pixel point;

[0049] A setting module, configured to set the RGB color of each pixel point in the screenshot to (0, 0, 0), (255, 255, 255), or (P, P, P) according to a preset binarization method, where P is a preset value greater than 0 and less than 255, so as to obtain a temporary picture composed of three colors;

[0050] A calculation module, configured to calculate the areas occupied by the three colors in the temporary picture, and adopt a preset text segmentation method for the areas occupied by the two colors with smaller areas to obtain separated single characters;

[0051] An identification module, configured to identify the sizes of each of the single characters through the canvas;

[0052] A setting module, configured to set a scaling factor for each of the single characters based on the sizes of each of the single characters according to a correspondence table between the text size and the magnification factor;

[0053] A scaling module, configured to perform a scaling process on each of the single characters based on the scaling factors of each of the single characters;

[0054] A sorting module, configured to sort the text in the canvas file through a PDF plugin to generate a PDF file;

[0055] A sending module, configured to send the PDF file to the user.

[0056] The present invention further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the above are implemented.

[0057] The present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.

[0058] Advantages of the present invention: Obtain the HyperText Markup Language file of the page to be downloaded through the website language, convert the HyperText Markup Language file into a canvas file through a canvas plugin, then scale and sort the text in the canvas file through a PDF plugin to generate a PDF file, and send the PDF file to the user. Thus, the adjustment of the font size in the converted PDF file is realized, making it more fluent for the user to read the PDF file. In addition, the colors in the page to be downloaded can also be retained in the form of a canvas, further improving the user browsing experience. Brief Description of the Drawings

[0059] Figure 1 is a schematic flowchart of a method for generating a PDF file based on website language according to an embodiment of the present invention;

[0060] Figure 2 is a schematic block diagram of the structure of a device for generating a PDF file based on website language according to an embodiment of the present invention;

[0061] Figure 3 is a schematic block diagram of the structure of a computer device according to an embodiment of the present application.

[0062] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0064] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly. The connections described herein can be direct connections or indirect connections.

[0065] The term "and / or" in this article is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and B can represent: A exists alone, A and B exist simultaneously, and B exists alone.

[0066] In addition, in the present invention, descriptions such as "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement it. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0067] Referring to Figure 1 , the present invention provides a method for generating a PDF file based on website language, including:

[0068] S1: Detect whether there is a download instruction for the page to be downloaded uploaded by the user;

[0069] S2: Obtain the hypertext markup language file of the page to be downloaded through the website language based on the download instruction;

[0070] S3: Convert the hypertext markup language file into a canvas file through a canvas plugin;

[0071] S4: Obtain the first screenshot of the page to be downloaded;

[0072] S5: Collect the values of the R color channel, G color channel, and B color channel in the RGB color model of the pixel points in the first screenshot to obtain the RGB color of each pixel point;

[0073] S6: Set the RGB color of each pixel point in the screenshot to (0, 0, 0), (255, 255, 255), or (P, P, P) according to a preset binarization method, where P is a preset value greater than 0 and less than 255, so as to obtain a temporary picture composed of three colors;

[0074] S7: Calculate the areas occupied by the three colors in the temporary picture, and use a preset text segmentation method for the areas occupied by the two colors with smaller areas to obtain separated individual characters;

[0075] S8: Identify the size of each of the individual characters through the canvas;

[0076] S9: Based on the size of each of the individual characters, set a scaling factor for each of the individual characters according to the correspondence table between the text size and the magnification factor;

[0077] S10: Perform scaling processing on each of the individual characters based on the scaling factor of each of the individual characters;

[0078] S11: Sort the text in the canvas file through a PDF plugin to generate a PDF file;

[0079] S12: Send the PDF file to the user.

[0080] As described in step S1 above, detect whether there is a download instruction for the page to be downloaded uploaded by the user. Among them, the way of the download instruction uploaded by the user can be to click on the corresponding download link to trigger the download instruction, or directly input the corresponding download instruction. It should be understood that the download instruction includes the page information of the page to be downloaded, so that the system can determine the web page to be downloaded.

[0081] As described in step S2 above, based on the download instruction, obtain the Hypertext Markup Language file of the page to be downloaded through the website language (JavaScript). Among them, the website language is a scripting language running on the browser, abbreviated as js, that is, the Hypertext Markup Language file of the page to be downloaded can be obtained based on this. Among them, Hypertext Markup Language, abbreviated as html, is an application under the Standard Generalized Markup Language. Hypertext Markup Language is a markup language, including a "head" part and a "body" part. The "head" part provides information about the web page, and the "body" part provides the specific content of the web page. The Hypertext Markup Language file in this application mainly refers to the "body" part.

[0082] As described in step S3 above, the canvas plugin tool is a tool for converting the Hypertext Markup Language file into a canvas object. Since the canvas plugin (in this application, the canvas plugin generally refers to html2canvas) tool can convert the Hypertext Markup Language (Hypertext Markup Language) page into a canvas object, the implementation principle of the canvas plugin is to convert the Hypertext Markup Language page into a canvas file, and the canvas file is a canvas file. Based on the built-in methods of the canvas, the data required for generating pictures can be generated. Among them, this data can be understood as copying the text in the Hypertext Markup Language page in the canvas, and the size format of each text in the page to be downloaded can be retained. In some embodiments, the font format of the copied text can also be set in the canvas plugin. For example, if the font in the page to be downloaded is Song typeface, but it can be copied as regular script in the canvas file. In addition, the color of the font can also be set in the canvas plugin.

[0083] As described in steps S4 - S49 above, it is realized to identify the size of a single text through the binarization method and set the scaling factor of each single text.

[0084] In steps S4 - S7, a screenshot of the page to be downloaded is taken. The way of taking the screenshot is not limited. It can be the system - built screenshot software or a third - party program, such as QQ, WeChat, etc., to take a screenshot of the page to be downloaded, obtaining a second screenshot. Then, by using the binarization method, the RGB colors of the pixel points in the second screenshot are set to (0, 0, 0), (255, 255, 255), or (P, P, P), where P is a preset value greater than 0 and less than 255, thereby obtaining a temporary picture composed of three colors. Calculate the areas occupied by the three colors in the temporary picture, and respectively adopt a preset text segmentation method for the regions occupied by the two colors with smaller areas (since the area with the largest area must be the background, there is no need to analyze the region with the largest area), thereby obtaining separated single characters. Among them, the support vector machine is a type of generalized linear classifier that classifies data in a supervised learning manner for binary classification, and is suitable for comparing the characters to be recognized with the pre - stored characters to output the most similar characters.

[0085] In steps S8 - S9, the sizes of each of the single characters are recognized on the canvas. In the canvas, it is possible to measure the data of the sizes of each single character and compare them with the corresponding font size formats, thereby obtaining the sizes of each single character.

[0086] In step S10, a scaling factor is set for each of the single characters according to the correspondence table between the character size and the magnification factor. Among them, a standard value can be set, and the ratio of the standard value to the size of each single character is used as the scaling factor. Or it can be that only the single characters that do not meet the requirements are scaled, and the rest are not scaled (equivalent to a scaling factor of 1), that is, a range of single characters is set, and the single characters outside this range are scaled. In some embodiments, the setting of the scaling factor can also be carried out in other ways, and this application does not make any limitations in this regard. In addition, the text can be directly scaled in the canvas file based on the scaling factor and then extracted. Or it can be that after extracting each character from the canvas file, the characters are scaled before sorting and then sorted. Thus, the text in the converted PDF file is convenient for users to read, further improving the user's reading experience.

[0087] As described in the above step S11, the text in the canvas file is sorted through a PDF plugin to generate a PDF file. Only the front-end technology of the PDF plugin is used to sort the text in the canvas file to generate a PDF file. Among them, the sorting method is not limited. Some layout templates, page margins of the PDF page, paragraph formats, etc. can be customized in the PDF plugin for the convenience of customers. The PDF plugin is a client solution based on HyperText Markup Language 5 for generating PDF documents for various purposes. Thus, it is realized that without sending the canvas file to the backend, the content in the page to be downloaded can be directly converted into a PDF at the front end, improving the conversion speed and providing a better user experience. In addition, the colors in the page to be downloaded can be retained in the form of a canvas, further improving the user browsing experience.

[0088] As described in the above step S12, the PDF file is sent to the user, that is, the generated PDF file is sent to the corresponding user. The sending method is not limited and can be any existing sending method, such as direct sending, sending through wireless communication technology, etc.

[0089] In one embodiment, step S3 of converting the HyperText Markup Language file into a canvas file through a canvas plugin includes:

[0090] S301: Provide a configuration window for the user through the variable style options in the canvas plugin;

[0091] S302: Obtain the style parameters set by the user in the configuration window;

[0092] S303: Convert the HyperText Markup Language file into a canvas file according to the style parameters.

[0093] As described in the above steps S301 - S302, the user's design of the style is realized, enabling the user to choose their favorite font and color to generate a PDF, thus improving the user experience.

[0094] In step S301, that is, the variable style options in the canvas plugin are used to provide a configuration window for the user. The variable style options include font, font color, etc.

[0095] In step S302, the style parameters set in the user configuration window are obtained. Based on the information configured by the user, corresponding parameter items are set in the canvas plugin, and the rendering effect can also be displayed to the user in real time based on this parameter item, enabling the user to continuously adjust the style in the configuration window based on this and improving the user experience.

[0096] In step S303, then based on the style parameters, the hypertext markup language file is converted into a canvas file. The canvas file is a canvas file, and the data required for generating an image can be generated based on the built-in methods of the canvas. The data can be understood as copying the text in the hypertext markup language page in the canvas, and the size format of each piece of text in the page to be downloaded can be retained.

[0097] In one embodiment, after step S3 of converting the hypertext markup language file into a canvas file through the canvas plug-in, the following steps are further included:

[0098] S401: Obtain the text size in the page to be downloaded according to the hypertext markup language file;

[0099] S402: Determine whether the text size is less than a preset size value;

[0100] S403: If it is less than the preset size value, magnify the content in the canvas file.

[0101] As described in the above steps S401 - S403, the magnification process for text with a smaller font size is realized, which is convenient for users to view subsequently and improves the user's browsing experience.

[0102] In step S401, the text size in the page to be downloaded is obtained according to the hypertext markup language file. In some embodiments, the text information of each piece of text is recorded in the hypertext markup language file, that is, the corresponding text size can be directly extracted from the text information. In some embodiments, if the text information of each piece of text is not recorded in the hypertext markup language file, the page to be downloaded corresponding to the hypertext markup language file can be screenshot, and then based on the picture, it is recognized through a preset text size recognition method. The specific size recognition method will be described in detail later and will not be elaborated here.

[0103] In step S402, it is determined whether the text size is less than a preset size value. Among them, the preset size value is the text size corresponding to the text size that does not affect people's reading in advance. When the text is too small, it is considered that it will affect the user's reading and give the user a bad experience, so magnification processing is required.

[0104] In step S403, if it is less than the preset size value, the content in the canvas file is magnified. The magnification method is in the canvas file, and the content is magnified through the canvas. It should be noted that only some of the font sizes in the page to be downloaded may be less than the preset size value. Therefore, only the part of the font needs to be magnified.

[0105] In one embodiment, step S11 of sorting the text in the canvas file through the PDF plugin to generate a PDF file includes:

[0106] S1101: Obtain a second screenshot of the page to be downloaded;

[0107] S1102: Analyze the paragraph format in the second screenshot;

[0108] S1103: Sort the text in the canvas file based on the paragraph format to generate the PDF file.

[0109] As described in steps S1101 - S1103 above, the setting of the PDF file paragraphs is realized, making the generated PDF file as identical as possible to the display content on the page to be downloaded. Therefore, a second screenshot of the page to be downloaded can be obtained, and the paragraph format in the second screenshot can be recognized through the OCR text recognition method, so that the text in the canvas file can be sorted based on the paragraph format to generate the PDF file.

[0110] In one embodiment, step S11 of sorting the text in the canvas file through the PDF plugin to generate a PDF file includes:

[0111] S1111: Sort the text in the canvas file through the PDF plugin and generate a preview interface in real time;

[0112] S1112: Send the preview interface to the user;

[0113] S1113: Determine whether the confirmation instruction from the user is received;

[0114] S1114: If the confirmation instruction from the user is received, generate the PDF file based on the current sorting.

[0115] As described in the above steps S1111 - S1114, the preview of the PDF file by the user is realized, so that the generated PDF file is convenient for the user to read and can meet the different reading needs of different users. In practice, different users have different reading needs. Therefore, before sending it to the customer, a preview interface needs to be generated for the user to view. That is, during the sorting process, real - time screenshots are taken. The time for taking the screenshot is not limited, as long as the sorting situation of the PDF file can be seen. For example, a screenshot is taken 5 seconds after the sorting instruction. At this time, part or all of the sorting has been completed. Therefore, this preview interface can be sent to the user. The user can determine whether it is suitable for their own reading based on this. If satisfied, a confirmation instruction can be input. Therefore, after receiving the confirmation instruction input by the user, a PDF file can be generated based on the current sorting. This enables the user to select their favorite file format based on the preview situation, improving the user experience.

[0116] In one embodiment, step S11 of generating a PDF file by sorting the text in the canvas file through a PDF plug - in includes:

[0117] S1121: Initially sort the text in the canvas file according to a preset template in the PDF plug - in to obtain a first temporary PDF file;

[0118] S1122: Determine whether the layout content on the last page of the first temporary PDF file is less than a preset ratio in the page;

[0119] S1123: If it is less than the preset ratio, re - select a template in the PDF plug - in for rearrangement until the obtained PDF file reaches the preset ratio.

[0120] As described in the above steps S1121 - S1123, the optimization and adjustment of the layout in the PDF file are achieved. Generally speaking, if there is only very little content on the last page, when the user reads the content on the last page, it is very difficult to connect it with the previous content, resulting in a poor user experience. Therefore, based on the initially set template, which includes the layout in the PDF page, such as the page margins in a PDF page. If only 10 preset - sized characters can be accommodated in one line of the PDF page in the PDF file, it will switch to the next line. Since the number of characters in the file to be downloaded is uncertain, it is impossible to determine the proportion of the sorted layout content on the last page that occupies one PDF page. Therefore, it can be detected whether the sorted layout content on the last page in the first temporary PDF file is less than the preset proportion in the page. When it is less than the preset proportion, a template in the PDF plugin is re - selected for rearrangement until the obtained PDF file reaches the preset proportion. The selection method of the re - selected PDF plugin is not limited, and preferably a template similar to the preset template is selected.

[0121] In one embodiment, step S11 of generating a PDF file by sorting the text in the canvas file through a PDF plugin includes:

[0122] S1131: Initially sort the text in the canvas file according to the preset template in the PDF plugin to obtain a second temporary PDF file;

[0123] S1132: Extract the reference color tone in the second temporary PDF file through a preset extraction algorithm; where the preset extraction algorithm is any one of the octree, median cut, K - means, fuzzy, C - means algorithms;

[0124] S1133: Set the theme color corresponding to the reference color tone by using a preset theme - color setting model; where the theme - color setting model is used to represent the correlation relationship between the reference color tone and the theme color of the image in terms of color characteristics.

[0125] As described in the above steps S1131 - S1133, the setting of the theme color in the PDF is achieved, making the user's reading experience better. The reference color tone suitable as the theme color of the PDF is determined through a theme - color setting model that represents the correlation relationship between the reference color tone and the theme color of the image at least in terms of color characteristics. The theme - color setting model can adopt any one of a linear fitting model, a non - linear fitting model, a regression model, and a deep - learning algorithm model. The training method of the model will be described in detail later and will not be elaborated here.

[0126] It should be noted that the color feature is a general term for various color-related features, which may include various specific color features such as color distribution, color difference, color saliency, etc. Since the theme color is also a color extracted from the PDF, the theme color setting model should at least be able to characterize the correlation between the reference color tone and the theme color in terms of color features, so as to determine which reference color tone is closer to the theme color and most suitable as the theme color. Specifically, when setting the theme color of the PDF using the theme color setting model, the actual scores of each reference color tone can be output by the theme color setting model respectively. The magnitude of the actual score represents the degree of approximation between the reference color tone and the theme color. Then, the reference color tone with the highest score among the actual scores is determined as the theme color. Of course, other methods such as grading can also be used to represent the degree of approximation between the reference color tone and the theme color.

[0127] In one embodiment, the training method of the theme color setting model includes:

[0128] S4421: respectively obtain the sample reference color tone and the sample theme color of the sample image;

[0129] S4422: obtain the Euclidean distance between the sample reference color tone and the sample theme color of the same sample image, and convert it to an approximation score according to the Euclidean distance;

[0130] S4423: extract the color distribution parameter, color difference parameter, and color saliency parameter from the sample reference color tone and the sample theme color of the same sample image;

[0131] S4424: based on the differences in the color distribution parameter, color difference parameter, and color saliency parameter, fit the corresponding approximation score to obtain a theme color setting model for characterizing the correlation between the reference color tone and the theme color of the image in terms of color features.

[0132] As described in step S4421 above, from each sample image, the sample reference color tone and the sample theme color of the sample image are respectively obtained. Among them, the sample reference color tone can be directly extracted through algorithms such as octree and K-means. The sample theme color is a certain color given by professional designers based on their relevant knowledge of theme color determination combined with long-term user experience. Further, the accuracy of the determined sample theme color can be improved by increasing the number of designers and determining the theme color in different application scenarios respectively.

[0133] As described in step S4422 above, this step aims to obtain the Euclidean distance between the sample reference hue and the sample theme color of each sample image by the above-mentioned execution entity, and convert the Euclidean distance to obtain the approximation score between each sample reference hue and the unique sample theme color. Among them, the Euclidean distance is a description method of the vector difference in the vector description space. Since both the sample reference hue and the sample theme color are colors, a suitable color space can be selected to describe the sample reference hue and the sample theme color in vector form respectively. Specifically, the Euclidean distance can be composed of multiple sub-vectors, and each sub-vector represents a certain color feature of the sample reference hue and the sample theme color in the color space, such as hue, saturation, brightness, etc.

[0134] As described in step S4423 above, this step aims to extract the respective color distribution parameters, color difference parameters, and color saliency parameters from the sample reference hue and the sample theme color of each sample image by the above-mentioned execution entity. For example, the color distribution parameters, color difference parameters, and color saliency parameters extracted from the sample reference hue are respectively named the first color distribution parameter, the first color difference parameter, and the first color saliency parameter, while the parameters extracted from the sample quantization are named the second color distribution parameter, the second color difference parameter, and the second color saliency parameter respectively. Among them, the color distribution parameter is used to describe the area ratio of each reference hue in the recolored image after recoloring the corresponding pixel positions in the original image with the reference hue; the color difference parameter is used to describe the color difference between the recolored image and the original image after recoloring the corresponding pixel positions in the original image with the reference hue; the color saliency parameter is used to describe the different degrees of attraction of different colors to the user's visual corner points.

[0135] As described in step S4421 above, this step aims to fit the approximation score between the sample theme color and the sample reference hue of the corresponding sample image based on the differences in the above three specific color features, in order to find the general reason for the approximation score through fitting. For example, the approximation score between the sample reference hue score A and the sample theme color B of sample image X is 95 points (out of 100), but the approximation score between the sample reference hue C and the sample theme color B of sample image X is only 50 points. Through the above fitting, it is found that the lower approximation score is mainly because the difference in the color distribution parameters between the sample reference hue A and the sample theme color B is small, and the two colors have similar color saliency parameters.

[0136] Advantages of the present invention: Obtain the Hypertext Markup Language file of the page to be downloaded through the website language, convert the Hypertext Markup Language file into a canvas file through a canvas plugin, then scale and sort the text in the canvas file through a PDF plugin to generate a PDF file, and send the PDF file to the user. Thus, the adjustment of the font size in the converted PDF file is realized, making it more fluent for the user to read the PDF file. In addition, the colors in the page to be downloaded can also be retained in the form of a canvas, further improving the user browsing experience.

[0137] Referring to Figure 2 , the present invention also provides a PDF file generation device based on website language, including:

[0138] A detection module 10, configured to detect whether there is a download instruction for the page to be downloaded uploaded by the user;

[0139] A first acquisition module 20, configured to obtain the Hypertext Markup Language file of the page to be downloaded through the website language based on the download instruction;

[0140] A conversion module 30, configured to convert the Hypertext Markup Language file into a canvas file through a canvas plugin;

[0141] A second acquisition module 40, configured to obtain a first screenshot of the page to be downloaded;

[0142] An acquisition module 50, configured to acquire the values of the R color channel, the G color channel, and the B color channel in the RGB color model of the pixel points in the first screenshot to obtain the RGB colors of each pixel point;

[0143] A setting module 60, configured to set the RGB color of each pixel point in the screenshot to (0, 0, 0), (255, 255, 255), or (P, P, P) according to a preset binarization method, where P is a preset value greater than 0 and less than 255, so as to obtain a temporary picture composed of three colors;

[0144] A calculation module 70, configured to calculate the areas occupied by the three colors in the temporary picture, and adopt a preset text segmentation method for the areas occupied by the two colors with smaller areas to obtain separated single characters;

[0145] An identification module 80, configured to identify the sizes of each of the single characters through the canvas;

[0146] A setting module 90, configured to set a scaling factor for each of the single characters based on the sizes of each of the single characters according to a correspondence table between text size and magnification factor;

[0147] A scaling module 100 for scaling each individual character based on the scaling factor of each individual character;

[0148] A sorting module 110 for sorting the text in the canvas file through a PDF plugin to generate a PDF file;

[0149] A sending module 120 for sending the PDF file to the user.

[0150] Refer to Figure 3 , in an embodiment of the present application, a computer device is further provided. This computer device can be a server, and its internal structure can be as Figure 3 shown. This computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of this computer design is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. This memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of this computer device is used to store various PDF files, etc. The network interface of this computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it can implement the method for generating a PDF file based on the website language described in any of the above embodiments.

[0151] Those skilled in the art can understand that Figure 3 the structure shown in

[0152] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. In an embodiment of the present application, relevant data can be acquired and processed based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, technologies, and application systems.

[0153] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0154] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for generating a PDF file based on website language described in any of the above embodiments can be implemented.

[0155] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in the present application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, there are various forms of RAM, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0156] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, apparatus, article, or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, apparatus, article, or method. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, apparatus, article, or method including that element.

[0157] Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.

[0158] The underlying blockchain platform may include processing modules such as user management, basic services, smart contracts, and operation monitoring. Among them, the user management module is responsible for the identity information management of all blockchain participants, including maintaining the generation of public and private keys (account management), key management, and the maintenance of the correspondence between the real identity of users and blockchain addresses (permission management). And under authorization, it supervises and audits the transaction situations of certain real identities, and provides the rule configuration for risk control (risk control audit); the basic service module is deployed on all blockchain node devices, used to verify the validity of business requests, and record them on the storage after consensus on valid requests. For a new business request, the basic service first performs interface adaptation parsing and authentication processing (interface adaptation), then encrypts the business information through a consensus algorithm (consensus management), transmits it to the shared ledger completely and consistently after encryption (network communication), and performs record storage; the smart contract module is responsible for the registration and issuance of contracts, as well as contract triggering and contract execution. Developers can define contract logic through a certain programming language, publish it to the blockchain (contract registration), trigger the execution by calling keys or other events according to the logic of the contract terms, complete the contract logic, and at the same time provide functions for contract upgrade and cancellation; the operation monitoring module is mainly responsible for the deployment, configuration modification, contract setting, cloud adaptation during the product release process, and the visual output of the real-time status during product operation, such as: alarm, monitoring network conditions, monitoring the health status of node devices, etc.

[0159] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A method for generating a PDF file based on website language, characterized in that, it includes: Detect whether there is a download instruction for the page to be downloaded uploaded by the user; Obtain the hypertext markup language file of the page to be downloaded through the website language based on the download instruction; Convert the hypertext markup language file into a canvas file through a canvas plugin; Obtain the first screenshot of the page to be downloaded; Collect the numerical values of the R color channel, G color channel, and B color channel in the RGB color model of the pixel points in the first screenshot to obtain the RGB colors of each pixel point; According to the preset binarization method, set the RGB color of each pixel point in the screenshot to (0, 0, 0), (255, 255, 255), or (P, P, P), where P is a preset value greater than 0 and less than 255, so as to obtain a temporary picture composed of three colors; Calculate the areas occupied by the three colors in the temporary picture, and use a preset text segmentation method for the areas occupied by the two colors with smaller areas to obtain separated individual characters; Identify the sizes of each of the individual characters through the canvas; Based on the sizes of each of the individual characters, set a scaling factor for each of the individual characters according to the correspondence table between the text size and the magnification factor; Perform scaling processing on each of the individual characters based on the scaling factors of each of the individual characters; Sort the text in the canvas file through a PDF plugin to generate a PDF file; Send the PDF file to the user; The step of converting the hypertext markup language file into a canvas file through the canvas plugin includes: Provide a configuration window for the user through the variable style options in the canvas plugin; Obtain the style parameters set by the user in the configuration window; Convert the hypertext markup language file into a canvas file according to the style parameters; The step of sorting the text in the canvas file through a PDF plugin to generate a PDF file includes: Obtain the second screenshot of the page to be downloaded; Analyze the paragraph format in the second screenshot; Sort the text in the canvas file based on the paragraph format to generate the PDF file.

2. The method for generating a PDF file based on website language according to claim 1, characterized in that, after the step of converting the hypertext markup language file into a canvas file through the canvas plugin, it further includes: Obtain the text size in the page to be downloaded according to the hypertext markup language file; Judge whether the text size is less than a preset size value; If it is less than the preset size value, magnify the content in the canvas file.

3. The method for generating a PDF file based on website language according to claim 1, characterized in that, The step of sorting the text in the canvas file through a PDF plugin to generate a PDF file includes: Sort the text in the canvas file through a PDF plugin and generate a preview interface in real time; Send the preview interface to the user; Judge whether the confirmation instruction of the user is received; If the confirmation instruction of the user is received, the PDF file is generated based on the current sorting.

4. The method for generating a PDF file based on website language according to claim 1, wherein, the step of sorting the text in the canvas file through a PDF plugin to generate a PDF file includes: preliminarily sorting the text in the canvas file according to a preset template in the PDF plugin to obtain a first temporary PDF file; judging whether the layout content on the last page in the first temporary PDF file is less than a preset ratio in the page; if it is less than the preset ratio, reselect a template in the PDF plugin for rearrangement until the obtained PDF file reaches the preset ratio.

5. The method for generating a PDF file based on website language according to claim 1, wherein, the step of sorting the text in the canvas file through a PDF plugin to generate a PDF file includes: preliminarily sorting the text in the canvas file according to a preset template in the PDF plugin to obtain a second temporary PDF file; extracting a reference color tone in the second temporary PDF file through a preset extraction algorithm; wherein, the preset extraction algorithm is any one of octree, median cut, K-means, fuzzy, C-means algorithms; setting a theme color corresponding to the reference color tone by using a preset theme color setting model; wherein, the theme color setting model is used to represent the correlation between the reference color tone and the theme color of an image in terms of color characteristics.

6. A device for generating a PDF file based on website language, wherein, it includes: a detection module, configured to detect whether there is a download instruction for a page to be downloaded uploaded by a user; a first acquisition module, configured to obtain a hypertext markup language file of the page to be downloaded through website language based on the download instruction; a conversion module, configured to convert the hypertext markup language file into a canvas file through a canvas plugin; a second acquisition module, configured to obtain a first screenshot of the page to be downloaded; a collection module, configured to collect the values of the R color channel, the G color channel, and the B color channel in the RGB color model of the pixel points in the first screenshot to obtain the RGB color of each pixel point; a setting module, configured to set the RGB color of each pixel point in the screenshot to (0, 0, 0), (255, 255, 255), or (P, P, P) according to a preset binarization method, where P is a preset value greater than 0 and less than 255, so as to obtain a temporary picture composed of three colors; a calculation module, configured to calculate the areas occupied by the three colors in the temporary picture, and adopt a preset text segmentation method for the areas occupied by the two colors with smaller areas to obtain separated single characters; an identification module, configured to identify the size of each of the single characters through a canvas; a setting module, configured to set a scaling multiple for each of the single characters based on the size of each of the single characters according to a correspondence table between text size and magnification factor. A scaling module, configured to scale each of the individual characters based on the scaling factor of each of the individual characters; A sorting module, configured to sort the characters in the canvas file through a PDF plugin to generate a PDF file; A sending module, configured to send the PDF file to the user; The step of converting the hypertext markup language file into a canvas file through a canvas plugin includes: Providing a configuration window for the user through variable style options in the canvas plugin; Obtaining the style parameters set by the user in the configuration window; Converting the hypertext markup language file into a canvas file according to the style parameters; The step of sorting the characters in the canvas file through a PDF plugin to generate a PDF file includes: Obtaining a second screenshot of the page to be downloaded; Analyzing the paragraph format in the second screenshot; Sorting the characters in the canvas file based on the paragraph format to generate the PDF file.

7. A computer device, including a memory and a processor, where the memory stores a computer program, characterized in that when the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Method and device for enhancing image display

    WO2018072270A1

  • Text region obtaining method and apparatus, storage medium and terminal device

    WO2020107866A1