Markdown source code paging method and device, electronic equipment and storage medium

By converting the Markdown source code into HTML source code and color marking, and paging it with the color information of PDF files, the problems of inaccurate page paging and layout change are solved, and the pagination effect with high accuracy and layout maintenance is achieved.

CN120493870APending Publication Date: 2025-08-15YANGTZE RIVER DELTA (ANHUI) KEXUN SMART PARK OPERATION CENTER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510304045.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing Markdown source code paging technology has problems such as paging inaccuracy and layout information changes, especially based on text matching methods and pre-rendering solutions, which cannot effectively solve the problem of spreading elements and layout maintenance.

Method used

By converting the original Markdown source code into HTML source code and color marking its elements, generating mark HTML source code and PDF files, using color information of mark PDF files and HTML source code to determine the single page Markdown source code to ensure that the layout information remains unchanged.

Benefits of technology

Improve the accuracy of Markdown source code paging, keep the layout information of the original source code unchanged, and reduce the occurrence of paging errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493870A_ABST
    Figure CN120493870A_ABST
Patent Text Reader

Abstract

The invention provides a Markdown source code paging method and device, electronic equipment and a storage medium, and relates to the technical field of document processing.The method comprises the steps that firstly, an original Markdown source code is obtained, and the original Markdown source code is converted into an original HTML source code; performing color marking on elements in the original HTML source code to obtain a marked HTML source code, and converting the marked HTML source code into a marked PDF (Portable Document Format) file; and finally, determining a single-page original Markdown source code based on the color information in the marked PDF file and the color information in the marked HTML source code. According to the method, the color mark is added into the original HTML source code, so that the original Markdown source code can be paged according to the color information in the marked HTML source code and the color information in the marked PDF file under the condition of keeping the layout information of the original Markdown source code basically unchanged, and the paging accuracy of the source code is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of document processing technology, and in particular to a Markdown source code paging method, device, electronic device and storage medium. Background Art

[0002] The emergence of large language models (LLMs), particularly large visual models, has made it possible to output structured information from entire pages of document images. Training large document models requires paired data: the input is a document image, and the output is the corresponding structured data. Training large document models requires a large amount of high-quality data, which is crucial for the accuracy and generalization of large document models. However, data annotation faces challenges such as high data complexity, the need for specialized annotation personnel, and high annotation costs.

[0003] Markdown, as a lightweight markup language, is also an output format that the document model needs to support. To enable the document model to output in Markdown format, a large number of images and corresponding Markdown source code are required as paired data to support the training of the document model. However, many Portable Document Format (PDF) files rendered from Markdown source code are currently multi-page, requiring both pagination of the PDF file and pagination of the Markdown source code to obtain paired data.

[0004] At present, Markdown source code paging technology is mainly divided into two categories: (1) Based on text matching: First, characters in the Markdown source code that are not related to text matching, such as comments, control symbols, etc., are filtered. Then, text is extracted from each page of the PDF file, and the content of the first and last page of each page is obtained based on the relative position of the text. These texts provide key clues for locating the paging position. Subsequently, the text of each page is spliced in a certain order, and finally the paging position in the Markdown source code is determined by text matching or directly using a pre-trained classifier to obtain a single page of Markdown source code. (2) Based on pre-rendering: By manually calling the browser application programming interface (Application Programming Interface, API) element acquisition tool, the page container Document Object Model (DOM) element of the PDF file to be generated is extracted, and the width and height of the element and its position information in the page are calculated. Based on the height and width of the element and the layout information of the current page, it is determined whether the element should be retained on the current page or the layout needs to be adjusted and placed on the next page. By using real-time pre-rendering results, we adjust the layout of elements within the page container or fill the page container with blank elements, ensuring that each element is displayed as completely as possible within a single page, thus maintaining information coherence. This pre-rendering method makes it easy to obtain elements located at page boundaries and extract the source code information of a single page.

[0005] However, based on the text matching method, the spatial layout of text in PDF files does not always follow a simple left-to-right and top-to-bottom pattern. This uncertainty may lead to deviations in the order of extracted text, thereby causing matching errors and paging problems; the representation of certain special symbols in the source code is quite different from the characters extracted after PDF rendering, and this difference will significantly affect the accuracy of paging.

[0006] Pre-rendering solutions, on the other hand, can't completely prevent elements from crossing pages. Once they do, paginating the source code becomes more difficult. Furthermore, this method changes the layout of the original Markdown source code, affecting the quality and diversity of the rendered layout. Furthermore, this method requires pre-rendering results, which places limitations on some existing tools and may not meet your needs. Summary of the Invention

[0007] The present invention provides a Markdown source code paging method, device, electronic device and storage medium to solve the defects existing in the related art.

[0008] The present invention provides a Markdown source code paging method, comprising: Obtaining original Markdown source code, and converting the original Markdown source code into original HTML source code; Color-marking the elements in the original HTML source code to obtain a marked HTML source code, and converting the marked HTML source code into a marked PDF file; The original Markdown source code of a single page is determined based on the color information in the marked PDF file and the color information in the marked HTML source code.

[0009] According to a Markdown source code paging method provided by the present invention, the elements in the original HTML source code include text elements and image elements; the color marking of the elements in the original HTML source code to obtain the marked HTML source code includes: Splitting the text element by characters to obtain target text elements; The target text element and the image element are color-marked to obtain the marked HTML source code.

[0010] According to a Markdown source code paging method provided by the present invention, color marking of the image elements is performed, including: Obtaining the color value of the marker of the image element; generating a full color image corresponding to the image-like element based on the marked color value; The path of the image element is replaced by the path of the full-color image.

[0011] According to a Markdown source code paging method provided by the present invention, determining a single page of original Markdown source code based on the color information in the marked PDF file and the color information in the marked HTML source code includes: Based on the color information in the marked PDF file and the color information in the marked HTML source code, the marked HTML source code is segmented to obtain a single page of marked HTML source code; Determining the original HTML source code of the single page based on the single page markup HTML source code; Based on the single-page original HTML source code, determine the single-page original Markdown source code.

[0012] According to a Markdown source code paging method provided by the present invention, the Markdown source code is segmented based on the color information in the Markdown PDF file and the color information in the Markdown HTML source code to obtain a single page of Markdown HTML source code, including: Based on the color information in the marked PDF file and the color information in the marked HTML source code, determining whether the element at the page break in the marked PDF file is a cross-page element; Based on the judgment result, the markup HTML source code is segmented to obtain the single-page markup HTML source code.

[0013] According to a Markdown source code paging method provided by the present invention, the elements in the marked HTML source code include original elements and newly added elements; and determining the original HTML source code of a single page based on the single-page marked HTML source code includes: Determine the target element included in the HTML source code of the single page markup; Restoring the color information of the original element in the target element and removing the color information of the newly added element in the target element; The newly added elements are merged to obtain the original HTML source code of the single page.

[0014] According to a Markdown source code paging method provided by the present invention, the color marking of elements in the original HTML source code includes: An identifier is added to each element in the original HTML source code.

[0015] According to a Markdown source code paging method provided by the present invention, the method determines a single page of original Markdown source code based on the color information in the marked PDF file and the color information in the marked HTML source code, and then includes: Converting the original HTML source code into an original PDF file; Paginate the original PDF file to obtain a single-page PDF, and convert the single-page PDF into a single-page image; Bind the single-page original Markdown source code with the single-page image.

[0016] The present invention also provides a Markdown source code paging device, comprising: A source code conversion module is used to obtain the original Markdown source code and convert the original Markdown source code into original HTML source code; A color marking module is used to color mark the elements in the original HTML source code to obtain marked HTML source code, and convert the marked HTML source code into a marked PDF file; The source code paging module is used to determine a single page of original Markdown source code based on the color information in the marked PDF file and the color information in the marked HTML source code.

[0017] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the Markdown source code paging method described above is implemented.

[0018] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described Markdown source code paging methods.

[0019] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-described Markdown source code paging methods.

[0020] The present invention provides a Markdown source code paging method, device, electronic device, and storage medium. The method first obtains original Markdown source code and converts the original Markdown source code into original HTML source code; then, color-tags elements in the original HTML source code to obtain tagged HTML source code, and converts the tagged HTML source code into a tagged PDF file; and finally, based on the color information in the tagged PDF file and the color information in the tagged HTML source code, determines a single page of original Markdown source code. By adding color tags to the original HTML source code, the method can paginate the original Markdown source code based on the color information in the tagged HTML source code and the color information in the tagged PDF file, while keeping the layout information of the original Markdown source code substantially unchanged, thereby achieving higher paging accuracy for the source code. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 This is one of the flow charts of the Markdown source code paging method provided by the present invention.

[0023] Figure 2 This is a tree structure diagram of each element in the original HTML source code in the Markdown source code paging method provided by the present invention.

[0024] Figure 3This is a schematic diagram of splitting text elements by characters in the Markdown source code paging method provided by the present invention.

[0025] Figure 4 This is a schematic diagram of nodes retained after deleting a node in the Markdown source code paging method provided by the present invention.

[0026] Figure 5 This is a schematic diagram of merging newly added nodes in the Markdown source code paging method provided by the present invention.

[0027] Figure 6 This is the second flow chart of the Markdown source code paging method provided by the present invention.

[0028] Figure 7 This is a schematic diagram of the structure of Markdown source code paging provided by the present invention.

[0029] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0030] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0031] Since existing Markdown source code paging technologies all have various problems, based on this, an embodiment of the present invention provides a Markdown source code paging method.

[0032] Figure 1 This is a flow chart of the Markdown source code paging method provided in an embodiment of the present invention. Figure 1 As shown, the method includes: S1, obtaining original Markdown source code, and converting the original Markdown source code into original HTML source code; S2, color-marking the elements in the original HTML source code to obtain a marked HTML source code, and converting the marked HTML source code into a marked PDF file; S3, determining a single page of original Markdown source code based on the color information in the marked PDF file and the color information in the marked HTML source code.

[0033] Specifically, the Markdown source code paging method provided in the embodiment of the present invention is executed by a Markdown source code paging device, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., which is not specifically limited here.

[0034] First, step S1 is executed to obtain the original Markdown source code. The original Markdown source code can be directly uploaded by the user or obtained through other means.

[0035] Next, the original Lightweight Markup Language (Markdown) source code is converted to original Hypertext Markup Language (HTML) source code. This conversion process can be accomplished in a variety of ways. Common implementations include using the Pandoc command-line tool, the JavaScript library Markdown-it, online editors such as Dillinger, and the Python library Markdown. Once the original HTML source code is obtained, it can be further modified using structured parsing tools.

[0036] The elements in the original HTML source code can include text elements and image elements according to the format. Text elements are related elements that describe text content and can include paragraph elements. , title element <h1> to< / h1> <h6>, inline container , code block <pre>,official 、强调元素 <em>和强烈强调元素<strong>等。图像类元素是描述图像的相关元素。

[0037] 然后执行步骤S2,可以先记录原始HTML源码中各元素的颜色属性,然后去除各元素的颜色属性,重新对原始HTML源码中各元素进行颜色标记,得到标记HTML源码。该颜色标记是指对原始HTML源码中各元素统一进行全局唯一颜色标记,且各元素重新标记的颜色均不相同,如此可以通过各元素的颜色属性,对各元素进行区分。

[0038] 可以理解的是,HTML源码支持多种颜色表示方式,包括但不限于预定义颜色名称、十六进制颜色码、RGB、RGBA、HSL和HSLA等。其中,RGB颜色值作为一种常用的颜色表示方法,通过红、绿、蓝三种颜色的强度来定义色彩,每种颜色的取值范围为0至255,可组合出16777215种不同的颜色。在实际操作中,可以使用rgb()函数指定颜色,例如,红色表示为rgb(255,0,0),绿色为rgb(0,255,0),蓝色为rgb(0,0,255)。

[0039] 颜色标记的方式会根据元素的类型有所不同。对于文本类元素,可以直接添加颜色修饰符进行颜色标记。对于图像类元素,则需要利用其标记颜色值确定对应的全色图实现颜色标记。

[0040] 此后,可以对标记HTML源码进行渲染,将标记HTML源码转换为标记PDF文件。此处,可以利用浏览器打印功能,也可以利用wkHTMLtoPDF、Puppeteer等渲染工具,对标记HTML源码进行处理,以生成标记PDF文件。

[0041] 可以理解的是,利用同样的渲染工具,添加颜色标记的标记HTML源码生成的标记PDF文件,与未添加颜色标记的原始HTML源码生成的原始PDF文件相比,二者在排版上一致,只是对应位置处的元素的颜色属性存在差异而已。

[0042] 最后,执行步骤S3,利用标记PDF文件中的颜色信息以及标记HTML源码中的颜色信息,确定单页原始Markdown源码。此处,可以通过PDF解析工具,对标记PDF文件进行解析,获取标记PDF文件中各元素的内容、颜色和位置。

[0043] 本发明实施例中,标记PDF文件的每一页中的元素可以包括字符和图像。标记PDF文件中的颜色信息可以包括每一页中各元素的颜色,标记HTML源码中的颜色信息可以包括标记HTML源码中各元素重新标记的颜色。标记HTML源码中的颜色信息与标记PDF文件中的颜色信息按元素一一对应。

[0044] 由于标记PDF文件已经完成分页,则可以利用标记PDF文件各页中的颜色信息与标记HTML源码中的颜色信息进行比对,确定标记PDF文件中分页处对应的标记HTML源码中的元素,并将标记HTML源码恢复为原始HTML源码,进而对原始HTML源码进行分页,得到单页原始HTML源码,将单页原始HTML源码转换为单页原始Markdown源码。此处,PDF文件中的分页处可以包括前一页的末位和后一页的首位。

[0045] 此外,还可以交换分页动作与源码恢复动作的顺序,即先对标记HTML源码进行分页,得到单页标记HTML源码,进而将单页标记HTML源码恢复为单页原始HTML源码。

[0046] 本发明实施例中提供的Markdown源码分页方法,首先获取原始Markdown源码,并将原始Markdown源码转换为原始HTML源码;然后对原始HTML源码中的元素进行颜色标记,得到标记HTML源码,并将标记HTML源码转换为标记PDF文件;最后基于标记PDF文件中的颜色信息以及标记HTML源码中的颜色信息,确定单页原始Markdown源码。该方法通过在原始HTML源码中加入颜色标记,即可根据标记HTML源码中的颜色信息和标记PDF文件中的颜色信息,在保持原始Markdown源码的布局信息基本不变的情况下,对原始Markdown源码实现分页,源码的分页准确性更高。

[0047] 在上述实施例的基础上,所述原始HTML源码中的元素包括文本类元素和图像类元素;所述对所述原始HTML源码中的元素进行颜色标记,得到标记HTML源码,包括:对所述文本类元素按字符进行拆分,得到目标文本类元素;对所述目标文本类元素与所述图像类元素进行颜色标记,得到所述标记HTML源码。

[0048] 具体地,由于原始Markdown源码的基本元素包括标题、段落、换行、强调、列表、链接、图片和代码等。将原始Markdown源码转换为目标PDF文件时,目标PDF文件中分页处的元素主要取决于原始Markdown源码的内容和格式,以及转换工具的处理方式。具体分析如下:(1)原始Markdown源码中的文本:在分页处,可能会有章节标题、段落或列表项。通常,转换工具会自动避免将这些元素截断,确保每个章节的标题完全显示在新的一页上。(2)原始Markdown源码中的图像和表格:大型图像或表格可能需要跨越多个页面。良好的转换工具会确保这些元素不会被截断,而是完整地跨越所需的页数。例如,一个大图表可能会从第1页开始,连续到第3页。(3)原始Markdown源码中的页脚和页眉:转换工具可能允许添加页脚和页眉,这些元素会在每一页的顶部和底部出现,包括页码、文档标题或章节名。(4)原始Markdown源码中的代码块:如果代码块在分页处被截断,会影响阅读和理解。高质量的转换工具会尽量处理这一问题,可能通过调整页边距或微调内容位置来实现。

[0049] 在处理原始Markdown源码中可能存在的跨页元素时,通过前述分析可知,页码、页眉和页脚等元素不会受到跨页问题的影响,并且它们属于全局相关元素,因此可以统一归类为special_element。其他元素则可能存在跨页问题,需要进行进一步拆分,以确保每个元素的渲染结果尽量避免跨栏现象。除了文本和图像之外,表格和代码本质上也由文本和图像元素组成,因此可以将所有元素简化为两大类:文本和图像。由于图像元素已被视为最细粒度的元素,因此无需进一步拆分,拆分的重点应放在文本类元素上。

[0050] 因此,对于原始HTML源码中的文本类元素,需要对其进行划分,其中普通文本按照字符分开,特殊字符需要保持独立,部分特殊字符如表1所示。

[0051] 表1 特殊字符然后,可以利用对每个独立的字符进行修饰。通过这种方式,可以将字符按最小单元进行拆分,从而使当前的元素生成新的元素。

[0052] 对文本类元素进行拆分后,即可得到目标文本类元素。该目标文本类元素中包含有未进行拆分的原始的文本类元素以及拆分得到的新增的文本类元素。

[0053] 最后,对目标文本类元素与图像类元素进行颜色标记,得到标记HTML源码。该标记HTML源码中可以包括原始元素和新增元素,原始元素是原始HTML源码中的元素,包括文本类元素和图像类元素,新增元素为经过拆分新增的元素,例如新增的文本类元素。

[0054] 本发明实施例中,先对文本类元素按字符进行拆分,然后再进行颜色标记,可以使标记HTML源码中的元素更加细粒度化,以降低标记PDF文件中的跨页元素的出现概率,从而提高分页质量。

[0055] 在上述实施例的基础上,对所述图像类元素进行颜色标记,包括:获取所述图像类元素的标记颜色值;基于所述标记颜色值生成与所述图像类元素对应的全色图;采用全色图的路径替换所述图像类元素的路径。

[0056] 具体地,在对图像类元素进行颜色标记时,可以先获取图像类元素的标记颜色值,例如某一图像类元素的标记颜色值可以表示为rgb(0,255,0);然后利用标记颜色值生成与图像类元素对应的全色图,例如纯绿色图。该全色图需要与对应的图像类元素尺寸和格式均一致,以确保渲染后布局不发生变化。

[0057] 最终,采用全色图的路径替换图像类元素的路径,即将图像类元素进行颜色标记得到对应的全色图,可以将标签的基于稀疏表达的分类(sparse representation-based classifier,SRC)属性值更新为全色图的统一资源定位地址(Uniform ResourceLocator,URL)实现。

[0058] 在上述实施例的基础上,所述基于所述标记PDF文件中的颜色信息以及所述标记HTML源码中的颜色信息,确定单页原始Markdown源码,包括:基于所述标记PDF文件中的颜色信息以及所述标记HTML源码中的颜色信息,将所述标记HTML源码进行切分,得到单页标记HTML源码;基于所述单页标记HTML源码,确定单页原始HTML源码;基于所述单页原始HTML源码,确定单页原始Markdown源码。

[0059] 具体地,在确定单页原始Markdown源码时,可以先利用标记PDF文件中的颜色信息以及标记HTML源码中的颜色信息,将标记HTML源码进行切分,得到单页标记HTML源码。即通过标记PDF文件中的颜色信息以及标记HTML源码中的颜色信息的对应关系,结合标记PDF文件的分页信息,可以确定标记HTML源码的分页信息,即确定标记HTML源码的单页标记HTML源码。

[0060] 此后,可以将单页标记HTML源码恢复为单页原始HTML源码,该恢复过程为颜色标记的逆过程。

[0061] 最终,可以对单页原始HTML源码进行转换得到单页Markdown源码。该转换过程既可以通过手动编写程序实现,也可以借助Pandoc等工具或在线平台进行转换,此处不作具体限定。

[0062] 本发明实施例中,给出了确定单页原始Markdown源码的具体步骤,通过HTML源码与Markdown源码之间的转换,可以提高单页原始Markdown源码的确定效率。

[0063] 由于现有的文本匹配无法区分分页处的图像元素,无法准确判断它们是位于上一页末尾还是当前页开头。基于此,在上述实施例的基础上,所述基于所述标记PDF文件中的颜色信息以及所述标记HTML源码中的颜色信息,将所述标记HTML源码进行切分,得到单页标记HTML源码,包括:基于所述标记PDF文件中的颜色信息以及所述标记HTML源码中的颜色信息,判断所述标记PDF文件中分页处的元素是否为跨页元素;基于判断结果,将所述标记HTML源码进行切分,得到所述单页标记HTML源码。

[0064] 具体地,在确定单页标记HTML源码时,可以先利用标记PDF文件中的颜色信息以及标记HTML源码中的颜色信息,判断标记PDF文件中分页处的元素是否为跨页元素。即可以判断分页处前一页的末位元素与后一页的首位元素是否为同一元素,若为同一元素,则说明该元素为跨页元素。否则,说明标记PDF文件中分页处的元素不是跨页元素。

[0065] 此后,可以利用标记PDF文件中各分页处的判断结果,可以确定标记HTML源码中每一页包含的目标元素,进而可以对标记HTML源码进行切分,得到单页标记HTML源码,如此可以提高单页标记HTML源码的分页准确性。

[0066] 在上述实施例的基础上,所述标记HTML源码中的元素包括原始元素和新增元素;所述基于所述单页标记HTML源码,确定单页原始HTML源码,包括:确定所述单页标记HTML源码中包含的目标元素;恢复所述目标元素中所述原始元素的颜色信息,并去除所述目标元素中的新增元素的颜色信息;对所述新增元素进行合并,得到所述单页原始HTML源码。

[0067] 具体地,在确定单页原始HTML源码的过程中,可以先利用上述判断结果,确定出标记HTML源码中每一页面包含的目标元素,即单页标记HTML源码中包含的目标元素。

[0068] 然后,对目标元素进行去颜色标记,即恢复目标元素中原始元素的颜色信息,去除目标元素中的新增元素的颜色信息,进而通过对新增元素进行合并,即可得到单页HTML源码。

[0069] 在上述实施例的基础上,所述对所述原始HTML源码中的元素进行颜色标记,之前包括:对所述原始HTML源码中的各元素添加标识符。

[0070] 具体地,本发明实施例中,还可以对原始HTML源码中的各元素添加标识符,原始HTML源码中每个元素可以通过添加唯一标识符进行标记。该标识符可以通过标识(Identity document,ID)全局属性的取值进行定义,该ID全局属性的取值具有全局唯一性。例如,对于原始HTML源码中的元素HTML®,可以为其添加一个ID全局属性,如HTML®。其中,ID全局属性的取值为"0",使得该段落元素在整个原始HTML源码中唯一地标识为"0"。为了实现对每个元素添加标识符,可以采用广度优先遍历的方式,假如原始HTML源码中每个元素的树形结构如图2所示,根节点被标记为"0",有三个直接子节点,按照广度优先原则,依次访问这些子节点,标记为"0-0"、"0-1"和"0-2"。由于根节点的标识符为"0",因此这些子节点的最终标记符分别为"0-0"、"0-1"和"0-2"。其他节点可以通过类似方式,通过父节点的标记符进行扩展。例如,节点"0-0"有两个子节点,它们的最终标记符分别为"0-0-0"和"0-0-1"。这种标记方式的一个显著优点是,可以通过一个元素的标记符推断出其父节点、祖父节点,甚至一直追溯到根节点,从而方便后续的操作和处理。

[0071] 在对文本类元素按字符进行拆分时,可以利用对每个独立的字符进行修饰。例如,HTML®,可拆解为HTML®。通过这种方式,可以将字符按最小单元进行拆分,从而使文本类元素生成新增的文本类元素。新增的文本类元素仍然需要添加唯一标识符,也就是说,对于标记HTML源码中的新增元素也需要添加唯一标识符,即HTML®。

[0072] 图2经过拆分后可能会变为图3,其中加粗的叶子节点表示新生成的叶子节点,里面的新增的节点同样需要需要添加唯一标识符,节点"0-0-0”拆分后由"0-0-0-0”,"0-0-0-1”,"0-0-0-2”和"0-0-0-3”4个子节点表示,4个子节点均可进行单字符拆分处理。

[0073] 在对原始HTML源码中的元素进行颜色标记之前,可以先根据各元素的标识符记录各元素的颜色属性,方便后续根据各元素的标识符恢复各元素的颜色属性。然后,去除各元素的颜色属性,便于后续进行颜色标记。

[0074] 在对原始HTML源码中的文本类元素进行颜色标记时,可以直接添加颜色修饰符,如HTML®。

[0075] 通过对标记HTML源码进行深度优先遍历,可以记录每个叶子节点的标识符(ID)及其对应的颜色标记值(color),并以(ID,color)的形式存储在列表中,记为pair_list。此后,可以从左到右访问pair_list,如果当前叶子节点的颜色标记值同时出现在标记PDF文件的相邻两个页面中,则表明当前叶子节点对应的元素位于标记PDF文件的分页处,且可能存在跨栏问题(例如超大图片),属于跨页元素。如果相邻的两个叶子节点的颜色值分别对应于标记PDF文件的相邻两个页面,则可以判定此处为一个分界点,两个叶子节点分别属于不同的页面。

[0076] 通过跨页元素的ID信息,可以确定每一页所包含的叶子节点。根据每一页的叶子节点,保留从叶子节点到根节点路径上的所有节点,其余节点则删除。若图3中加粗的叶子节点”0-0-0-2”,”0-0-0-3”,”0-0-1”,"0-1-0”,”0-2-0-0”和”0-2-0-1"属于同一页,则其保留的节点如图4所示,剩余的节点即表示当前页的内容。对于这些剩余节点,去除新增的节点的颜色属性,而非新增的节点则恢复其原始的颜色属性。最后,将新增的节点进行合并,得到最终的单页原始HTML源码。合并示例如图5所示,只对新增的节点进行合并,合并可以看成拆分的逆向操作,唯一的区别是合并时部分叶子节点被删除。合并方式例如,ML®可以合并为ML®。

[0077] 本发明实施例中,通过添加标识符,可以简化后续操作的复杂度。

[0078] 在上述实施例的基础上,所述基于所述标记PDF文件中的颜色信息以及所述标记HTML源码中的颜色信息,确定单页原始Markdown源码,之后包括:将所述原始HTML源码转换为原始PDF文件;对所述原始PDF文件进行分页,得到单页PDF,并将所述单页PDF转换为单页图片;将所述单页原始Markdown源码与所述单页图片进行绑定。

[0079] 具体地,在确定单页原始Markdown源码之后,还可以将原始HTML源码转换为原始PDF文件。此处,可以利用浏览器打印功能,也可以利用wkHTMLtoPDF、Puppeteer等渲染工具,对原始HTML源码进行处理,以生成原始PDF文件。

[0080] 此后,对原始PDF文件进行分页,得到单页PDF,并将单页PDF转换为单页图片,将单页原始Markdown源码与单页图片进行绑定,以获取配对数据支持文档大模型的训练。

[0081] 图6为本发明实施例中提供的Markdown源码分页方法的完整流程示意图,如图6所示,该方法包括:获取原始Markdown源码;将原始Markdown源码转换为原始HTML源码;对原始HTML源码中的各元素添加标识符;一方面,将原始HTML源码中的文本类元素按字符进行拆分,得到目标文本类元素,对目标文本类元素与图像类元素进行颜色标记,得到标记HTML源码。

[0082] 此后,将标记HTML源码转换为标记PDF文件,提取标记PDF文件中的颜色信息,并确定标记PDF文件中每一页包含的颜色信息,进而确定单页标记HTML源码中包含的目标元素。

[0083] 对目标元素进行去颜色标记,即恢复目标元素中原始元素的颜色信息,并去除目标元素中的新增元素的颜色信息,对新增元素进行合并,得到单页原始HTML源码。

[0084] 将单页原始HTML源码进行源码转换,得到单页原始Markdown源码。

[0085] 另一方面,将原始HTML源码转换为原始PDF文件,然后对原始PDF文件进行分页,得到单页PDF,并将单页PDF转换为单页图片,将单页原始Markdown源码与单页图片进行绑定,得到一个配对数据(pair)。

[0086] 如图7所示,在上述实施例的基础上,本发明实施例中提供了一种Markdown源码分页装置,包括:源码转换模块71,用于获取原始Markdown源码,并将所述原始Markdown源码转换为原始HTML源码;颜色标记模块72,用于对所述原始HTML源码中的元素进行颜色标记,得到标记HTML源码,并将所述标记HTML源码转换为标记PDF文件;源码分页模块73,用于基于所述标记PDF文件中的颜色信息以及所述标记HTML源码中的颜色信息,确定单页原始Markdown源码。

[0087] 在上述实施例的基础上,本发明实施例中提供的Markdown源码分页装置,所述原始HTML源码中的元素包括文本类元素和图像类元素;所述对所述原始HTML源码中的元素进行颜色标记,得到标记HTML源码,包括:对所述文本类元素按字符进行拆分,得到目标文本类元素;对所述目标文本类元素与所述图像类元素进行颜色标记,得到所述标记HTML源码。

[0088] 在上述实施例的基础上,本发明实施例中提供的Markdown源码分页装置,对所述图像类元素进行颜色标记,包括:获取所述图像类元素的标记颜色值;基于所述标记颜色值生成与所述图像类元素对应的全色图;采用全色图的路径替换所述图像类元素的路径。

[0089] 在上述实施例的基础上,本发明实施例中提供的Markdown源码分页装置,所述基于所述标记PDF文件中的颜色信息以及所述标记HTML源码中的颜色信息,确定单页原始Markdown源码,包括:基于所述标记PDF文件中的颜色信息以及所述标记HTML源码中的颜色信息,将所述标记HTML源码进行切分,得到单页标记HTML源码;基于所述单页标记HTML源码,确定单页原始HTML源码;基于所述单页原始HTML源码,确定单页原始Markdown源码。

[0090] 在上述实施例的基础上,本发明实施例中提供的Markdown源码分页装置,所述基于所述标记PDF文件中的颜色信息以及所述标记HTML源码中的颜色信息,将所述标记HTML源码进行切分,得到单页标记HTML源码,包括:基于所述标记PDF文件中的颜色信息以及所述标记HTML源码中的颜色信息,判断所述标记PDF文件中分页处的元素是否为跨页元素;基于判断结果,将所述标记HTML源码进行切分,得到所述单页标记HTML源码。

[0091] 在上述实施例的基础上,本发明实施例中提供的Markdown源码分页装置,所述标记HTML源码中的元素包括原始元素和新增元素;所述基于所述单页标记HTML源码,确定单页原始HTML源码,包括:确定所述单页标记HTML源码中包含的目标元素;恢复所述目标元素中所述原始元素的颜色信息,并去除所述目标元素中的新增元素的颜色信息;对所述新增元素进行合并,得到所述单页原始HTML源码。

[0092] 在上述实施例的基础上,本发明实施例中提供的Markdown源码分页装置,所述对所述原始HTML源码中的元素进行颜色标记,之前包括:对所述原始HTML源码中的各元素添加标识符。

[0093] 在上述实施例的基础上,本发明实施例中提供的Markdown源码分页装置,所述基于所述标记PDF文件中的颜色信息以及所述标记HTML源码中的颜色信息,确定单页原始Markdown源码,之后包括:将所述原始HTML源码转换为原始PDF文件;对所述原始PDF文件进行分页,得到单页PDF,并将所述单页PDF转换为单页图片;将所述单页原始Markdown源码与所述单页图片进行绑定。

[0094] 具体地,本发明实施例中提供的Markdown源码分页装置中各模块的作用与上述方法类实施例中各步骤的操作流程是一一对应的,实现的效果也是一致的,具体参见上述实施例,本发明实施例中对此不再赘述。

[0095] 图8示例了一种电子设备的实体结构示意图,如图8所示,该电子设备可以包括:处理器(processor)810、通信接口(Communications Interface)820、存储器(memory)830和通信总线840,其中,处理器810,通信接口820,存储器830通过通信总线840完成相互间的通信。处理器810可以调用存储器830中的逻辑指令,以执行上述各实施例中提供的Markdown源码分页方法。

[0096] 此外,上述的存储器830中的逻辑指令可以通过软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本发明的技术方案本质上或者说对相关技术做出贡献的部分或者该技术方案的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行本发明各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、磁碟或者光盘等各种可以存储程序代码的介质。

[0097] 另一方面,本发明还提供一种计算机程序产品,所述计算机程序产品包括计算机程序,计算机程序可存储在计算机可读存储介质上,所述计算机程序被处理器执行时,计算机能够执行上述各实施例中提供的Markdown源码分页方法。

[0098] 又一方面,本发明还提供一种计算机可读存储介质,其上存储有计算机程序,该计算机程序被处理器执行时实现以执行上述各实施例中提供的Markdown源码分页方法。该计算机可读存储介质既可以是非暂态计算机可读存储介质,也可以是暂态计算机可读存储介质,此处不作具体限定。

[0099] 以上所描述的装置实施例仅仅是示意性的,其中所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部模块来实现本实施例方案的目的。本领域普通技术人员在不付出创造性的劳动的情况下,即可以理解并实施。

[0100] 通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到各实施方式可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件。基于这样的理解,上述技术方案本质上或者说对相关技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品可以存储在计算机可读存储介质中,如ROM / RAM、磁碟、光盘等,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行各个实施例或者实施例的某些部分所述的方法。

[0101] 最后应说明的是:以上实施例仅用以说明本发明的技术方案,而非对其限制;尽管参照前述实施例对本发明进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本发明各实施例技术方案的精神和范围。< / strong>< / em> < / pre> < / h6>

Claims

1. A Markdown source code paging method, characterized in that: include: Obtaining original Markdown source code, and converting the original Markdown source code into original HTML source code; Color-marking the elements in the original HTML source code to obtain a marked HTML source code, and converting the marked HTML source code into a marked PDF file; The original Markdown source code of a single page is determined based on the color information in the marked PDF file and the color information in the marked HTML source code.

2. The Markdown source code paging method according to claim 1, characterized in that: The elements in the original HTML source code include text elements and image elements; the color marking of the elements in the original HTML source code to obtain the marked HTML source code includes: Splitting the text element by characters to obtain target text elements; The target text element and the image element are color-marked to obtain the marked HTML source code.

3. The Markdown source code paging method according to claim 2, characterized in that: Marking the image elements with colors includes: Obtaining the color value of the marker of the image element; generating a full color image corresponding to the image-like element based on the marked color value; The path of the image element is replaced by the path of the full-color image.

4. The Markdown source code paging method according to claim 1, characterized in that: The determining of the single-page original Markdown source code based on the color information in the marked PDF file and the color information in the marked HTML source code includes: Based on the color information in the marked PDF file and the color information in the marked HTML source code, the marked HTML source code is segmented to obtain a single page of marked HTML source code; Determining the original HTML source code of the single page based on the single page markup HTML source code; Based on the single-page original HTML source code, determine the single-page original Markdown source code.

5. The Markdown source code paging method according to claim 4, characterized in that: The step of segmenting the markup HTML source code based on the color information in the markup PDF file and the color information in the markup HTML source code to obtain a single-page markup HTML source code includes: Based on the color information in the marked PDF file and the color information in the marked HTML source code, determining whether the element at the page break in the marked PDF file is a cross-page element; Based on the judgment result, the markup HTML source code is segmented to obtain the single-page markup HTML source code.

6. The Markdown source code paging method according to claim 4, characterized in that: The elements in the marked HTML source code include original elements and newly added elements; and determining the original HTML source code of a single page based on the marked HTML source code of the single page includes: Determine the target element included in the HTML source code of the single page markup; Restoring the color information of the original element in the target element and removing the color information of the newly added element in the target element; The newly added elements are merged to obtain the original HTML source code of the single page.

7. The Markdown source code paging method according to any one of claims 1 to 5, characterized in that: The color marking of the elements in the original HTML source code previously includes: An identifier is added to each element in the original HTML source code.

8. The Markdown source code paging method according to any one of claims 1 to 5, characterized in that: The step of determining the original Markdown source code of a single page based on the color information in the marked PDF file and the color information in the marked HTML source code further includes: Converting the original HTML source code into an original PDF file; Paginate the original PDF file to obtain a single-page PDF, and convert the single-page PDF into a single-page image; Bind the single-page original Markdown source code with the single-page image.

9. A Markdown source code paging device, characterized in that: include: A source code conversion module is used to obtain the original Markdown source code and convert the original Markdown source code into original HTML source code; A color marking module is used to color mark the elements in the original HTML source code to obtain marked HTML source code, and convert the marked HTML source code into a marked PDF file; The source code paging module is used to determine a single page of original Markdown source code based on the color information in the marked PDF file and the color information in the marked HTML source code.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the Markdown source code paging method according to any one of claims 1 to 8 is implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the Markdown source code paging method according to any one of claims 1 to 8 is implemented.

12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the Markdown source code paging method according to any one of claims 1 to 8 is implemented.