Data processing method and device, computer equipment and computer readable storage medium

CN122797461APending Publication Date: 2026-09-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510344708.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0004]但是,相关技术中,在页面设计工具中还原PDF元素时元素位置数据处理的准确性不高

Benefits of technology

[0056]本公开实施例提供的数据处理方法,通过获取便携式文档格式文件,并对便携式文档格式文件进行文件解析,得到便携式文档格式文件中多个字符元素的字符元素信息,字符元素信息包括第一字符位置信息和字符属性信息;根据字符属性信息检测多个字符元素之间的行关联关系,并根据行关联关系对多个字符元素的第一字符位置信息进行调整,得到每一字符元素信息对应的第二字符位置信息;基于字符属性信息、第二字符位置信息以及预设的行间距阈值确定多个字符元素之间的段落关系,并根据段落关系对多个字符元素的第二字符位置信息进行调整,得到每一字符元素信息对应的第三字符位置信息;根据第三字符位置信息与字符属性信息生成便携式文档格式文件对应的界面设计数据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797461A_ABST
    Figure CN122797461A_ABST
Patent Text Reader

Abstract

The present disclosure provides a data processing method and device, computer equipment and computer readable storage medium. The method comprises: obtaining a portable document format file and parsing the portable document format file to obtain character element information of a plurality of character elements; detecting a line association relationship between the plurality of character elements according to the character attribute information, and adjusting first character position information of the plurality of character elements according to the line association relationship to obtain second character position information; determining a paragraph relationship between the plurality of character elements based on the character attribute information, the second character position information and a preset line spacing threshold, and adjusting the second character position information according to the paragraph relationship to obtain third character position information; and generating interface design data corresponding to the portable document format file according to the third character position information and the character attribute information. The method can improve the accuracy of processing element position data when restoring a portable document format file in a vector graphics editing application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a data processing method, apparatus, computer equipment, and computer-readable storage medium. Background Technology

[0002] With the rapid development of internet technology, various internet applications have emerged, providing convenient services and rich functions for all aspects of people's lives.

[0003] Internet applications interact with users through application pages. To pursue a better user experience, the demand for fully functional and highly interactive interfaces is constantly growing. When developing application interfaces, developers can draw each component individually in an interface design tool, but this process is time-consuming. To improve efficiency, elements from other document formats (such as Portable Document Format, PDF) can be reused to generate interface components. Therefore, it is necessary to recreate PDF elements in an interface design tool and convert them into editable components within the tool.

[0004] However, in related technologies, the accuracy of element position data processing is not high when restoring PDF elements in page design tools. Summary of the Invention

[0005] This disclosure provides a data processing method, apparatus, computer device, and computer-readable storage medium. The data processing method can improve the accuracy of element position data processing when restoring PDF files in vector graphics editing applications, thereby achieving better PDF file restoration results.

[0006] The first aspect of this disclosure provides a data processing method, the method comprising:

[0007] A portable document format file is obtained, and the portable document format file is parsed to obtain character element information of multiple character elements in the portable document format file. The character element information includes first character position information and character attribute information.

[0008] The row association relationship between the multiple character elements is detected based on the character attribute information, and the first character position information of the multiple character elements is adjusted according to the row association relationship to obtain the second character position information corresponding to each character element information;

[0009] Based on the character attribute information, the second character position information, and the preset line spacing threshold, the paragraph relationship between the multiple character elements is determined, and the second character position information of the multiple character elements is adjusted according to the paragraph relationship to obtain the third character position information corresponding to each character element information;

[0010] The interface design data corresponding to the portable document format file is generated based on the third character position information and the character attribute information.

[0011] A second aspect of this disclosure provides a data processing apparatus, the apparatus comprising:

[0012] The acquisition unit is used to acquire a portable document format file and parse the portable document format file to obtain character element information of multiple character elements in the portable document format file. The character element information includes first character position information and character attribute information.

[0013] The first adjustment unit is used to detect the row association relationship between the plurality of character elements according to the character attribute information, and adjust the first character position information of the plurality of character elements according to the row association relationship to obtain the second character position information corresponding to each character element information;

[0014] The second adjustment unit is used to determine the paragraph relationship between the plurality of character elements based on the character attribute information, the second character position information and the preset line spacing threshold, and to adjust the second character position information of the plurality of character elements according to the paragraph relationship to obtain the third character position information corresponding to each character element information.

[0015] The generation unit is used to generate interface design data corresponding to the portable document format file based on the third character position information and the character attribute information.

[0016] Optionally, in some embodiments, the first adjustment unit includes:

[0017] The calculation subunit is used to extract the character deformation parameters corresponding to each character element from the character attribute information, and calculate the rotation angle corresponding to each character element based on the character deformation parameters.

[0018] The detection subunit is used to detect the row association relationship between the plurality of character elements according to the rotation angle.

[0019] Optionally, in some embodiments, the detection subunit includes:

[0020] The first calculation module is used to extract the drawing origin coordinates corresponding to each character element from the character attribute information, and calculate the slope angle between any two character elements among the plurality of character elements based on the drawing origin coordinates.

[0021] The second calculation module is used to calculate the first distance value between any two character elements among the plurality of character elements based on the coordinates of the drawing origin.

[0022] The first detection module is used to detect the row association relationship between the plurality of character elements based on the slope angle, the first distance value, and the rotation angle.

[0023] Optionally, in some embodiments, the first detection module includes:

[0024] The first determining submodule is used to determine the first angle difference between the rotation angles of any two character elements, and to determine the second angle difference between the rotation angle of each character element and the corresponding slope angle in any two character elements.

[0025] The second determining submodule is used to determine a first comparison result between the first angle difference and a preset first angle difference threshold, a second comparison result between the second angle difference and a preset second angle difference threshold, and a third comparison result between the first distance value and a preset first distance threshold.

[0026] The detection submodule is used to detect the row association relationship between the plurality of character elements based on the first comparison result, the second comparison result, and the third comparison result.

[0027] Optionally, in some embodiments, the first adjustment unit includes:

[0028] The first determining subunit is used to divide the multiple character elements according to the row association relationship to obtain multiple text lines, and to determine the reference vertical coordinate corresponding to the text line based on the first character position information of the character elements in each text line.

[0029] The adjustment subunit is used to adjust the character ordinate corresponding to the first character position information of the character element in the text line according to the reference ordinate, so as to obtain the second character position information corresponding to each character element information.

[0030] Optionally, in some embodiments, the data processing apparatus provided in this disclosure further includes:

[0031] The extraction subunit is used to extract the font information corresponding to the character element in each text line from the character attribute information, and to extract the character horizontal coordinate corresponding to the character element in each text line from the first character position information;

[0032] The merging subunit is used to merge multiple character elements in each text line using the font information, the character horizontal coordinates, and the rotation angle.

[0033] Optionally, in some embodiments, the rotation angle and the first character position information are determined in a first coordinate system. The data processing apparatus provided in this disclosure further includes:

[0034] The transformation subunit is used to perform coordinate transformation based on the rotation angle corresponding to each character element and the adjusted position information of the first character to obtain the rotation angle and position information of the first character in the second coordinate system.

[0035] The second determining subunit is used to determine the second character position information corresponding to each character element information based on the rotation angle in the second coordinate system and the first character position information.

[0036] Optionally, in some embodiments, the second adjustment unit includes:

[0037] The third determining subunit is used to determine the line spacing between any two text lines based on the second character position information;

[0038] The fourth determining subunit is used to extract the font information corresponding to the character element from the character attribute information, and determine the paragraph relationship between the multiple character elements according to the line spacing, the font information and the preset line spacing threshold.

[0039] Optionally, in some embodiments, the second adjustment unit includes:

[0040] The sub-unit is used to divide the multiple text lines into multiple text paragraphs based on the paragraph relationship.

[0041] The adjustment subunit is used to adjust the character ordinate in the second character position information of the multiple character elements based on the preset reference line spacing and the line spacing of adjacent text lines in each of the text paragraphs, so as to obtain the third character position information corresponding to each character element information.

[0042] Optionally, in some embodiments, the data processing apparatus provided in this disclosure further includes:

[0043] The first acquisition subunit is used to acquire a second distance value and a third distance value between the boundary of each text paragraph and the boundary of each text line in the text paragraph, wherein the second distance value indicates the distance between the left side of the text line and the left side of the text paragraph, and the third distance value indicates the distance between the right side of the text line and the right side of the target merged paragraph;

[0044] The second acquisition subunit is used to acquire a fourth distance value between the midline of each text paragraph and the midline of each text line in the text paragraph;

[0045] The comparison subunit is used to compare the second distance value, the third distance value, and the fourth distance value with a preset distance threshold, and determine the alignment of each text segment based on the three comparison results.

[0046] Optionally, in some embodiments, the acquisition unit includes:

[0047] The parsing subunit is used to parse the portable document format file to obtain first character element data in Extensible Markup Language format and second character element data in Scalable Vector Graphics format.

[0048] The fifth determining subunit is used to determine the character element information of multiple character elements in the portable document format file based on the first character element data and the second character element data.

[0049] Optionally, in some embodiments, the fifth determining subunit includes:

[0050] The second detection module is used to detect text character elements and vector text elements in the first character element data and the second character element data, and obtain detection results;

[0051] The filtering module is used to filter the vector text elements when the detection result indicates that both text character elements and vector text elements exist in either the first character element data or the second character element data, and to determine the character element information of multiple character elements in the portable document format file based on the text character elements.

[0052] The determination module is used to determine the character element information of multiple character elements in the portable document format file based on the vector text elements when the detection result indicates that only vector text elements exist in both the first character element data and the second character element data.

[0053] A third aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the data processing method as described in the first aspect.

[0054] This fourth aspect of the disclosure provides a computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the data processing method as described in the first aspect.

[0055] The fifth aspect of this disclosure provides a computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the data processing method as described in the first aspect.

[0056] The data processing method provided in this embodiment obtains a portable document format file and parses the file to obtain character element information of multiple character elements in the portable document format file. The character element information includes first character position information and character attribute information. The method detects line relationships between the multiple character elements based on the character attribute information and adjusts the first character position information of the multiple character elements according to the line relationships to obtain second character position information corresponding to each character element. Based on the character attribute information, the second character position information, and a preset line spacing threshold, the method determines the paragraph relationship between the multiple character elements and adjusts the second character position information of the multiple character elements according to the paragraph relationship to obtain third character position information corresponding to each character element. Finally, the method generates interface design data corresponding to the portable document format file based on the third character position information and the character attribute information.

[0057] This disclosure improves the accuracy of character position information by parsing the character attribute information of character elements in a portable document format file to detect the line association relationship between multiple character elements. Then, based on the line association relationship, the first character position information obtained from the file parsing is adjusted. Furthermore, the adjusted first character position information is further adjusted by determining the paragraph relationship between multiple character elements, thereby improving the accuracy of the character position information. Compared to related technologies that directly generate interface design data from the character position information obtained from file parsing, this disclosure improves the accuracy of processing character position information by adjusting the parsed character position information multiple times, ultimately improving the accuracy of the generated interface design data.

[0058] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objectives and other advantages of this disclosure may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0059] The accompanying drawings are provided to further understand the technical solutions of this disclosure and constitute a part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.

[0060] Figure 1 This is a flowchart illustrating a data processing method in a related technology.

[0061] Figure 2 This is a comparison chart of the PDF reproduction and the original PDF in related technologies;

[0062] Figure 3 A system architecture diagram for the data processing method provided in the embodiments of this disclosure;

[0063] Figure 4 A schematic flowchart of a data processing method provided in an embodiment of this disclosure;

[0064] Figure 5 A schematic diagram of the vector graphics editing application plugin interface provided in the embodiments of this disclosure;

[0065] Figure 6 A schematic diagram illustrating the effective text detection results and corresponding processing methods provided in the embodiments of this disclosure;

[0066] Figure 7 A schematic diagram of the character rotation angle provided in an embodiment of this disclosure;

[0067] Figure 8 A schematic diagram of an slanted character element provided in an embodiment of this disclosure;

[0068] Figure 9 A schematic diagram of the process for detecting the association relationship of character element rows provided in this embodiment of the disclosure;

[0069] Figure 10 A comparison diagram of the SVG coordinate system and the vector graphics editing application coordinate system provided in this embodiment;

[0070] Figure 11 This is a schematic diagram illustrating the process of character elements transforming according to the transform field in this embodiment;

[0071] Figure 12 This is a schematic diagram illustrating the detection of text paragraph alignment provided in this embodiment;

[0072] Figure 13 A comparison diagram of directory nodes in the vector graphics editing application provided in this embodiment;

[0073] Figure 14 Another flowchart illustrating the data processing method provided in this disclosure;

[0074] Figure 15 Another flowchart illustrating the data processing method provided in this disclosure;

[0075] Figure 16 This is a schematic diagram illustrating the content composition of the PDF parsing data provided in this disclosure;

[0076] Figure 17 A schematic diagram of the position and style mapping table of character elements provided in this disclosure;

[0077] Figure 18 This is a schematic diagram of the structure of the data processing apparatus provided in the embodiments of this disclosure;

[0078] Figure 19 This is a terminal structure diagram for implementing various methods according to an embodiment of the present disclosure;

[0079] Figure 20 This is a server structure diagram illustrating the implementation of various methods according to an embodiment of the present disclosure. Detailed Implementation

[0080] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this disclosure.

[0081] Before providing a further detailed description of the embodiments of this disclosure, the terms and concepts used in these embodiments are explained, and they are subject to the following interpretations:

[0082] Portable Document Format (PDF): PDF is a file format used to present documents that include elements such as text, images, and tables. It preserves the original document's appearance and layout. This file format is independent of applications, hardware, and operating systems. Documents in this format typically consist of vector graphics, text, and bitmap graphics and can be used for engineering drawings, interface design drafts, and academic papers.

[0083] Extensible Markup Language (XML): XML is a markup language that can be used to mark up data and define data types. It is commonly used for data transmission and storage and can operate independently of software and hardware.

[0084] PyMuPDF is a high-performance Python library for extracting, analyzing, converting, and manipulating data from PDF documents.

[0085] Scalable Vector Graphics (SVG): SVG is an open standard format based on XML for describing two-dimensional vector graphics. In this embodiment of the disclosure, SVG format image information is extracted from a PDF document using the get_svg_image method in the PyMuPDF library. The extracted SVG format information may include element information of various elements in the document, such as text elements, image elements, and table elements. Specifically, for image elements, it may include the image length and width, as well as information such as the image's Uniform Resource Identifier (URL).

[0086] Domain-Specific Language (DSL): A DSL is a computer language specifically designed to solve problems in a particular domain. Compared to general-purpose programming languages ​​(such as Java and C++), DSLs have stronger expressive power within a specific domain, allowing developers to solve problems more efficiently. In this embodiment of the disclosure, a DSL is a computer language specifically used to define interface layout and behavior design, used to carry the structural data of PDF elements obtained after parsing a PDF document.

[0087] Vector graphics editing applications are applications used to create, modify, and manipulate vector graphics. Unlike pixel-based bitmap editing tools, vector graphics editing applications define graphic elements based on mathematical formulas, thus allowing for infinite scaling of graphic elements without loss of quality. They are suitable for creating high-quality icons, illustrations, logos, and typography designs. Therefore, vector graphics editing applications can be used to draw interface design drafts and design user interfaces (UI).

[0088] With the development of internet technology and the increasing participation of people in the internet, the amount of data in internet applications has grown massively. Data storage and transmission require suitable data carriers, among which PDF files are an important one. PDF files can serve as carriers for various forms of data, including text, images, icons, tables, and resource links. Because PDF documents are independent of various applications, hardware, and operating systems, data formats that require specific applications to open and operate can be converted to PDF for easier manipulation across different applications. For example, plain text academic papers can be converted to PDF documents, as can presentation documents (such as PowerPoint presentations).

[0089] As internet applications increasingly permeate all aspects of social life, people's demand for aesthetically pleasing, fully functional, and highly interactive internet application interfaces is becoming more urgent. Interface designers can use vector graphics editing applications to draw various components of the application interface, such as icons, controls, and text, one by one; however, this process is time-consuming. To improve interface development efficiency, developers can reuse existing design materials by converting them into editable components in vector graphics editing applications, and then fine-tune them, thereby significantly improving interface development efficiency. Design materials can be stored as PDF files. In this case, the PDF file needs to be converted to its original format, and then restored in the vector graphics editing application. The elements in the PDF file are then directly converted into components in the vector graphics editing application to improve application interface development efficiency.

[0090] The method for restoring PDF files in vector graphics editing applications in related technologies is as follows: The PDF document is parsed using a PDF parsing library to obtain element information of the PDF elements. This obtained element information can include element content information, element attribute information, and element position information. For example, for text elements, the parsed element information can include text content information, text attribute information (such as font, font size, and color), and text position information. Furthermore, components in the vector graphics editing application can be rendered one by one based on the obtained PDF elements and their element information. For example, please refer to... Figure 1 The method for converting PDF documents in related technologies includes: step 110, parsing portable document format data, and step 120, rendering vector graphics editing application components. After step 110, portable document format element data is obtained, including a portable document format element set 101. This set includes text element 1, text element 2, text element 3, and image element 1. Each PDF element is rendered individually to obtain the corresponding editable component. After step 120, a vector graphics editing application component set 102 is obtained, including text component 1 corresponding to text element 1, text component 2 corresponding to text element 2, text component 3 corresponding to text element 3, and image component 1 corresponding to image element 1.

[0091] like Figure 2As shown in the comparison image between the PDF restored version and the original PDF obtained according to the PDF format conversion method in related technologies, in the original PDF, all characters belonging to the same line are on the same horizontal line. However, in the restored PDF, multiple characters belonging to the same line do not appear on the same horizontal line, presenting an uneven display effect. Therefore, the method in related technologies has low accuracy in restoring PDF files in vector graphics editing applications, resulting in poor restoration effects. This is because PDF files determine the vertical coordinate of each character based on its height. However, even for character elements with the same font size, their heights differ, leading to inconsistent vertical coordinates for characters belonging to the same line. For example, the letters H and e, both with a font size of 4, have different heights and corresponding vertical coordinates; for instance, the vertical coordinate of letter H is 2, while that of letter e is 1. Thus, if the corresponding components in a vector graphics editing application are rendered directly based on the element position data parsed from the PDF file, inaccurate element position data will result in... Figure 2 The restoration effect in the image.

[0092] To address the issue of poor accuracy in processing PDF element position data (i.e., data processing) when restoring PDF files in vector graphics editing applications using related technologies, this disclosure provides a data processing method to improve the accuracy of element position data processing under the aforementioned conditions and achieve better PDF file restoration results.

[0093] System architecture and scenario description of the embodiments disclosed herein

[0094] Figure 3 This is a system architecture diagram for a data processing method according to an embodiment of the present disclosure. It includes a terminal 340, an Internet 330, a gateway 320, a server 310, etc.

[0095] Terminal 340 includes various device forms such as desktop computers, laptops, PDAs (personal digital assistants), mobile phones, in-vehicle terminals, home theater terminals, dedicated terminals, intelligent voice interaction devices, smart home appliances, or aircraft. Furthermore, it can be a single device or a collection of multiple devices. Terminal 340 can communicate with the Internet 330 via wired or wireless means to exchange data.

[0096] Server 310 refers to a computer system capable of providing certain services to terminal 340. Compared to ordinary terminal 340, server 310 has higher requirements in terms of stability, security, and performance. Server 310 can be a single high-performance computer in a network platform, a cluster of multiple high-performance computers, a portion of a single high-performance computer (e.g., a virtual machine), or a combination of portions of multiple high-performance computers (e.g., virtual machines). In this embodiment, server 310 specifically provides storage functionality; that is, server 310 can be a database node in a distributed database, and this disclosure includes a cluster of multiple servers 310 forming a distributed database.

[0097] Gateway 320, also known as an internetwork connector or protocol converter, is a computer system or device that acts as a translator, enabling network interconnection at the transport layer. It translates between two systems using different communication protocols, data formats, languages, or even completely different architectures. Gateways can also provide filtering and security functions. Messages sent from terminal 340 to server 310 are forwarded to the corresponding server 310 via gateway 320. Messages sent from server 310 to terminal 340 are also forwarded to the corresponding terminal 340 via gateway 320. In this embodiment, terminal 340 sends a data access request to server 310 via gateway 320, and server 310 returns the data access result to terminal 340 via gateway 320.

[0098] The data processing method provided in this embodiment can be implemented in terminal 340, server 310, or partly in terminal 340 and partly in server 310.

[0099] When the data processing method provided in this embodiment is implemented in terminal 340, terminal 340 acquires a portable document format file and parses the portable document format file to obtain character element information of multiple character elements in the portable document format file. The character element information includes first character position information and character attribute information. Then, terminal 340 detects the line association relationship between multiple character elements based on the character attribute information and adjusts the first character position information of multiple character elements according to the line association relationship to obtain second character position information corresponding to each character element information. Further, terminal 340 determines the paragraph relationship between multiple character elements based on the character attribute information, the second character position information, and a preset line spacing threshold, and adjusts the second character position information of multiple character elements according to the paragraph relationship to obtain third character position information corresponding to each character element information. Finally, terminal 340 generates interface design data corresponding to the portable document format file based on the third character position information and the character attribute information.

[0100] When the data processing method provided in this embodiment is implemented in server 310, server 310 obtains a portable document format file and parses the portable document format file to obtain character element information of multiple character elements in the portable document format file. The character element information includes first character position information and character attribute information. Then, server 310 detects the line association relationship between multiple character elements based on the character attribute information and adjusts the first character position information of multiple character elements according to the line association relationship to obtain second character position information corresponding to each character element information. Further, server 310 determines the paragraph relationship between multiple character elements based on character attribute information, second character position information, and a preset line spacing threshold, and adjusts the second character position information of multiple character elements according to the paragraph relationship to obtain third character position information corresponding to each character element information. Finally, server 310 generates interface design data corresponding to the portable document format file based on the third character position information and character attribute information.

[0101] When the data processing method provided in this embodiment is partially implemented in the terminal 340 and partially implemented in the server 310, the terminal 340 acquires a portable document format file and then sends the acquired portable document format file to the server 310. The server 310 then parses the portable document format file to obtain character element information for multiple character elements within the file. The character element information includes first character position information and character attribute information. Further, the server 310 detects the line association relationship between the multiple character elements based on the character attribute information and adjusts the first character position information of the multiple character elements according to the line association relationship to obtain second character position information corresponding to each character element. Based on the character attribute information, the second character position information, and a preset line spacing threshold, the server determines the paragraph relationship between the multiple character elements and adjusts the second character position information of the multiple character elements according to the paragraph relationship to obtain third character position information corresponding to each character element. Finally, the server 310 generates interface design data corresponding to the portable document format file based on the third character position information and the character attribute information, and then sends the generated interface design data to the terminal 340.

[0102] The data processing method provided in this disclosure can be applied to the task of restoring element data of PDF files in various vector graphics editing applications. Specifically, it can be applied to restoring PDF files of academic papers in vector graphics editing applications, as well as restoring PDF files of UI design drafts and presentation documents in vector graphics editing applications.

[0103] For example, when the data processing method provided in this embodiment is applied to reconstructing a PDF file of an academic paper in a vector graphics editing application, the academic paper PDF file can be obtained, and then the academic paper PDF file can be parsed to obtain character element information of multiple character elements in the academic paper PDF file. The character element information includes first position character information and character attribute information. Then, the line association relationship between multiple character elements can be detected based on the character attribute information, and the first character position information of multiple character elements can be adjusted according to the line association relationship to obtain the second character position information corresponding to each character element information. Further, the paragraph relationship between multiple character elements can be determined based on the character attribute information, the second character position information, and a preset line spacing threshold, and the second character position information of multiple character elements can be adjusted according to the paragraph relationship to obtain the third character position information corresponding to each character element information. The interface design data corresponding to the academic paper PDF file can be generated based on the third character position information and the character attribute information. In this way, the reconstruction of an academic paper PDF file can be realized in a vector graphics editing application.

[0104] For example, when the data processing method provided in this embodiment is applied to reconstructing a PDF file of an application interface design draft in a vector graphics editing application, the design draft PDF file can be obtained, and then the design draft PDF file can be parsed to obtain character element information of multiple character elements in the design draft PDF file. The character element information includes first position character information and character attribute information. Then, the line association relationship between multiple character elements can be detected based on the character attribute information, and the first character position information of multiple character elements can be adjusted according to the line association relationship to obtain the second character position information corresponding to each character element information. Further, the paragraph relationship between multiple character elements can be determined based on the character attribute information, the second character position information, and a preset line spacing threshold, and the second character position information of multiple character elements can be adjusted according to the paragraph relationship to obtain the third character position information corresponding to each character element information. The interface design data corresponding to the design draft PDF file can be generated based on the third character position information and the character attribute information. In this way, the design draft PDF file can be reconstructed in a vector graphics editing application.

[0105] The above examples do not limit the scope of protection in this case.

[0106] General Description of Embodiments in this Disclosure

[0107] According to one embodiment of this disclosure, a data processing method is provided. For example... Figure 4The diagram shown is a flowchart illustrating a data processing method provided in this disclosure. This method can be applied to a data processing apparatus, which can be integrated into a computer device, specifically a terminal or a server. The data processing method may include:

[0108] Step 410: Obtain the portable document format file and parse the portable document format file to obtain the character element information of multiple character elements in the portable document format file.

[0109] As described above, the data processing method provided in this disclosure can be specifically applied to data processing scenarios involving the restoration of PDF files in various vector graphics editing applications, including but not limited to restoring academic paper PDF files and interface design draft PDF files in various vector graphics editing applications. The portable document format file in this disclosure can be a PDF file of any content type, such as academic papers, interface design drafts, or interface design materials (e.g., icons and controls used in application interface design). The portable document format file in this disclosure can also be obtained by converting other document format files, such as image format files, plain text format (e.g., ".txt" format) files, presentation format (e.g., PPT) files, table (e.g., Excel) files, or drawing tool files.

[0110] In this embodiment of the disclosure, a vector graphics editing application plugin is provided, which allows users to select and upload PDF files that need to be converted on the plugin's plugin page. For example... Figure 5 As shown, the vector graphics editing application plugin interface 500 includes a PDF file selection area 510, a conversion control 520, a reference image page number setting area 530, and a reference image preview area 540. The PDF file selection area 510 includes a file selection control 511. In response to the triggering operation of the file selection control 511, a list of candidate PDF files can be displayed, allowing selection of the target PDF file for conversion. After selecting the target PDF file, its name is displayed in the PDF file selection area 510. In response to the triggering operation of the conversion control 520, conversion of the target PDF file can begin; in response to the triggering operation of the cancel conversion control, the conversion process can be terminated. The page number setting area 530 allows setting the page number range for the reference image, where the reference image is the original image of the PDF file, used to provide a reference for the reconstructed PDF file. Furthermore, after acquiring the PDF file, file parsing is performed on it.

[0111] In some embodiments, parsing a portable document format file to obtain character element information of multiple character elements in the portable document format file includes the following steps:

[0112] The portable document format file is parsed to obtain the first character element data in Extensible Markup Language format and the second character element data in Scalable Vector Graphics format.

[0113] The character element information of multiple character elements in a portable document format file is determined based on the first character element data and the second character element data.

[0114] Specifically, this embodiment of the disclosure uses the PyMuPDF library to parse PDF files, obtaining first character element data in Extensible Markup Language (XML) format and second character element data in Scalable Vector Graphics (SGT) format. PDF files have complex structures, and to ensure security, PDF creators may encrypt or obfuscate the files to prevent tampering. For such PDF files, the parsed data may contain garbled characters, obfuscated code, or special characters, such as those specified by "%" or "&". These garbled characters, obfuscated code, or special characters cannot be parsed, causing the PDF parsing process to interrupt, ultimately leading to parsing failure and termination of the PDF conversion process. This embodiment of the disclosure processes invalid character data obtained during PDF parsing, including replacement, character escaping, and filtering. Specifically, this embodiment can replace invalid XML characters obtained during parsing with spaces, recording the space positions, and replace "&" characters with "&". Character escaping ensures that the obtained characters can be correctly displayed in vector graphics editing applications, thus avoiding parsing failure and preventing invalid text information in the parsed data.

[0115] In the embodiments of the present disclosure, first character element data in extensible markup language format and second character element data in scalable vector graphics format are obtained by parsing a PDF file, and the aforementioned character escape processing is performed on the first character element data and the second character element data. The first character element data is XML-format data, which may specifically include character content information and character attribute information of text characters. For example, the character content information may include "Qinyuanchun Snow", and the character attribute information may include character style information such as font (fontFamily, style), font size and color, and may also include URL, position information of characters in the PDF file and space position information. The space position information may include the space position corresponding to an actual space in the PDF file and the space position corresponding to a space when an invalid XML character is replaced with a space. Recording space positions can ensure the integrity of file parsed data, and XML-format data can also store content information and position information corresponding to vector text.

[0116] In this embodiment, a text character is a text element that allows editing of a single character and modification of its style, while a vector text is a vector graphic composed of a plurality of drawing points forming the text outline. A vector text is displayed in the form of an image, and it is impossible to edit a single character or adjust its style therein, and only color filling can be performed on it. The second character element data is SVG-format data. Different from the first character element data, the first character element data contains detailed character element information but does not contain hierarchical relationship information between character elements, while SVG-format data is a tree-shaped data that can be used to store the hierarchical relationship between character elements in a PDF file, so that in the subsequent rendering process, rendering can be performed layer by layer according to the element hierarchical relationship, which avoids the problem of element misalignment or element occlusion existing in the related art. SVG-format data can be used to store content information and position information corresponding to text characters or vector text. For text characters, SVG-format data may further include a transformation matrix for controlling character deformation and character color. For vector text, SVG-format data may further include text attribute information of the vector text. The vector information of the vector text may include coordinate information corresponding to each contour point, and the text attribute information may include the filling color of the vector text. A code example of SVG-format data is as follows:

[0117] text="水" transform="matrix{a,b,c,d,e,f}" fill="#098245"

[0118] The `text` field stores the character content, which is "water". The transformation matrix for this character is "matrix{a,b,c,d,e,f}". In the SVG coordinate system, the `transform` field controls the scaling, rotation, and translation of the character. `a` and `d` control scaling, `b` and `c` control rotation, and `e` and `f` control translation. The `fill` field represents the character color, encoded as "#098245". In addition to the above information, the SVG format data may also include the `xlink:href` field, which defines the attribute of a resource link. This is typically used to reference external resources (such as images, icons, or other SVG files) within the SVG. In this embodiment, XML and SVG format data corresponding to each character element can be parsed. Furthermore, character element information for multiple character elements in the PDF file can be determined based on the XML and SVG format data.

[0119] In some embodiments, determining character element information of multiple character elements in a portable document format file based on first character element data and second character element data includes the following steps:

[0120] The text character elements and vector text elements in the first and second character element data are detected to obtain the detection results.

[0121] When the detection result indicates that both text character elements and vector text elements exist in either the first character element data or the second character element data, the vector text elements are filtered, and the character element information of multiple character elements in the portable document format file is determined based on the text character elements.

[0122] When the detection result indicates that only vector text elements exist in both the first and second character element data, the character element information of multiple character elements in the portable document format file is determined based on the vector text elements.

[0123] Specifically, in this embodiment of the disclosure, a PDF file is parsed to obtain SVG format data and XML format data. Since both SVG format data and XML format data can include text character elements and vector text elements, if both text character elements and vector text elements exist simultaneously, the position information of the two types of elements is the same. In order to avoid the problem of text overlap (ghosting) caused when vector text elements and text character elements exist and are displayed together, this embodiment of the disclosure detects the text character elements and vector text elements in the parsed XML format data and SVG format data.

[0124] In this embodiment, detection can be performed in the following ways: First, by detecting the display position information of text character elements and vector text elements, since the display position information of vector text elements and text character elements is the same; second, by determining the element identifier (ID), since the element identifiers of both are the same; third, by detecting the results returned from file parsing, where vector text elements and text character elements, if present simultaneously, are usually returned consecutively. The detection results are as follows... Figure 6 As shown, the detection results can be divided into three cases: The first case is that both XML and SVG format data contain only vector text elements and no text character elements; the second case is that either XML or SVG format data contains both vector text elements and text character elements; the third case is that both XML and SVG format data contain only text character elements. For the first case, vector text elements are retained, and the character element information of multiple character elements in the PDF file is determined based on the obtained vector text elements and their corresponding element information. This allows vector text to be displayed in vector graphics editing applications, avoiding blank spaces. For the second case, vector text elements can be filtered to retain text character elements, thus avoiding the problem of overlapping text display caused by the simultaneous presence of vector text and text character elements. For the third case, which is the most common scenario, text character elements can be directly retained and then processed.

[0125] Thus, this embodiment of the present disclosure detects vector text elements and text character elements in the data obtained by parsing the PDF file, and performs compatibility processing for different detection situations. This ensures that text content is displayed during the conversion process, while avoiding the ghosting problem caused by the simultaneous display of vector text elements and text character elements, thus ensuring a better restoration effect of the PDF file.

[0126] Furthermore, the PDF file parsing process is further processed through the aforementioned steps, including character replacement and character escaping, as well as filtering of the XML and SVG format data parsed from the file, to obtain character element information corresponding to multiple character elements in the PDF file. Specifically, this embodiment uses character elements as the smallest granularity, performing character escaping and special character replacement one by one to achieve reorganization and optimization of the PDF file parsing data, obtaining character element information corresponding to multiple character elements. Among these, the multiple character elements in the PDF file are text character elements. In cases where only vector text elements are parsed and not text character elements are parsed normally, the vector text elements and their corresponding element information are retained. This embodiment focuses on the case where only text character elements are parsed. The character element information corresponding to multiple character elements in a PDF file can include element content information, first character position information, and character attribute information. The element content information is as described above in the `text` field. The first character position information includes the drawing coordinates of the character element in the SVG coordinate system. This first character position information can be determined based on the coordinates of the SVG coordinate system origin and the aforementioned translation control parameters, specifically the translation control parameters in the transformation matrix. For example, in the `transform` matrix, if the translation parameters are 180 and 170, the first character element position information can be determined as (180, 170), in pixels. The character attribute information can include the character element's style attributes, element identifier, URL, and the aforementioned transformation matrix information. This embodiment of the present disclosure reorganizes and optimizes the data obtained from parsing the PDF file through the aforementioned steps, thereby improving the accuracy of the obtained character element information. Furthermore, this embodiment of the present disclosure simultaneously extracts XML and SVG format data. By combining XML and SVG format data, it can obtain both the hierarchical relationship between character elements and complete and detailed character element information for rendering, thus greatly improving the accuracy of restoring the PDF file in vector graphics editing applications.

[0127] Step 420: Detect the row association relationship between multiple character elements based on the character attribute information, and adjust the first character position information of multiple character elements according to the row association relationship to obtain the second character position information corresponding to each character element information.

[0128] In this embodiment of the disclosure, after obtaining the character element information corresponding to multiple character elements in a PDF file, line association detection can be performed on the multiple character elements. Due to the complexity of the PDF file structure itself, if there are differences in character styles, such as differences in font color or font style, when returning the parsed data of the PDF file, on the one hand, multiple text lines belonging to the same paragraph may be returned separately, and multiple characters belonging to the same text line may be returned as individual characters. The parsed data loses the association between character elements, resulting in characters that should belong to the same string being displayed separately and unable to display the correct semantics. For example, for the English string "abcbca", if it is split into scattered characters for display, such as being split into "a", "bcb", and "ca", these scattered characters or strings are displayed separately in the same line or scattered across different lines. The degree of aggregation of character elements is low, resulting in the directory nodes obtained in vector graphics editing applications being very scattered. For example, the directory node should be "abcbca", but according to related technologies, the corresponding components are rendered one by one according to the returned character elements, and the resulting directory nodes become "a", "bcb", and "ca". The usability of multiple scattered nodes is extremely low. On the other hand, as described above, due to differences in character height, if rendering is performed directly based on the parsed character position information, the characters belonging to the same line will not be displayed on the same horizontal line because the drawing coordinates are not on the same horizontal line. This results in staggered characters in the drawn text lines. To solve this problem, this embodiment uses a single character as the processing unit to detect the line association relationship between multiple character elements in the PDF file. This allows it to determine whether character elements can be merged into a line and adjust the position information of the first character of each character element in each line, so that character elements belonging to the same text line are displayed on the same horizontal line. This restores the display effect of the PDF file at the pixel level, achieving high fidelity.

[0129] In some embodiments, detecting the row association relationship between multiple character elements based on character attribute information includes the following steps:

[0130] Extract the character deformation parameters corresponding to each character element from the character attribute information, and calculate the rotation angle corresponding to each character element based on the character deformation parameters;

[0131] Detect the row association relationship between multiple character elements based on the rotation angle.

[0132] In this embodiment of the disclosure, the character deformation parameters corresponding to each character element can be extracted from the character attribute information. These character deformation parameters are the matrix parameters in the transform matrix described above. Furthermore, the rotation angle corresponding to each character element can be calculated based on the character deformation parameters. The rotation angle can include the rotation angle along the x-axis and the rotation angle along the y-axis. Please refer to... Figure 7 The rotation angle along the x-axis is α, and the rotation angle along the y-axis is β. As mentioned earlier, the matrix parameters in the transform matrix are specifically matrix{a,b,c,d,e,f}. The corresponding rotation angles can be calculated based on these matrix parameters. The formula for calculating the rotation angle α along the x-axis is as follows:

[0133]

[0134] Where b is the parameter in the transformation matrix used to control rotation, and a is the parameter in the transformation matrix used to control scaling. The formula for calculating the rotation angle β along the y-axis is as follows:

[0135]

[0136] Here, c is another parameter in the transformation matrix used to control rotation. Thus, the rotation angle corresponding to each character element can be calculated. It should be noted that the rotation angle calculated above is the rotation angle of the character element in the SVG coordinate system.

[0137] By calculating the rotation angle of character elements, it can be determined whether character elements in a PDF file have a slant style set. When rendering the component corresponding to the character element in a vector graphics editing application, the slant effect of the character element can be restored based on the rotation angle. Furthermore, the row association relationship between multiple character elements can be detected based on their corresponding rotation angles. Specifically, the row association relationship between two character elements can be detected by checking if any two characters in multiple character elements in the PDF file have the same rotation angle. This row association relationship detection step needs to be performed on each pair of multiple character elements in the PDF file.

[0138] In some embodiments, detecting the row association relationship between multiple character elements based on the rotation angle includes the following steps:

[0139] Extract the origin coordinates of each character element from the character attribute information, and calculate the slope angle between any two character elements among multiple character elements based on the origin coordinates.

[0140] Calculate the first distance value between any two character elements among multiple character elements based on the coordinates of the origin point.

[0141] Detecting the line association relationship between a plurality of character elements according to the slope angle, the first distance value and the rotation angle.

[0142] In the embodiments of the present disclosure, as introduced above, the character attribute information includes drawing coordinate information of character elements, thus the drawing origin coordinate corresponding to each character element can be extracted from the character attribute information. It should be noted that the drawing origin coordinate is the drawing origin coordinate of the character element in the SVG coordinate system. Further, the slope angle and the first distance value between any two character elements in the PDF file can be calculated, wherein the slope angle represents the included angle between the connecting line of the drawing origins corresponding to the two character elements and the x-axis, and the first distance value is the length value of the connecting line between the drawing origins corresponding to the two character elements. As introduced above, the line association relationship between a plurality of character elements can be detected through the rotation angle corresponding to the character elements, and further, the line association relationship between a plurality of character elements can also be detected through the slope angle and the first distance value. As Figure 8 shows, the distance value between the drawing origins of the character element "Wo (I)" and the character element "He (and)" is the first distance value, the rotation angle of the character element "Wo (I)" is α, and the slope angle therebetween is θ. At this time, α is exactly equal to θ, indicating that the two inclined character elements "Wo (I)" and "He (and)" are on the same straight line. If the first distance value therebetween satisfies the preset merging rule, these two character elements can be merged into one line.

[0143] In some embodiments, detecting the line association relationship between a plurality of character elements according to the slope angle, the first distance value and the rotation angle comprises the following steps:

[0144] determining a first angle difference of rotation angles corresponding to any two character elements, and determining a second angle difference between the rotation angle of each character element in any two character elements and the corresponding slope angle;

[0145] determining a first comparison result between the first angle difference and a preset first angle difference threshold, a second comparison result between the second angle difference and a preset second angle difference threshold, and a third comparison result between the first distance value and a preset first distance threshold;

[0146] detecting the line association relationship between the plurality of character elements based on the first comparison result, the second comparison result and the third comparison result.

[0147] In this embodiment, the rotation angle corresponding to each character element in the PDF file can be calculated first. The rotation angles of the character elements are then compared pairwise. A preset first angle difference threshold is used to determine whether the rotation angles of two character elements are the same or differ significantly. If the rotation angles of two character elements are the same (i.e., the first angle difference is zero), it indicates that they can be merged, and the next step of detection can proceed. If the first angle difference of the rotation angles of two character elements is not zero (i.e., the rotation angles are different), but the angle difference is still within the preset angle difference threshold range, the next step of detection can proceed. If the first angle difference is not within the angle difference threshold range, the next step of detection is unnecessary.

[0148] Furthermore, the second angle difference between the rotation angle and slope angle of each of any two character elements can be calculated. Alternatively, only the second angle difference between the rotation angle of one character element and the slope angle of the two characters can be calculated. Then, the second angle difference is compared with a preset second angle difference threshold, where the second angle difference threshold is a threshold range, such as [-1°, 1°]. If the second angle difference is within the second angle difference threshold range, the next step of row association detection can be performed. This embodiment of the present disclosure determines whether to merge two character elements by calculating the slope angle between two character elements and then using the difference between the rotation angle of one character element and the slope angle. For some character elements that are not on the same line but belong to the same area, this method can merge these characters together. For example, some character elements displayed according to an arc trajectory (e.g., character elements in a design draft), although the rotation angles of the character elements are not the same, if the difference between the rotation angle and slope angle of each character element remains within the threshold range, it indicates that these character elements, although not on the same line, may belong to the same area or trajectory, and it can be further determined whether to merge them.

[0149] Furthermore, a first distance value is compared between each pair of multiple character elements with a preset first distance threshold, where the first distance threshold is the Euclidean distance between the drawing origins of the two character elements. For example, the first distance threshold ranges from 10 to 20 pixels. If the first distance value between two character elements is within the first distance threshold range, the two character elements can be merged. When all three comparison results meet the line merging requirements, the character elements that meet the requirements can be merged into a line. It should be noted that the preset threshold can be dynamically adjusted according to the font size.

[0150] In this embodiment of the disclosure, detecting the line association relationship between multiple character elements in a PDF file may specifically include, for example: Figure 9The steps shown are as follows: Step 910, calculate the slope angle between two character elements based on the coordinates of the origin; Step 920, calculate the rotation angle between two character elements based on the character deformation parameters; Step 930, calculate the first distance value between two character elements based on the coordinates of the origin; Step 940, perform line association detection on the two character elements based on the rotation angle, slope angle, and first distance value. By performing line association detection on multiple character elements in the PDF file through the above steps, multiple character elements that can be merged into a single line are identified, and the line association relationships between multiple character elements are determined. The line association relationships can be represented by a list, with each line corresponding to a list containing multiple character elements in each line and the index position of each character element within its line. After determining the line association relationships between multiple character elements in the PDF file, the first character position information of the character elements can be further adjusted, including adjusting the vertical and horizontal coordinates of the characters in the first character position information.

[0151] In some embodiments, after detecting the row association relationship between multiple character elements based on the first comparison result, the second comparison result, and the third comparison result, the following steps are further included:

[0152] Extract the font information corresponding to the character element in each text line from the character attribute information, and extract the horizontal coordinate of the character element in each text line from the first character position information;

[0153] The font information, character x-coordinate, and rotation angle are used to merge multiple character elements in each text line into strings.

[0154] In this embodiment, after determining the row association, the character elements in each text line can be horizontally merged to ensure complete semantics. This guarantees that when rendering the components corresponding to the character elements in the vector graphics editing application, complete directory nodes can be generated, avoiding the problem of low node usability caused by splitting a directory node into multiple fragmented directory nodes. Specifically, the font information corresponding to the character elements in each text line can be extracted from the character attribute information. The font information can include the font size information corresponding to the character elements, such as the font size. Then, the horizontal coordinate of the character elements in each text line can be extracted from the first character position information. Starting from the first character element in each text line, the font size difference of the character elements is calculated sequentially. For example, the font size difference between the first and second character elements is calculated, then the font size difference between the second and third character elements is calculated, and so on, until the last character element in the list. The first character element in each text line can be the starting character element with an index position of 0 or 1 in the list. Furthermore, starting from the first character element in each text line, the difference in the horizontal coordinates of each character element can be calculated sequentially. For example, calculate the difference in the horizontal coordinates between the first and second character elements, then between the second and third character elements, and so on, until the last character element in the list. Further, the rotation angles of every two character elements in each text line can be determined using the aforementioned method. If the font size difference and horizontal coordinate difference determined through the aforementioned steps are compared with preset thresholds, and both are within the threshold range and the rotation angles of the two character elements are the same, they can be merged into a single string. For example, for the character elements "A", "B", and "C", satisfying the aforementioned three conditions, they can be merged into the string "ABC". It should be noted that the preset thresholds can be dynamically adjusted according to the font size.

[0155] In other embodiments, the vertical coordinate of each character element can be extracted from the first character position information. Based on the aforementioned three conditions, the difference in the vertical coordinates of the character elements is compared with a preset threshold. The result of this comparison further determines whether to merge the character elements in each row into a string. This step is performed after determining the row relationships but before adjusting the vertical coordinates of the character elements in each row. If the vertical coordinates of the character elements in each row have been adjusted according to the row relationships, this step is unnecessary; the text can be horizontally merged based on the aforementioned three conditions.

[0156] In other embodiments, if the PDF file contains paragraph numbers or table of contents numbers, a mapping relationship can be established between the paragraph number or table of contents number and the first character element in the corresponding text. Specifically, a mapping relationship can be established between the paragraph number and the character element index. After determining the line association relationship, the paragraph number and the corresponding text can be horizontally merged according to the mapping relationship to make it a complete line.

[0157] This embodiment of the disclosure improves the semantic correlation between character elements by judging the correlation between character elements in each text line and merging character elements that should belong to the same string horizontally, thereby improving the degree of aggregation of character elements and ultimately improving the restoration effect of PDF files.

[0158] In some embodiments, adjusting the first character position information of multiple character elements according to the row association relationship to obtain the second character position information corresponding to each character element information includes the following steps:

[0159] Multiple character elements are divided according to the line association relationship to obtain multiple text lines, and the reference vertical coordinate of the text line is determined based on the position information of the first character of the character element in each text line.

[0160] Adjust the character ordinate corresponding to the first character position information of each character element in the text line based on the baseline ordinate to obtain the second character position information corresponding to each character element.

[0161] In this embodiment, multiple character elements in a PDF file can be divided into multiple text lines based on line association relationships. Each text line can be represented by a list, which may include the index position of the character element within the text line. Further, the reference ordinate of the text line can be determined based on the first character position information of the character elements in each text line. Specifically, the reference ordinate can be the ordinate of the first character element in each text line, or it can be the median or average of the ordinates of all character elements in the text line. Then, the ordinate of each character element in the text line can be adjusted based on the determined reference ordinate, and the first character position information is updated based on the adjusted ordinate to obtain the second character position information corresponding to each character element.

[0162] In some embodiments, as described above, the rotation angle of the character element and the position information of the first character are determined based on the transformation matrix in the SVG format data, that is, determined in the SVG coordinate system. Therefore, after adjusting the character ordinate corresponding to the position information of the first character element in the text line according to the reference ordinate, the following steps are also included:

[0163] Based on the rotation angle corresponding to each character element and the adjusted position information of the first character, coordinate transformation is performed to obtain the rotation angle and the position information of the first character in the second coordinate system.

[0164] The second character position information corresponding to each character element information is determined based on the rotation angle in the second coordinate system and the first character position information.

[0165] In the embodiments disclosed herein, such as Figure 10 As shown, the SVG coordinate system is not the same as the coordinate system in vector graphics editing applications. The left side is the SVG coordinate system, and the right side is the coordinate system in the vector graphics editing application. The SVG coordinate system uses the top-left corner of the character element's reflection as the origin, while the vector graphics editing application uses the top-left corner of the text box as the origin. Therefore, the transform matrix cannot be directly applied to the vector graphics editing application. First, the origin coordinates and rotation angle of the character element in the first coordinate system (i.e., the SVG coordinate system) can be obtained from the transform matrix. For example, the transform matrix is ​​as follows:

[0166] transform=matrix(36,-10,-10,-36,50,50)

[0167] Thus, the coordinates of the drawing origin in the SVG coordinate system can be calculated as (50, 50), and the rotation angle of the character element can be calculated using the following formula:

[0168]

[0169] The character element is rotated by -15° in the SVG coordinate system, such as... Figure 11 As shown, the original character is vertically flipped and then rotated 15° to obtain the transformed character. Since the origin coordinates in the SVG coordinate system are the same as those in the vector graphics editing application, the origin coordinates in the vector graphics editing application are also (50, 50). The rotation angle of the character element in the vector graphics editing application and the corresponding rotation angle of the character element in the SVG coordinate system are opposites. Therefore, the rotation angle of the character element in the vector graphics editing application is 15°. Thus, the second character position information corresponding to each character element can be determined based on the rotation angle in the second coordinate system (i.e., the vector graphics editing application) and the first character position information.

[0170] Step 430: Determine the paragraph relationship between multiple character elements based on character attribute information, second character position information, and a preset line spacing threshold, and adjust the second character position information of multiple character elements according to the paragraph relationship to obtain the third character position information corresponding to each character element information.

[0171] After determining the line relationships between multiple character elements in the PDF file through the aforementioned steps, the vertical coordinates of the characters in the first character position information of the multiple character elements can be adjusted according to the line relationships. Alternatively, the horizontal coordinates of the characters in the first character position information can be adjusted by horizontally merging each text line to obtain the second character position information. Furthermore, paragraph relationships can be detected and paragraphs can be merged from the multiple text lines determined based on the line relationships.

[0172] In some embodiments, determining the paragraph relationship between multiple character elements based on character attribute information, second character position information, and a preset line spacing threshold includes:

[0173] The line spacing between any two text lines is determined based on the position information of the second character;

[0174] Extract the font information corresponding to the character element from the character attribute information, and determine the paragraph relationship between multiple character elements based on the line spacing, font information, and preset line spacing threshold.

[0175] In this embodiment, since the vertical coordinates of the characters in the second character position information corresponding to the character elements in each text line are the same, the vertical coordinate of each text line can be determined based on the vertical coordinates of the characters in the second character position information, and the font information corresponding to each character element can be extracted from the character attribute information, wherein the font information may include font color, font style, and font size. The line spacing between any two text lines is calculated based on the display vertical coordinate corresponding to each text line.

[0176] Furthermore, the paragraph relationship between character elements in a PDF file can be determined based on the line spacing and font information. Specifically, the line spacing can be compared with a preset line spacing threshold to detect whether to merge two text lines into a paragraph. The first few text lines in the PDF file can be merged into paragraphs based on the preset line spacing threshold. These first few text lines can be determined using the index and coordinates of each text line; for example, a text line with index 0 is the first line, and one with index 1 is the second line. After merging the first few text lines into paragraphs, the coordinates of each subsequent text line are obtained sequentially. The line spacing between this text line and the last text line of the merged paragraph is compared with the preset line spacing threshold. If the line spacing exceeds the threshold, the paragraphs are not merged; otherwise, they are merged into a single paragraph. If the font size, font color, and font style of the character elements in two text lines differ significantly, the preset line spacing threshold can be reduced to minimize the merging of these two text lines. The preset line spacing threshold can be dynamically adjusted based on font size, font color, and font style information.

[0177] In this embodiment of the disclosure, as described above, each text paragraph may contain a paragraph number or a table of contents number. When performing paragraph relationship detection, text lines that begin with a paragraph number or a table of contents number are not merged with previous text lines. This is because the paragraph number or paragraph number is used to identify a new text paragraph, that is, the beginning of a text paragraph.

[0178] Thus, this embodiment of the present disclosure, based on merging lines of character elements in a PDF file, further merges paragraphs between multiple text lines, improving the aggregation degree of multiple character elements in the PDF file, achieving a better PDF file restoration effect, and making the generated directory nodes in vector graphics editing applications more singular and more usable. Furthermore, after determining the paragraph relationships between multiple text lines, the line spacing of each text line within a paragraph and the spacing between paragraphs can also be adjusted.

[0179] In some embodiments, the second character position information of multiple character elements is adjusted according to paragraph relationships to obtain the third character position information corresponding to each character element information, including:

[0180] Divide multiple lines of text into paragraphs based on paragraph relationships to obtain multiple text paragraphs;

[0181] Based on the preset reference line spacing and the line spacing between adjacent text lines in each text paragraph, the character ordinates in the second character position information of multiple character elements are adjusted to obtain the third character position information corresponding to each character element.

[0182] In this embodiment of the disclosure, the multiple text lines determined in the aforementioned steps can be divided into multiple text paragraphs according to paragraph relationships. First, the line spacing between the corresponding text lines of the first text paragraph can be adjusted. For example, for the first paragraph, the line spacing of the first few text lines in the paragraph can be adjusted based on a preset reference line spacing. After determining the line spacing of the first few lines, the remaining text lines in the paragraph can be adjusted sequentially with reference to the line spacing of the first few lines and the line spacing between adjacent text lines. If the line spacing is equal to the reference line spacing, there is no need to adjust the character ordinate; if not, the character ordinate in the second character position information of multiple character elements is adjusted. Thus, the line spacing between text lines in each paragraph is determined sequentially, and the character ordinate in the second character position information of each character element is adjusted according to the adjusted line spacing to obtain the third character position information corresponding to each character element.

[0183] In some embodiments, after determining the paragraph relationship between multiple character elements based on character attribute information, second character position information, and a preset line spacing threshold, the method further includes:

[0184] Obtain the second and third distance values ​​between the boundary of each text paragraph and the boundary of each text line within the text paragraph;

[0185] Get the fourth distance value between the midline of each text paragraph and the midline of each text line in the text paragraph;

[0186] The second, third, and fourth distance values ​​are compared with preset distance thresholds, and the alignment of each text paragraph is determined based on the three comparison results.

[0187] In this embodiment of the disclosure, the alignment of each text paragraph can be determined, such as... Figure 12 As shown, a second and third distance value can be obtained between the boundary of each text paragraph and the boundary of each text line within the text paragraph. The second distance value indicates the distance between the left side of the text line and the left side of the text paragraph, and the third distance value indicates the distance between the right side of the text line and the right side of the target merged paragraph. Figure 12 As shown, the left coordinate value (here, the horizontal coordinate) of the first line of text is L1, and the right coordinate value is R1. The left coordinate value of the second line of text is L2, and the right coordinate value is R2. The left coordinate value of the third line of text is L3, and the right coordinate value is R3. The left coordinate value of the text paragraph (let's assume it's represented by PL) is the same as the left coordinate value of the first line of text. The right coordinate value of the text paragraph (let's assume it's represented by PR) is the same as the right coordinate value of the second line of text. The coordinate value of the midline of the text paragraph (let's assume it's represented by PM) is the same as the midpoint coordinate value of L1 and R2.

[0188] Furthermore, the second and third distance values ​​for each text line can be calculated. For example, the second distance value of the first text line can be calculated using the difference between L1 and PL, the second distance value of the second text line can be calculated using the difference between L2 and PL, and the second distance value of the third text line can be calculated using the difference between L3 and PL. The third and fourth distance values ​​are calculated similarly. After calculating the second, third, and fourth distance values, they are compared with a threshold according to the priority of the paragraph alignment. For example, if the paragraph alignment priority is center alignment > left alignment > right alignment, then the fourth distance value can be compared with a preset distance threshold first. If it does not meet the threshold requirement, then the second distance value is compared with the preset threshold. It should be noted that the preset threshold can be dynamically adjusted based on information such as the font size and font style of the character elements. After determining the alignment method, the position information of the character elements can be fine-tuned to obtain the third character position information.

[0189] Step 440: Generate interface design data corresponding to the portable document format file based on the third character position information and character attribute information.

[0190] This disclosure improves the aggregation of character elements in PDF files by performing coordinate transformation, rotation angle processing, line merging, and paragraph merging, resulting in better restoration effects. The generated directory nodes in vector graphics editing applications are also more streamlined and usable. Please refer to... Figure 13 Based on the original PDF document 1310, the conversion process yields a PDF restoration document with corresponding directory nodes, as shown in the first directory node display area 1320. The string "SONNENESTRASSE" in the original PDF document 1310 is split into multiple fragmented directory nodes "SONNENESTR", "A", "S", and "SE", and the string "EINGANG" is split into multiple fragmented directory nodes "EING" and "ANG". However, the directory nodes obtained through this embodiment are shown in the second directory node display area 1330, where "SONNENESTRASSE" and "EINGANG" each correspond to a complete node, resulting in higher node usability.

[0191] After determining the position information of the third character corresponding to each character element, further, interface design data corresponding to the portable document format file can be generated based on the third character position information and character attribute information. This interface design data is obtained by reorganizing and optimizing the data parsed from the PDF file as described above. The interface design data is a DSL structure tree data used to render components in vector graphics editing applications. It can be used to reproduce design drafts, academic papers, and presentations in vector graphics editing applications. After the server generates the interface design data, it can be sent back to the vector graphics editing application plugin for rendering to obtain the final PDF reproduction.

[0192] The data processing method provided in this embodiment obtains a portable document format file and parses the file to obtain character element information of multiple character elements in the portable document format file. The character element information includes first character position information and character attribute information. The method detects line relationships between the multiple character elements based on the character attribute information and adjusts the first character position information of the multiple character elements according to the line relationships to obtain second character position information corresponding to each character element. Based on the character attribute information, the second character position information, and a preset line spacing threshold, the method determines the paragraph relationship between the multiple character elements and adjusts the second character position information of the multiple character elements according to the paragraph relationship to obtain third character position information corresponding to each character element. Finally, the method generates interface design data corresponding to the portable document format file based on the third character position information and the character attribute information.

[0193] This disclosure improves the accuracy of character position information by parsing the character attribute information of character elements in a portable document format file to detect the line association relationship between multiple character elements. Then, based on the line association relationship, the first character position information obtained from the file parsing is adjusted. Furthermore, the adjusted first character position information is further adjusted by determining the paragraph relationship between multiple character elements, thereby improving the accuracy of the character position information. Compared to related technologies that directly generate interface design data from the character position information obtained from file parsing, this disclosure improves the accuracy of processing character position information by adjusting the parsed character position information multiple times, ultimately improving the accuracy of the generated interface design data.

[0194] This disclosure provides a detailed description of embodiments in conjunction with specific application scenarios.

[0195] like Figure 14 The diagram shown is another flowchart illustrating the data processing method provided in this disclosure. This embodiment will use the restoration of an academic paper PDF file in a vector graphics editing application as an example to describe the data processing method in detail, focusing on the executing entities of each step. The method specifically includes the following steps:

[0196] Step 1401: The terminal obtains a portable document format file and sends the portable document format file to the server.

[0197] In this embodiment, the data processing method provided by this disclosure will be described in detail using the restoration of an academic paper PDF file in a vector graphics editing application as an example. It should be noted that this embodiment mainly addresses the technical problem of low accuracy in text position data processing when restoring PDF files in vector graphics editing applications. Using a PDF file whose main content is text as an example better demonstrates the beneficial effects of this embodiment. Those skilled in the art will understand that the method provided by this embodiment is still applicable to PDF files of design drafts or presentations. In this embodiment, the terminal can be any one of a personal computer, mobile device, wearable smart device, in-vehicle device, or server device.

[0198] To improve the accuracy of restoring PDF files in vector graphics editing applications, specifically by improving the accuracy of element position data processing during the PDF file restoration process, this disclosure provides a data processing method. Please refer to... Figure 15 The method mainly includes: the user uploads a portable document format file from the terminal; the server downloads the portable document format file; the server parses the portable document format file; further, the server reorganizes and optimizes the parsed data to generate interface design data, and sends the generated interface design data to the terminal. Then, the terminal parses the interface design data, renders components based on the interface design data, and finally performs rendering optimization. Each step will be described in detail later.

[0199] In this embodiment of the disclosure, the object can be accessed through, for example... Figure 5 The plugin interface shown above allows you to upload the PDF file that needs to be converted. The terminal then sends the uploaded PDF file to the server.

[0200] Step 1402: The server parses the portable document format file to obtain Extensible Markup Language (XML) format data and Scalable Vector Graphics (SGT) format data.

[0201] Furthermore, the server parses the portable document format file to obtain XML and SVG format data. To ensure document security, PDF file creators encrypt or obfuscate the PDF file to prevent tampering. Parsing such PDF files may result in garbled or obfuscated text information, lacking the original text. The parsed data may contain some corrupted data, therefore special processing is required. This can include replacement, escaping, and filtering. Specifically, this includes replacing invalid XML characters, replacing "&" characters with "&", character escaping, and data filtering to optimize the parsing process. This can prevent PDF file parsing failures or invalid data.

[0202] This embodiment of the disclosure uses the Pymupdf library to parse PDF files, enabling the separate acquisition of XML format data for text and image elements. Since XML format data cannot understand the hierarchical relationships between elements, it can cause element rendering overlay issues. This embodiment of the disclosure uses SVG format data to obtain the hierarchical order of PDF elements. However, SVG format data lacks some detailed attributes of elements, such as the font style of text characters. Therefore, this embodiment of the disclosure combines XML and SVG format data to obtain the hierarchical relationships between elements during rendering, while ensuring sufficient information to handle the details of the element rendering process.

[0203] In this embodiment of the disclosure, compatibility processing is also performed for XML format data and SVG format data in scenarios where there is no valid text. The case of no valid text means that the parsed data does not contain text characters, only vector text, in which case the vector text can be displayed. Please refer to... Figure 6 In addition to the case where there is no valid text, this embodiment also handles the case where both vector text and text characters exist simultaneously. In this case, vector text can be filtered out while text characters are retained, avoiding text overlap caused by displaying both at the same time. Another case is where only text characters exist, which is the most common case, and the text characters can be processed directly.

[0204] Furthermore, after the server performs character escaping, replacement, and filtering on the parsed data of the PDF file through the aforementioned steps, it can extract the font style information and space positions of character elements from the XML format data (space padding is needed for subsequent rendering optimization), and obtain the vector information (text vector) and position information (text attributes) of character elements from the SVG format data. Please refer to... Figure 16The PDF file is parsed to obtain XML and SVG data. These data are then processed separately. Specifically, invalid XML characters are replaced and escaped in the XML data, while character escaping is performed on the SVG data. This results in processed XML and SVG data, which are more accurate. Furthermore, font styles and space positions can be extracted from the XML data, and text vectors and text attributes can be extracted from the SVG data.

[0205] This embodiment of the disclosure performs the above processing on the parsed data of the PDF file, which improves the accuracy of the parsed data and avoids the technical problem in the related technology that the parsed data contains special characters that cause parsing failure. At the same time, it solves the problem in the related technology that the parsed data contains garbled characters, obfuscated characters or escape characters, which causes the PDF restoration effect to be completely different from the display effect of the original PDF. It also solves the problems of double text overlay display and element misalignment in the related technology.

[0206] Step 1403: The server determines multiple character elements and their corresponding character attribute information in the portable document format file based on Extensible Markup Language (XML) format data and Scalable Vector Graphics (SGT) format data.

[0207] In this embodiment of the disclosure, the server identifies the text characters and vector text in the parsed PDF file data, referring to... Figure 6 The processing method for different text scenarios involves identifying multiple character elements in the PDF file based on the extracted XML and SVG format data, and then performing subsequent processing at the granularity of individual character elements. The server can extract character attribute information from the XML and SVG format data, such as extracting font style information from the XML format data and text vector information from the SVG format data.

[0208] Step 1404: The server extracts the transformation matrix information corresponding to each character element from the scalable vector graphics format data, and calculates the rotation angle corresponding to each character element, the slope angle between any two character elements, and the distance from the drawing origin based on the transformation matrix information.

[0209] In this embodiment, the server can extract the transformation matrix information corresponding to each character element from the SVG format data, where the transformation matrix information is the information corresponding to the aforementioned transform field. Further, the server can calculate the rotation angle corresponding to each character element based on the transformation matrix information, obtaining the rotation angle of the character element in the SVG coordinate system. If the rotation angle is zero, it indicates that the character element has not set a tilt style, and no rotation angle conversion is needed. If the rotation angle is not zero, for example, a rotation angle of -15°, it indicates that the character element has set a tilt style, and in this case, coordinate system conversion is required to obtain the rotation angle of the character element in the vector graphics editing application. Specifically, the rotation angle of the character element in the vector graphics editing application is determined based on the negative of the rotation angle in the SVG coordinate system. Thus, this embodiment can correctly handle the rotation effect of character elements when restoring PDF files.

[0210] Furthermore, the server can determine the coordinates of the drawing origin point corresponding to each character element based on the transformation matrix information, and calculate the distance between the drawing origin points of any two character elements and the slope angle of the line connecting the drawing origin points of the two character elements based on the drawing origin point coordinates. The specific calculation method has been introduced in the preceding steps and will not be repeated here.

[0211] Step 1405: The server detects the row association relationship of multiple character elements based on the rotation angle, slope angle and distance from the drawing origin, and determines multiple text lines based on the determined row association relationship.

[0212] Due to the inherent complexity of PDF structure, even relevant parsed data may not be returned in the same parsing code area. This is especially true when character elements on the same line have different font colors or styles, often resulting in the parsed data being returned separately. This causes text in the same paragraph to be split into multiple text lines, or even text within the same line to be split into individual character elements. Furthermore, if the drawing origin coordinates of each character element are not on the same horizontal line, the resulting text on the same line will appear uneven. To address these issues, this embodiment processes PDF files at the individual character element level, performing special character escaping, coordinate transformation, and other processing on each character element. It also combines font style information to determine the correlation between character elements, merging multiple character elements into text lines, and further merging multiple text lines into text paragraphs.

[0213] The overall text optimization process is as follows: The server can first process the character element information in the PDF file, and then perform coordinate transformation on the coordinate position of the character elements; further, it can determine whether the character elements are on the same line based on the character element information, and if the line merging requirement is met, the character elements are merged into a text line; finally, multiple text lines are merged into paragraphs.

[0214] In this embodiment, the server can determine whether to merge two character elements into a single text line based on the rotation angle of each character element, the slope angle between two character elements, and the corresponding distance between the drawing origins. This involves detecting the line association relationships between multiple character elements in a PDF file. The specific line merging rules are as follows: The server first calculates the difference in rotation angles between two character elements and compares this difference with a preset threshold. Then, it compares the distance between the drawing origins of the two character elements with a preset distance threshold range. Further, it calculates the difference between the rotation angle of each character element and the slope angle of the line connecting the two drawing origins, and compares this angle difference with a preset threshold. If all three threshold comparison results meet the conditions, the two character elements involved can be merged.

[0215] Since character elements within the same text line may have different styles, merging character elements into a single text line requires additional processing of their font style information. After detecting line relationships, the server can establish a mapping table that includes the mapping relationship between character position information and font style information, such as `<position, style>`. Here, the character position information represents the index position of the character element within its text line, for example, index 0. Specifically, the server can first record the character style of the first character element in the text line in the mapping table. During text line merging, the server can extract the font style information corresponding to each character element from the XML format data. If a different font style is detected between the current character element and the previous character element within the same text line, the font style of that character element can be recorded in the mapping table so that text line attributes can be set for the character element when rendering the corresponding component. Please refer to [link to relevant documentation]. Figure 17,"A thousand miles of northern land, wrapped in ice, a thousand miles of snow drifting", the corresponding individual character elements are merged into a text line, and the font style corresponding to the first character element in the text line is recorded. For example, the font style corresponding to the character element at index 0 in the text line is regular, the font style of the character element "国" (country) at index 1 is the same as that of "北" (north), so rendering is performed with reference to the font style of the previous character element, the font style of the character element "风" (wind) at index 2 is set to italic, which can be recorded as <index 2, italic> in the mapping table, the font style of the character element "光" (light) is the same as that of "风" (wind), so rendering is performed with reference to its style. In this way, the font styles in the text line can be sequentially recorded in the mapping table, specifically including <index 5, light gray>, <index 7, bold>, <index 10, font size>, <index 12, dark gray>. In this way, the embodiments of the present disclosure can better restore the PDF file by setting the font style of individual character elements.

[0216] In step 1406, the server adjusts the character ordinate in the first character position information of the character elements according to the line association relationship, and obtains the second character position information.

[0217] In the embodiments of the present disclosure, the server may create a mapping relationship list of each text line and the character elements in the text line according to the line association relationship, which may specifically include information such as the index position of the character element in the text line, font style, rotation angle, and character position. The reference ordinate of each text line may be determined first, then the character ordinate in the first character position information of each character element is updated according to the reference ordinate to obtain the second character position information of each character element, and the updated second character position information is recorded in the mapping relationship list of the text line and the character elements.

[0218] Furthermore, after determining the line relationships, the server can also perform horizontal merging of character elements in text lines, adjusting the horizontal coordinates of character elements in each text line. Specifically, character elements belonging to the same string can be merged. The general merging rules are as follows: character elements with significantly different font sizes are not merged; distance values ​​or y-differences exceeding a threshold range are not merged; and the first text is not a directory number. Horizontal merging can also include merging paragraph numbers and body text. Since the directory or number of paragraphs is usually some distance from the body text, it cannot be merged using the above character element horizontal merging rules. For merging paragraph numbers and body text, the merging rules are as follows: the font size difference between the two texts (a character element in the same line and the character element preceding it) is within a certain threshold; the x-difference and y-difference between the two texts are within a certain threshold; the rotation angles of the two texts are the same; and the first text is within the directory number map mapping. Character elements that meet the above requirements can be merged horizontally. Thus, the embodiments of this disclosure can handle the special case of horizontal merging of paragraph numbers and body text, and can merge character elements that originally belong to the same string, so that the semantics of the string are complete, and ultimately achieve a better PDF file restoration effect.

[0219] Step 1407: The server obtains the line spacing between any two text lines and determines the paragraph relationship of multiple character elements based on font style information, line spacing, and a preset line spacing threshold.

[0220] Furthermore, after dividing multiple character elements of a PDF file into multiple text lines based on line relationships, the server can also perform paragraph merging on these text lines. The server can set a line spacing threshold, obtain the line spacing between any two text lines, and retrieve the font style information corresponding to each character element from the XML format data. It then determines whether paragraph merging is necessary based on the following rules: if the font color or font style is different, or the font size differs significantly, the line spacing threshold is reduced to minimize merging; if the line spacing exceeds the preset threshold, merging is not performed; if the paragraph number indicates the start of a paragraph, text lines starting with a paragraph number are not merged with the preceding text lines. When merging paragraph numbers with the main text, space positions can be extracted from the XML data and used to fill in the spaces between the paragraph numbers and the main text for better restoration.

[0221] Step 1408: The server merges multiple text lines into paragraphs based on paragraph relationships to obtain multiple text paragraphs, and adjusts the character ordinates in the second character position information to obtain the third character position information.

[0222] After determining the paragraph relationships among multiple character elements, the server can merge multiple text lines into multiple text paragraphs based on these relationships. Similarly, a mapping table can be established between text paragraphs and their constituent text lines and character elements. Furthermore, the vertical coordinate of the character in the second character position information of a character element can be adjusted to update the third character position information, and the updated character position information can be recorded in a list.

[0223] In other embodiments, after dividing multiple text lines into multiple paragraphs, the alignment of each paragraph can be judged and adjusted. The specific paragraph alignment judgment rules are as follows: calculate whether the distance from the left side of each line to the left side of the paragraph is within the left alignment threshold range; calculate whether the distance from the right side of each line to the right side of the paragraph is within the right alignment threshold range; calculate whether the distance from the center line of each line to the center line of the paragraph is within the center alignment threshold range; alignment priority: center alignment > left alignment > right alignment.

[0224] Step 1409: The server generates interface design data based on the third character position information and character attribute information of the character element, and sends the interface design data to the terminal.

[0225] In this embodiment, the server can reorganize data according to the character element-to-text line mapping list, the text line-to-paragraph mapping list, and the preset DSL data structure determined in the aforementioned steps to generate the final interface design data. This data may include the third character position information of each character element and the character attribute information of each character element. The character attribute information may include vector information of vector text or font style information and rotation angle of text characters, etc. After generating the interface design data, the server sends the interface design data to the terminal for component rendering.

[0226] Step 1410: The terminal renders components based on the interface design data and the preset font library to obtain the interface design draft corresponding to the portable document format file.

[0227] Since PDF files may be generated by various software programs, using a variety of fonts and font styles, and vector graphics editing applications typically support a limited number of fonts and font styles, compatibility processing is necessary during rendering to avoid rendering failures. This embodiment employs a mapping table approach. The terminal searches for a similar font library in the vector graphics editing application for rendering. If no matching font is found, a default font, such as "Inter", "Times New Roman", or "Arial", is used as a substitute. After the font library used changes, the font display effect also differs from the original PDF, exhibiting issues such as inconsistent font sizes, failure to fill text lines, or exceeding text line boundaries. This embodiment optimizes the font size, character spacing, and line spacing of character elements after font replacement. The specific optimization steps are as follows: First, the terminal checks whether the font size is too large or too small based on the width, height, and original font size of the text box. Then, it sets a font size range and iterates within this range using a binary search until a suitable font size is found that meets the width and height requirements of the drawn text box. Next, it recalculates and adjusts the character spacing using the same method to ensure that the character elements fill the text line. Finally, the terminal adjusts the line spacing using the same method based on the paragraph box height and the current drawing height. By performing rendering optimization through these steps during the component rendering process, even if a suitable font style is not matched, the terminal can still reproduce the PDF file to the greatest extent possible, obtaining the corresponding interface design draft.

[0228] Description of apparatus and devices according to embodiments of this disclosure

[0229] It is understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated in this embodiment, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.

[0230] It should be noted that in the various specific embodiments of this disclosure, when processing is required based on data related to the characteristics of the target object, such as target object attribute information or a set of attribute information, the permission or consent of the target object will be obtained first. Furthermore, the collection, use, and processing of this data will comply with the relevant laws, regulations, and standards of the relevant regions. In addition, when this application embodiment needs to obtain target object attribute information, separate permission or consent from the target object will be obtained through pop-up windows or redirection to a confirmation page. Only after obtaining the target object's separate permission or consent will the necessary target object-related data for the normal operation of this application embodiment be obtained.

[0231] Figure 18 A schematic diagram of the structure of a data processing apparatus 1800 provided in an embodiment of this disclosure. The apparatus includes:

[0232] The acquisition unit 1810 is used to acquire a portable document format file and parse the portable document format file to obtain character element information of multiple character elements in the portable document format file. The character element information includes the first character position information and character attribute information.

[0233] The first adjustment unit 1820 is used to detect the row association relationship between multiple character elements according to the character attribute information, and adjust the first character position information of multiple character elements according to the row association relationship to obtain the second character position information corresponding to each character element information;

[0234] The second adjustment unit 1830 is used to determine the paragraph relationship between multiple character elements based on character attribute information, second character position information and preset line spacing threshold, and to adjust the second character position information of multiple character elements according to the paragraph relationship to obtain the third character position information corresponding to each character element information.

[0235] The generation unit 1840 is used to generate interface design data corresponding to a portable document format file based on the third character position information and character attribute information.

[0236] Optionally, in some embodiments, the first adjustment unit includes:

[0237] The calculation subunit is used to extract the character deformation parameters corresponding to each character element from the character attribute information, and calculate the rotation angle corresponding to each character element based on the character deformation parameters.

[0238] The detection subunit is used to detect the row association relationship between multiple character elements based on the rotation angle.

[0239] Optionally, in some embodiments, the detection subunit includes:

[0240] The first calculation module is used to extract the coordinates of the drawing origin corresponding to each character element from the character attribute information, and calculate the slope angle between any two character elements among multiple character elements based on the coordinates of the drawing origin.

[0241] The second calculation module is used to calculate the first distance value between any two character elements among multiple character elements based on the coordinates of the drawing origin.

[0242] The first detection module is used to detect the row association relationship between multiple character elements based on the slope angle, the first distance value, and the rotation angle.

[0243] Optionally, in some embodiments, the first detection module includes:

[0244] The first determining submodule is used to determine the first angle difference between the rotation angles of any two character elements, and to determine the second angle difference between the rotation angle of each character element and the corresponding slope angle in any two character elements.

[0245] The second determining submodule is used to determine a first comparison result between the first angle difference and a preset first angle difference threshold, a second comparison result between the second angle difference and a preset second angle difference threshold, and a third comparison result between the first distance value and a preset first distance threshold.

[0246] The detection submodule is used to detect the row association relationship between multiple character elements based on the first comparison result, the second comparison result, and the third comparison result.

[0247] Optionally, in some embodiments, the first adjustment unit includes:

[0248] The first determining subunit is used to divide multiple character elements according to the line association relationship to obtain multiple text lines, and to determine the reference vertical coordinate of the text line based on the first character position information of the character elements in each text line.

[0249] The adjustment sub-unit is used to adjust the character ordinate corresponding to the first character position information of the character element in the text line according to the reference ordinate, so as to obtain the second character position information corresponding to each character element information.

[0250] Optionally, in some embodiments, the data processing apparatus provided in this disclosure further includes:

[0251] The extraction sub-unit is used to extract the font information corresponding to the character element in each text line from the character attribute information, and to extract the character horizontal coordinate corresponding to the character element in each text line from the first character position information;

[0252] The merge sub-unit is used to merge multiple character elements in each text line using font information, character x-coordinates, and rotation angles.

[0253] Optionally, in some embodiments, the rotation angle and the position information of the first character are determined in the first coordinate system. The data processing apparatus provided in this disclosure further includes:

[0254] The transformation subunit is used to perform coordinate transformation based on the rotation angle corresponding to each character element and the adjusted position information of the first character to obtain the rotation angle and the position information of the first character in the second coordinate system.

[0255] The second determining sub-unit is used to determine the second character position information corresponding to each character element information based on the rotation angle in the second coordinate system and the first character position information.

[0256] Optionally, in some embodiments, the second adjustment unit includes:

[0257] The third determining subunit is used to determine the line spacing between any two text lines based on the position information of the second character;

[0258] The fourth determination subunit is used to extract the font information corresponding to the character element from the character attribute information, and determine the paragraph relationship between multiple character elements based on the line spacing, font information and preset line spacing threshold.

[0259] Optionally, in some embodiments, the second adjustment unit includes:

[0260] Sub-units are used to divide multiple lines of text into paragraphs based on paragraph relationships, resulting in multiple text paragraphs.

[0261] The adjustment sub-unit is used to adjust the character ordinate in the second character position information of multiple character elements based on the preset reference line spacing and the line spacing of adjacent text lines in each text paragraph, so as to obtain the third character position information corresponding to each character element information.

[0262] Optionally, in some embodiments, the data processing apparatus provided in this disclosure further includes:

[0263] The first acquisition subunit is used to acquire a second distance value and a third distance value between the boundary of each text paragraph and the boundary of each text line in the text paragraph. The second distance value indicates the distance between the left side of the text line and the left side of the text paragraph, and the third distance value indicates the distance between the right side of the text line and the right side of the target merged paragraph.

[0264] The second acquisition subunit is used to acquire the fourth distance value between the center line of each text paragraph and the center line of each text line in the text paragraph;

[0265] The comparison sub-unit is used to compare the second distance value, the third distance value, and the fourth distance value with a preset distance threshold, and determine the alignment of each text paragraph based on the three comparison results.

[0266] Optionally, in some embodiments, the acquisition unit includes:

[0267] The parsing subunit is used to parse portable document format files to obtain first character element data in Extensible Markup Language format and second character element data in Scalable Vector Graphics format.

[0268] The fifth determining subunit is used to determine the character element information of multiple character elements in a portable document format file based on the first character element data and the second character element data.

[0269] Optionally, in some embodiments, the fifth determining subunit includes:

[0270] The second detection module is used to detect text character elements and vector text elements in the first character element data and the second character element data, and obtain the detection results;

[0271] The filtering module is used to filter vector text elements when the detection result indicates that both text character elements and vector text elements exist in either the first character element data or the second character element data, and to determine the character element information of multiple character elements in the portable document format file based on the text character elements.

[0272] The determination module is used to determine the character element information of multiple character elements in a portable document format file based on the vector text elements when the detection result indicates that only vector text elements exist in both the first character element data and the second character element data.

[0273] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0274] Reference Figure 19 , Figure 19To implement the structural block diagram of a portion of the terminal 340 of the data processing method according to an embodiment of this disclosure, the terminal 340 includes: a radio frequency (RF) circuit 1910, a memory 1915, an input unit 1930, a display unit 1940, a sensor 1150, an audio circuit 1960, a wireless fidelity (WiFi) module 1970, a processor 1980, and a power supply 1990, etc. Those skilled in the art will understand that... Figure 19 The terminal 340 structure shown does not constitute a limitation on a mobile phone or computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0275] The RF circuit 1910 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 1980; in addition, it transmits uplink data to the base station.

[0276] The memory 1915 can be used to store software programs and modules. The processor 1980 executes various terminal functions and document editing by running the software programs and modules stored in the memory 1915.

[0277] The input unit 1930 can be used to receive input numeric or character information, and to generate key signal inputs related to the settings and function control of the terminal. Specifically, the input unit 1930 may include a touch panel 1931 and other input devices 1932.

[0278] Display unit 1940 can be used to display input or provided information, as well as various menus of the terminal. Display unit 1940 may include display panel 1941.

[0279] Audio circuitry 1960, speaker 1961, and microphone 1962 provide an audio interface.

[0280] In this embodiment, the processor 1980 included in the terminal 340 can execute the data processing method of the previous embodiment.

[0281] The terminal 340 in this embodiment includes, but is not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc.

[0282] Figure 20This is a partial structural block diagram of a server 310 for implementing the data processing method of this disclosure embodiment. The server 310 can vary significantly due to different configurations or performance, and may include one or more central processing units (CPUs) 2022 (e.g., one or more processors) and storage devices 2032, and one or more storage media 2030 (e.g., one or more mass storage devices) for storing application programs 2042 or data 2044. The storage devices 2032 and storage media 2030 can be temporary or persistent storage. The program stored in the storage media 2030 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server 310. Furthermore, the central processing unit 2022 may be configured to communicate with the storage media 2030 and execute the series of instruction operations in the storage media 2030 on the server 310.

[0283] Server 310 may also include one or more power supplies 2026, one or more wired or wireless network interfaces 2050, one or more input / output interfaces 2058, and / or one or more operating systems 2041, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0284] The central processing unit 2022 in server 310 can be used to execute the data processing method of the embodiments of this disclosure.

[0285] This disclosure also provides a computer-readable storage medium for storing program code for executing the data processing methods of the foregoing embodiments.

[0286] This disclosure also provides a computer program product comprising a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the data processing method described above.

[0287] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “including,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.

[0288] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0289] It should be understood that in the description of the embodiments of this disclosure, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.

[0290] In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0291] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0292] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0293] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned computer-readable storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0294] It should also be understood that the various implementation methods provided in this disclosure can be combined arbitrarily to achieve different technical effects.

[0295] The above is a detailed description of the embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.

Claims

1. A data processing method, characterized in that, The method includes: A portable document format file is obtained, and the portable document format file is parsed to obtain character element information of multiple character elements in the portable document format file. The character element information includes first character position information and character attribute information. The row association relationship between the multiple character elements is detected based on the character attribute information, and the first character position information of the multiple character elements is adjusted according to the row association relationship to obtain the second character position information corresponding to each character element information; Based on the character attribute information, the second character position information, and the preset line spacing threshold, the paragraph relationship between the multiple character elements is determined, and the second character position information of the multiple character elements is adjusted according to the paragraph relationship to obtain the third character position information corresponding to each character element information; The interface design data corresponding to the portable document format file is generated based on the third character position information and the character attribute information.

2. The method according to claim 1, characterized in that, The step of detecting the row association relationship between the multiple character elements based on the character attribute information includes: Extract the character deformation parameters corresponding to each character element from the character attribute information, and calculate the rotation angle corresponding to each character element based on the character deformation parameters; The row association relationship between the multiple character elements is detected based on the rotation angle.

3. The method according to claim 2, characterized in that, The step of detecting the row association relationship between the multiple character elements based on the rotation angle includes: Extract the origin coordinates of each character element from the character attribute information, and calculate the slope angle between any two character elements among the plurality of character elements based on the origin coordinates. Calculate the first distance value between any two character elements among the plurality of character elements based on the coordinates of the origin point. The row association relationship between the plurality of character elements is detected based on the slope angle, the first distance value, and the rotation angle.

4. The method according to claim 3, characterized in that, The step of detecting the row association relationship between the multiple character elements based on the slope angle, the first distance value, and the rotation angle includes: Determine the first angle difference between the rotation angles corresponding to any two character elements, and determine the second angle difference between the rotation angle of each character element and the corresponding slope angle in any two character elements; The system determines a first comparison result between the first angle difference and a preset first angle difference threshold, a second comparison result between the second angle difference and a preset second angle difference threshold, and a third comparison result between the first distance value and a preset first distance threshold. Based on the first comparison result, the second comparison result, and the third comparison result, the row association relationship between the plurality of character elements is detected.

5. The method according to claim 4, characterized in that, The step of adjusting the first character position information of the plurality of character elements according to the row association relationship to obtain the second character position information corresponding to each character element information includes: The multiple character elements are divided according to the line association relationship to obtain multiple text lines, and the reference vertical coordinate of the text line is determined based on the first character position information of the character elements in each text line. The character ordinates corresponding to the first character position information of the character element in the text line are adjusted according to the reference ordinate to obtain the second character position information corresponding to each character element information.

6. The method according to claim 4, characterized in that, After detecting the row association relationship between the plurality of character elements based on the first comparison result, the second comparison result, and the third comparison result, the method further includes: The font information corresponding to the character element in each text line is extracted from the character attribute information, and the horizontal coordinate of the character element in each text line is extracted from the first character position information. The font information, the horizontal coordinate of the character, and the rotation angle are used to merge multiple character elements in each text line into strings.

7. The method according to claim 5, characterized in that, The rotation angle and the first character position information are determined in the first coordinate system. After adjusting the character ordinate corresponding to the first character position information of the character element in the text line according to the reference ordinate, the method further includes: Based on the rotation angle corresponding to each character element and the adjusted position information of the first character, coordinate transformation is performed to obtain the rotation angle and position information of the first character in the second coordinate system. The second character position information corresponding to each character element information is determined based on the rotation angle in the second coordinate system and the first character position information.

8. The method according to claim 1, characterized in that, The step of determining the paragraph relationship between the multiple character elements based on the character attribute information, the second character position information, and a preset line spacing threshold includes: The line spacing between any two text lines is determined based on the position information of the second character; The font information corresponding to the character elements is extracted from the character attribute information, and the paragraph relationship between the multiple character elements is determined based on the line spacing, the font information, and the preset line spacing threshold.

9. The method according to claim 8, characterized in that, The step of adjusting the second character position information of the plurality of character elements according to the paragraph relationship to obtain the third character position information corresponding to each character element information includes: The multiple text lines are divided into paragraphs based on the paragraph relationships to obtain multiple text paragraphs; Based on the preset reference line spacing and the line spacing between adjacent text lines in each of the text paragraphs, the character ordinates in the second character position information of the multiple character elements are adjusted to obtain the third character position information corresponding to each character element information.

10. The method according to claim 9, characterized in that, After determining the paragraph relationship between the multiple character elements based on the character attribute information, the second character position information, and a preset line spacing threshold, the method further includes: Obtain a second distance value and a third distance value between the boundary of each text paragraph and the boundary of each text line in the text paragraph, wherein the second distance value indicates the distance between the left side of the text line and the left side of the text paragraph, and the third distance value indicates the distance between the right side of the text line and the right side of the target merged paragraph; Obtain a fourth distance value between the midline of each of the text paragraphs and the midline of each of the text lines in the text paragraph; The second distance value, the third distance value, and the fourth distance value are compared with preset distance thresholds, and the alignment of each text segment is determined based on the three comparison results.

11. The method according to claim 1, characterized in that, The step of parsing the portable document format file to obtain character element information of multiple character elements in the portable document format file includes: The portable document format file is parsed to obtain first character element data in Extensible Markup Language format and second character element data in Scalable Vector Graphics format; The character element information of multiple character elements in the portable document format file is determined based on the first character element data and the second character element data.

12. The method according to claim 11, characterized in that, The step of determining the character element information of multiple character elements in the portable document format file based on the first character element data and the second character element data includes: The text character elements and vector text elements in the first character element data and the second character element data are detected to obtain the detection results. When the detection result indicates that both text character elements and vector text elements exist in either the first character element data or the second character element data, the vector text elements are filtered, and the character element information of multiple character elements in the portable document format file is determined based on the text character elements. When the detection result indicates that both the first character element data and the second character element data contain only vector text elements, the character element information of multiple character elements in the portable document format file is determined based on the vector text elements.

13. A data processing apparatus, characterized in that, The device includes: The acquisition unit is used to acquire a portable document format file and parse the portable document format file to obtain character element information of multiple character elements in the portable document format file. The character element information includes first character position information and character attribute information. The first adjustment unit is used to detect the row association relationship between the plurality of character elements according to the character attribute information, and adjust the first character position information of the plurality of character elements according to the row association relationship to obtain the second character position information corresponding to each character element information; The second adjustment unit is used to determine the paragraph relationship between the plurality of character elements based on the character attribute information, the second character position information and the preset line spacing threshold, and to adjust the second character position information of the plurality of character elements according to the paragraph relationship to obtain the third character position information corresponding to each character element information. The generation unit is used to generate interface design data corresponding to the portable document format file based on the third character position information and the character attribute information.

14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the data processing method according to any one of claims 1 to 12.

15. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the data processing method according to any one of claims 1 to 12.

16. A computer program product comprising a computer program that is read and executed by a processor of a computer device, causing the computer device to perform the data processing method according to any one of claims 1 to 12.