A document processing method and device based on semantic understanding

CN122596013APending Publication Date: 2026-08-18BEIJING 21VIANET DATA CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610659880.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

传统文档处理主要依赖人工实现文档调整,难以满足大批量、高标准的文档处理需求

Benefits of technology

[0006]上述方法中,由于初始结构标签能够准确标识文档的内容类型和布局结构,可以为后续转换提供结构依据。初始结构标签可以使文档结构可被程序理解,实现自动化处理。将结构化数据填充至目标结构标签对应的文档位置可以实现了内容与结构的结合,生成完整的目标文档结构。同时,本申请通过自动化填充数据,可以确保内容准确填入对应位置,避免人工排版错误。基于目标结构标签标识的内容类型与布局结构调整待处理文档的格式,生成目标文档,实现根据需求灵活生成不同格式的目标文档,满足用户的文档处理需求。相较于现有技术中人工实现文档格式转换,本申请通过自动化文档处理大幅减少人工成本,同时可以显著提升文档处理的效率,减少文档处理时间,并提升文档处理的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122596013A_ABST
    Figure CN122596013A_ABST
Patent Text Reader

Abstract

The application relates to the computer technical field, in particular to a document processing method and device based on semantic understanding. The method comprises the following steps: generating an initial structure label of a to-be-processed document based on a content type and a layout structure of the to-be-processed document, wherein the content type comprises digital information presented in different carrier forms, and the layout structure comprises a position relationship of content blocks, a page structure and a layout structure; replacing the initial structure label in the to-be-processed document with a target structure label based on a mapping relationship between the initial structure label and the target structure label; filling structured data to a document position corresponding to the target structure label, wherein the structured data is content data extracted from the to-be-processed document and subjected to structured processing; and adjusting a format of the to-be-processed document based on a content type and a layout structure identified by the target structure label, and generating a target document. The above method can improve the automation capability and accuracy of document processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a document processing method and apparatus based on semantic understanding. Background Technology

[0002] As a crucial medium for information transmission, the formatting and accuracy of documents significantly impact office efficiency. Traditional document processing relies heavily on manual adjustments, which struggles to meet the demands of processing large volumes of documents to high standards.

[0003] Although some basic format processing tools exist in existing technologies, such as format conversion software and format painters in office software, these tools lack intelligent processing capabilities and are still insufficient in terms of typesetting efficiency and control over format consistency. How to improve the automation and accuracy of document processing has become a question worth discussing. Summary of the Invention

[0004] This application provides a document processing method and apparatus based on semantic understanding, which improves the automation and accuracy of document processing.

[0005] In a first aspect, embodiments of this application provide a document processing method based on semantic understanding, the method comprising: Based on the content type and layout structure of the document to be processed, an initial structure tag is generated for the document to be processed. The initial structure tag is used to identify the content type and layout structure of the document to be processed. The content type includes digital information presented in different media formats, and the layout structure includes the positional relationship of content blocks, page structure, and layout structure. Based on the mapping relationship between the initial structure tags and the target structure tags, the initial structure tags in the document to be processed are replaced with the target structure tags. The target structure tags are used to identify the content type and layout structure of the target document. Fill the document position corresponding to the target structure tag with structured data. The structured data is the content data extracted from the document to be processed and processed in a structured manner. The structured data includes at least one or more of the following: text content, tables and images. The format of the document to be processed is adjusted based on the content type and layout structure identified by the target structure tags to generate the target document.

[0006] In the above method, the initial structure tags accurately identify the document's content type and layout structure, providing a structural basis for subsequent conversions. These initial structure tags enable the document structure to be understood by the program, facilitating automated processing. Filling the document positions corresponding to the target structure tags with structured data combines content and structure, generating a complete target document structure. Furthermore, this application ensures accurate content placement through automated data filling, avoiding human formatting errors. Adjusting the format of the document to be processed based on the content type and layout structure identified by the target structure tags generates the target document, allowing for flexible generation of target documents in different formats to meet user document processing needs. Compared to manual document format conversion in existing technologies, this application significantly reduces labor costs through automated document processing, while also significantly improving document processing efficiency, reducing processing time, and enhancing accuracy.

[0007] Optionally, before generating the initial structure tags of the document to be processed based on its content type and layout structure, the method further includes: Extracting structured data from initial documents in non-standard formats; Semantic analysis of structured data based on a large model is used to determine the document type of the initial document to be processed; The standard template is determined based on the document type, and there is a preset correspondence between document types and standard templates; The format of the initial document to be processed is adjusted based on the standard template to obtain a standardized document to be processed.

[0008] The above method extracts structured data from the initial non-standardized document to be processed, transforming the originally messy and unstructured document content into structured data that can be recognized and processed by the program. Because structured data has a unified format, it facilitates subsequent semantic analysis, format conversion, and content optimization, laying the foundation for automated document processing. Based on the powerful semantic understanding capabilities of the large model, document types can be accurately identified, improving classification accuracy and avoiding the subjectivity and errors caused by manual judgment. Standardizing the format of the initial document to be processed unifies the document format, facilitating subsequent tagging processing.

[0009] Optionally, the content types of the document to be processed mentioned above include at least one or more of the following: text content, table content, and image content. Text content includes at least one or more of the following: paragraph, title, font, and font size. Table content includes at least one or more of the following: table, table header, and cell. Image content includes at least one or more of the following: image and image description.

[0010] In the above methods, since the documents to be processed can include diverse content such as text, tables, and images, the document processing method of this application can also improve the efficiency and accuracy of document processing when dealing with complex documents.

[0011] Optionally, the method of filling the structured data of the document to be processed before the document position corresponding to the target structure tag further includes: Based on a large model, the text content is logically proofread to obtain proofreading results. The logical proofreading includes at least one or more of the following: grammar checking and logic optimization. Grammar checking is used to identify grammatical errors, punctuation errors and spelling errors. Logic optimization is used to analyze the logical structure of paragraphs. Logic optimization is also used to determine whether the expression of sentences conforms to preset standards. The text content is optimized based on the proofreading results to obtain the optimized text content; The structured data of the document to be processed is populated into the document positions corresponding to the target structure tags, specifically including: The optimized text content, along with other structured data excluding the text content in the document to be processed, is then filled into the document positions corresponding to the target structure tags.

[0012] In the above methods, language logic proofreading of text content based on a large model can automatically identify and correct grammatical errors, punctuation errors, and spelling errors, thereby improving the language quality of the document; through logic optimization, paragraph structure and sentence expression can be analyzed to make the document more in line with preset standards.

[0013] Optionally, the target document may be in one or more of the following formats: Hypertext Markup Language (HTML), Cascading Style Sheets (CSS), Portable Document (PDF), and Word.

[0014] The above methods, based on a variety of different target document formats, can meet the different document processing needs of users, thereby improving the user experience.

[0015] Optionally, the above methods also include: If an error occurs during document processing, an alarm message will be generated.

[0016] In the above method, if an abnormality occurs during document processing, an alarm message is generated, which can promptly notify those skilled in the art to handle the problem, avoid batch task failure, and improve the reliability and maintainability of the system.

[0017] Optionally, the above methods also include: Monitor document processing progress in real time and output progress prompts. The progress prompts should include at least the percentage of document processing progress, which is the percentage of the document content data that has been processed out of the total content data of the entire document to be processed.

[0018] The above method accurately quantifies the document processing progress based on the proportion of content data volume and outputs the progress percentage in real time, which can intuitively show the current status of document processing and improve the user experience.

[0019] Secondly, embodiments of this application provide a document processing apparatus based on semantic understanding, comprising: The processing module is used to generate initial structure tags for the document to be processed based on its content type and layout structure. The initial structure tags are used to identify the content type and layout structure of the document to be processed. The content type includes digital information presented in different media formats, and the layout structure includes the positional relationship of content blocks, page structure, and layout structure. The processing module is also used to replace the initial structure tags in the document to be processed with the target structure tags based on the mapping relationship between the initial structure tags and the target structure tags. The target structure tags are used to identify the content type and layout structure of the target document. The fill module is used to fill structured data into the document positions corresponding to the target structure tags. The structured data is content data extracted from the document to be processed and processed in a structured manner. The structured data includes at least one or more of the following: text content, tables, and images. The processing module is also used to adjust the format of the document to be processed based on the content type and layout structure identified by the target structure tag, and generate the target document.

[0020] Thirdly, embodiments of this application also provide a computer device, including: Memory, used to store program instructions; A processor is used to call program instructions stored in memory and execute any of the methods in the first aspect above according to the obtained program.

[0021] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the method described in any of the first aspects above.

[0022] Fifthly, embodiments of this application also provide a computer program product, which includes an executable program that is executed by a processor using the methods described in any of the first aspects above. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 An application scenario diagram of a document processing method based on semantic understanding provided in this application embodiment; Figure 2 A schematic diagram illustrating a document processing method based on semantic understanding, provided in an embodiment of this application; Figure 3 An exemplary flowchart illustrating a document processing method based on semantic understanding, provided for an embodiment of this application; Figure 4 A schematic diagram illustrating a semantic understanding-based document processing structure provided in this application embodiment; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] The application scenarios described in this application are for the purpose of more clearly illustrating the technical solutions of this application, and do not constitute a limitation on the technical solutions provided in this application. Those skilled in the art will understand that with the emergence of new application scenarios, the technical solutions provided in this application are also applicable to similar technical problems. In the description of this application, unless otherwise stated, "multiple" means two or more.

[0027] like Figure 1 As shown in the figure, this application provides an application scenario diagram of a document processing method based on semantic understanding. Figure 1The system includes a terminal 101 and a server 102. The terminal 101 and server 102 can communicate via a network to implement the document processing method of this application. The terminal 101 can have various client applications installed, such as programming applications, web browser applications, search applications, etc. The terminal 101 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, desktop computers, etc. The server 102 can be a standalone server or a server cluster consisting of multiple servers.

[0028] Terminal 101 generates initial structure tags for the document to be processed based on its content type and layout structure. These initial structure tags identify the content type and layout structure of the document. Content types include digital information presented in different formats, and layout structure includes the positional relationships of content blocks, page structure, and typesetting structure. Based on the mapping relationship between the initial structure tags and target structure tags, the initial structure tags in the document to be processed are replaced with target structure tags. These target structure tags identify the content type and layout structure of the target document. Structured data is then populated into the document positions corresponding to the target structure tags. This structured data is content data extracted from and structured from the document to be processed, and includes at least one or more of the following: text content, tables, and images. Finally, the format of the document to be processed is adjusted based on the content type and layout structure identified by the target structure tags to generate the target document.

[0029] It is understood that either terminal 101 or server 102 can be used to implement the document processing method provided in this application embodiment. In this application embodiment, server 102 can be a standalone server or a server cluster composed of multiple servers. Terminal 101 can be various electronic devices with a display screen and support web browsing, including but not limited to smartphones, tablets, desktop computers, etc.

[0030] The following explanation uses the example of a terminal executing the document processing method provided in this application. Figure 2 As shown in the flowchart, this application provides a document processing method based on semantic understanding, including the following steps: Step S201: Based on the content type and layout structure of the document to be processed, generate the initial structure tags of the document to be processed.

[0031] The initial structure tag identifies the content type and layout structure of the document to be processed. Content type includes digital information presented in different media formats. Layout structure includes the positional relationships of content blocks, page structure, and typesetting structure. The content type of the document to be processed must include at least one or more of the following: text content, table content, and image content. Text content must include at least one or more of the following: paragraphs, headings, fonts, and font sizes. Table content must include at least one or more of the following: tables, table headers, and cells. Image content must include at least one or more of the following: images and image captions. Image captions can be used to describe the insertion position and size of the image.

[0032] In one optional embodiment, the user can upload an initial document to be processed in a non-standard format through the terminal's interactive interface; wherein, the format of the initial document to be processed in a non-standard format can be HyperText Markup Language (HTML), Cascading Style Sheets (CSS), Portable Document Format (PDF), Microsoft Word Document (DOC), Microsoft Word Open XMLDocument (DOCX), or other document formats.

[0033] Optionally, to achieve batch document processing and reduce document processing time, users can also upload an initial set of documents to be processed through the terminal's interactive interface, and the terminal receives this initial set of documents. The initial set of documents to be processed includes at least two documents in non-standard formats. If the terminal receives an initial set of documents to be processed, the terminal executes the document processing method of this application on each document in the initial set of documents to achieve automated batch document processing.

[0034] After receiving the initial document to be processed, the terminal can extract structured data from the non-standardized format of the initial document. This structured data refers to content data extracted from the document and processed in a structured manner. Structured data includes at least one or more of the following: text content, tables, and images. For example, the terminal can use a document parsing engine to extract structured data from the non-standardized format of the initial document. This structured data may include information such as existing fonts, margins, paragraph structure, table headers, cell content, images, and the insertion position and size of images. The document parsing engine is used to identify and extract structured data (such as text content, tables, and images) from the document. The document parsing engine can be a proprietary engine; it can also be a third-party Optical Character Recognition (OCR) interface (such as Baidu OCR or Tencent OCR) combined with format parsing logic. Optionally, proprietary document parsing engines have stronger adaptability and higher parsing accuracy for common digital documents such as Word and PDF; third-party OCR recognition interfaces are suitable for recognizing non-digital documents such as scanned documents and image documents, and can handle scenarios where text cannot be directly read. The terminal can also determine a more suitable parsing engine based on the format of the initial document to be processed, so as to improve the accuracy and applicability of structured data extraction.

[0035] The method described above extracts structured data from the initial non-standardized document, transforming the originally messy and unstructured document content into structured data that can be recognized and processed by the program. Because structured data has a unified format, it facilitates subsequent semantic analysis, format conversion, and content optimization, laying the foundation for automated document processing.

[0036] After obtaining structured data, the terminal can input the structured data into the large model, enabling the model to perform semantic analysis on the structured data and determine the document type of the initial document to be processed. In this way, based on the powerful semantic understanding capabilities of the large model, document types can be accurately identified, improving classification accuracy, avoiding the subjectivity and errors caused by human judgment, providing a basis for subsequently selecting the correct standard template, and ensuring the accuracy of document format conversion.

[0037] After determining the document type of the initial document to be processed, the terminal determines a standard template based on the document type, where a preset correspondence exists between document types and standard templates. For example, the terminal can determine the standard template corresponding to a document type from a database; wherein the database stores standard templates corresponding to at least two different document types. In this way, by preset standard templates corresponding to different document types, it is possible to flexibly adapt to various business scenarios, unify document formats, and improve the universality and efficiency of document processing.

[0038] It is understandable that those skilled in the art can pre-set the parameter configuration of the standard template, and the parameters of the standard template can also be changed by those skilled in the art. For example, the standard template can be a standard template for a patent document, and the standard template for a patent document can be set to format parameters such as SimSun font size 12 and 1.27 cm page margin.

[0039] Optionally, to meet users' personalized document processing needs, the terminal can also receive a standard template defined by the user and use the user-defined standard template as the basis for adjusting the format of the initial document to be processed.

[0040] In one possible scenario, if the terminal receives a document processing request that converts an initial non-standardized document into another specific format, such as HTML, CSS, PDF, or Word, the terminal first determines a corresponding standard template from a database based on the document type (e.g., patent application, daily report, office document). A pre-defined correspondence exists between document types and standard templates. The database stores standard templates corresponding to at least two different document types. Subsequently, the terminal adjusts the format of the initial document based on the determined standard template, resulting in a standardized document. For example, format adjustments based on the standard template might include: unifying the font (e.g., replacing non-SimSun text with SimSun), adjusting the font size (e.g., changing from size 5 to 12pt), setting line spacing (e.g., fixing it to 20 points), and adjusting paragraph indentation (e.g., first-line indentation of 2 characters). Finally, subsequent document processing steps are performed based on the standardized document.

[0041] In another possible scenario, if the terminal receives a document processing request to convert an initial non-standard format document into a standard format and output it, the terminal can determine a standard template based on the document type, adjust the format of the initial document based on the determined standard template, and then directly output the target document in a standardized format.

[0042] The following explains how to format an initial document to be processed based on a standard template to obtain a standardized document format: The terminal can obtain the standard structure tags of the standard template from the database based on the document identifier of the standard template. Among them, the standard structure tags are used to identify the content type and layout structure of the standard template. The document identifier is used to identify different documents. The terminal can also generate the original structure tags of the initial document to be processed based on the content type and layout structure of the initial document to be processed. Among them, the original structure tags are used to identify the content type and layout structure of the initial document to be processed. First, the terminal can modify the original structure tags in the initial document to be processed into standard structure tags; subsequently, the terminal adjusts the initial document to be processed according to the content type and layout structure identified by the standard structure tags to obtain the document to be processed. For example, the terminal can receive the document identifier of the initial document to be processed, query and obtain the corresponding standard structure tags of the document from the data table in the database. The standard structure tags can include parameters such as font type (font_family), font size (font_size), top / right / bottom / left page margins (margin_top / margin_right / margin_bottom / margin_left), table border (table_border), and header font weight (header_font_weight). Subsequently, the original data queried in the initial document to be processed is converted into recognizable structured data, providing an effective configuration basis for the subsequent format adjustment of the initial document to be processed. For example, convert "Small Four" to "12pt", convert "1.27" to "1.27cm", convert "yes" to "1px solid #000", convert "bold" to "bold", and set default values for parameters without valid configurations (such as the default font is "SimSun" and the default page margin is "1.27cm").

[0043] Optionally, in order to save computer computing resources and improve document conversion efficiency, the terminal can also perform document format conversion based on a pre-configured code template: write the mapping relationship of different types of structure tags into the execution logic corresponding to the code template, and then trigger and execute the conversion code to achieve precise matching of heterogeneous tags and automatic conversion of document formats. For example, write the mapping relationship between the target structure tags and the initial structure tags into the execution logic corresponding to the code template. Another example is to write the mapping relationship between the standard structure tags and the original structure tags into the execution logic corresponding to the code template, and then trigger and execute the conversion code to achieve precise matching of heterogeneous tags and automatic conversion of document formats.

[0044] After receiving the document to be processed, the terminal generates initial structure tags for the document based on its content type and layout. For example, initial structure tags might include: Title: centered, font (bold, size 12); First paragraph: left-aligned, used to describe background technology; Table: two rows and three columns, used to list comparative data; Image: a flowchart located below the table. It should be noted that the initial structure tags in this step can be newly created by the terminal after real-time analysis of the document's content; or they can be preset standard structure tags retrieved from the database based on the document's content type and layout. By reusing standard tags from the database as initial structure tags, the terminal can ensure tag standardization and improve the efficiency of document standardization processing.

[0045] Step S202: Based on the mapping relationship between the initial structure tags and the target structure tags, replace the initial structure tags in the document to be processed with the target structure tags.

[0046] The target structure tag is used to identify the content type and layout structure of the target document.

[0047] It is understandable that target result tags are determined based on the content type and layout structure of the target document. For example, target structure tags for an HTML document may include: tables... Labels, headers ; where, th> is the exclusive identifier for the header cell; the content after style= is the specific format configuration items for the header: font-weight:bold: specifies that the header text is in bold; font-family:SimSun: specifies that the font of the header text is uniformly SimSun; font-size:12pt: specifies that the font size of the header text is uniformly 12 points. For example, the target structure tag of a CSS document can include body{font-family:SimSun;font-size:12pt;margin:1.27cm 1.27cm 1.27cm 1.27cm;}; where body indicates that the following rules apply to the entire body area of ​​the CSS document and are the global style base for the document content; font-family:SimSun specifies that the font of the document body is uniformly SimSun; font-size:12pt specifies that the font size of the document body is uniformly 12 points; margin:1.27cm1.27cm1.27cm1.27cm specifies that the margins of the document body area are uniformly 1.27 cm on the top, right, bottom, and left.

[0048] Step S203: Fill the structured data into the document position corresponding to the target structure tag.

[0049] Structured data refers to content data extracted from and structured from the document to be processed. Structured data includes at least one or more of the following: text content, tables, and images.

[0050] It is understandable that the document to be processed is a formatted version of the initial document to be processed, and the structured data contained in the document to be processed is exactly the same as the structured data in the initial document to be processed.

[0051] The above method, by filling structured data into the document positions corresponding to the target structure tags, combines the document content with the document structure, generating a complete target document structure. Simultaneously, automated data filling ensures that the document content is accurately filled into the corresponding positions, avoiding manual formatting errors. It supports various content types such as text, tables, and images, meeting the processing needs of complex documents.

[0052] For example, a terminal can populate structured data into the document location corresponding to the target structure tag using the following steps: After obtaining structured data containing text, tables, and images, the terminal can iterate through each piece of structured data, and according to the different data types, use the corresponding filling logic to insert the content into the Word document object, and configure the local format of the corresponding content based on standardized format parameters; for example, for text content filling: the terminal calls the paragraph addition method of the Word document object to create a new document paragraph and fill the text content in the element; at the same time, it configures the standardized format for the paragraph, sets the first line indentation to 2 characters (converted to inches that python-docx can recognize according to the conversion ratio of 1 character ≈ 0.423 cm), and sets 1.5 line spacing to ensure that the text content layout meets business specifications (such as the text layout requirements of patent documents). For example, when filling table content: the terminal first extracts the table header and body data, counts the number of rows (header + body rows) and columns (header columns), then calls the Word document object's table addition method to create a blank table with the corresponding number of rows and columns, and applies the Table Grid style to the table to achieve standardized display of table borders; next, it fills the table header, after obtaining the first row cell, writes the header content sequentially, and configures the format for the header text (for example, setting it to bold, while using the font name and size in format_params to ensure that the header font is consistent with the document's global format); finally, it fills the table body content, iterates through the body data and writes it to the cells row by row and column by column, and configures the font name and size for the body text to be consistent with the global format, ensuring that the table's overall format is standardized and the content is complete. For example, when filling in image content: the terminal first verifies the existence of the image file using a file path check method to avoid document generation failure due to an invalid path; if the image path is valid, it calls the image addition method of the Word document object to insert the image into the document, and converts the image's width and height from cm units to inches units that python-docx can recognize (using a conversion ratio of 1 cm = 0.3937 inches) to ensure that the size of the inserted image meets the preset requirements and is coordinated with the overall layout of the document.

[0053] In some embodiments, the structured data of the document to be processed is filled before the document position corresponding to the target structure tag. The terminal can input the text content in the structured data into the large model for language logic proofreading, and obtain the proofreading result output by the large model. The language logic proofreading includes at least one or more of grammar checking and logic optimization. Grammar checking is used to identify grammatical errors, punctuation errors, and spelling errors (such as typos). Logic optimization is used to analyze the logical structure of paragraphs (such as cause-effect reversal, content repetition), and also to determine whether the expression of sentences conforms to preset standards. The proofreading result output by the large model includes at least one or more of error type, error location, and optimization suggestions. For example, the proofreading result can indicate the error type (such as grammatical error, logical fallacy, etc.), error location (such as the second sentence of the third paragraph), and optimization suggestions (such as modifying the original sentence to "based on the semantic analysis capability of the large model, automatic format adjustment is achieved"). Optionally, the terminal can display the proofreading result to the user, who can submit optimization requests based on the proofreading result, and the terminal optimizes the text based on the received optimization requests. For example, users can choose "One-click application of optimization suggestions" to have the system automatically complete the optimization; or they can choose "Manual adjustment" to modify the text themselves, thus obtaining the optimized text content. After obtaining the optimized text content, the terminal can fill the document position corresponding to the target structure tag with the optimized text content and other structured data besides the text content in the document to be processed.

[0054] In the above methods, language logic proofreading of text content based on a large model can automatically identify and correct grammatical errors, punctuation errors, and spelling errors, thereby improving the language quality of the document; through logic optimization, paragraph structure and sentence expression can be analyzed to make the document more in line with preset standards.

[0055] Step S204: Adjust the format of the document to be processed based on the content type and layout structure identified by the target structure tag, and generate the target document.

[0056] In some embodiments, the terminal can traverse the structured data of the document to be processed. For each structured data, it adjusts the format according to the content type and layout requirements of the target structure label corresponding to the structured data, and integrates all the adjusted structured data to obtain the target file. For example, taking the generation of a target file in CSS format as an example: First, the terminal obtains the initial structure label of the document to be processed (the initial structure label can identify the content types and layout structures such as the font type, font size, page margin, table border style, and font weight of the table header of the document to be processed); among them, the document to be processed is obtained after formatting the initial document to be processed based on a standard template (such as converting "small four" in the initial document to "12pt", "1.27 centimeters" to "1.27cm", etc. for formatting adjustments); Subsequently, based on the mapping relationship between the initial structure label and the target structure label, the terminal replaces the initial structure label in the document to be processed with the target structure label, and fills the structured data into the document position corresponding to the target structure label; Finally, based on the content type and layout structure of the filled CSS format, the terminal adjusts the format of the document to be processed and generates a target file in CSS format.

[0057] For another example, taking the generation of a target file in HTML format as an example: First, the terminal obtains the initial structure label of the document to be processed (the initial structure label can identify the content types and layout structures such as the text paragraph format, table structure, and picture attributes of the document to be processed); among them, the document to be processed is obtained after formatting the initial document to be processed based on a standard template (such as standardizing the text paragraphs in the initial document into a unified segmented form, calibrating the table cell size to a standard size, and adjusting the width and height of the picture in proportion, etc. for formatting adjustments); Subsequently, based on the mapping relationship between the initial structure label and the target structure label, the terminal replaces the initial structure label in the document to be processed with the target structure label, and fills the structured data into the document position corresponding to the target structure label; Finally, based on the content type and layout structure of the filled HTML format, the terminal adjusts the format of the document to be processed and generates a target file in HTML format.

[0058] For example, taking the generation of a Word format target file as an example: First, the terminal obtains the initial structure tags of the document to be processed (the initial structure tags can identify the global font style, page margin rules, content types and layout structure such as text, tables, and images in the document to be processed); the document to be processed is obtained by adjusting the format of the initial document to be processed based on a standard template (such as converting the font units in the initial document to Pt values, converting the page margin units from centimeters to inches, and calibrating the paragraph alignment to the standard style, etc.); then, based on the mapping relationship between the initial structure tags and the target structure tags, the terminal replaces the initial structure tags in the document to be processed with the target structure tags, and fills the complete structured content containing text, tables, and images into the document positions corresponding to the target structure tags; finally, based on the content types and layout structure of the completed Word format, the terminal adjusts the format of the document to be processed and generates a Word format target file.

[0059] Optionally, after obtaining the target document, the terminal can save and output the target document, ensuring that the standardized document can be persistently stored and available to the user. Optionally, to ensure that the target document can be generated independently using standardized methods, the terminal can pre-set the storage location of the target document. For example, the target document can be saved locally on the terminal; or, for example, it can be saved on a server. To inform the user or subsequent processes of the document generation result, the terminal can simultaneously output a notification message indicating that the target document has been generated and its storage location, facilitating quick document location and verification by the user. In this way, through the fully automated processing of global format configuration, structured content filling, and final document saving and output, the entire document processing operation can be completed independently without relying on external assistance, achieving automated document generation and output.

[0060] Optionally, the terminal can store different format files (HTML, CSS, Word, etc.) generated from the same document, along with the document identifier and the Python code used for format conversion, in different databases. This categorized storage method facilitates quick tracing of the source, conversion process, and corresponding implementation code for each document, improving the standardization of data management. For example, the terminal can save HTML documents, CSS documents, Python code, and document identifiers in separate UTF-8 encoded storage paths.

[0061] Optionally, the terminal can monitor the document processing progress in real time, output progress prompts, and display document processing progress information. The progress prompts should include at least the percentage of document processing progress, which is the percentage of processed document content data out of the total content data of the entire document to be processed. Specifically, for the currently processing document, the terminal can calculate the total content data of the document (e.g., total bytes, total number of content segments). During document processing, the terminal calculates the amount of processed document content data. The terminal then determines the percentage of processed document content data out of the total content data of the entire document to be processed, obtaining the document processing progress percentage, thus generating and displaying the progress prompts, which are updated in real time on the interface. For example, if the document set to be processed includes one document, and the total content data of the document is equivalent to 100 content units; 75 content units have been processed; the document processing progress percentage = 75 / 100 × 100% = 75%. The terminal can display: The current document processing progress is 75%. For example, a document collection requiring document processing includes multiple documents. The terminal can process the documents in the document collection sequentially according to a preset order. Assume the current document to be processed is the 5th document in the document collection; the total data volume of the 5th document is 100 units, and 80 units have been processed; the progress is 80%; the terminal can display: "Currently processing the 5th document, the current document processing progress is 80%." It is understood that when the terminal processes the documents in the document collection sequentially according to a preset order, this preset order can be an order preset by those skilled in the art, a random order, or the order in which the terminal loads the documents, etc. This application does not specifically limit this.

[0062] For example, the terminal can automatically capture and record various exception logs during document processing, such as document import failures, format recognition errors, and code generation errors. Simultaneously, when an exception occurs during document processing (such as document format recognition failure), an alarm mechanism can be proactively triggered, pushing alarm information to the administrator via pop-ups (e.g., "The third document has an abnormal format recognition; please verify the file's validity.").

[0063] Optionally, the terminal can also determine statistical data during the document processing process; wherein, the statistical data is used to represent the operational indicators during the document processing process. For example, the statistical data may include a cumulative total of 100 documents processed today, with a processing success rate of 98%. In this way, based on the statistical data, it is convenient for those skilled in the art to perform subsequent optimizations, further improving the efficiency and accuracy of automated document processing.

[0064] Optionally, the document processing method in this application can be deployed either locally on the terminal or on a cloud server (e.g., based on a cloud-based Software as a Service (SaaS) deployment). Deploying the document processing method locally on the terminal enhances data security. When deployed on a cloud server, users can access and use it via a web interface, eliminating the need for local hardware maintenance, resulting in lower deployment costs and suitability for more application scenarios, such as cross-regional, multi-departmental, and multi-user collaborative document processing.

[0065] Based on the same technological concept, such as Figure 3 As shown in the figure, this application provides an exemplary flowchart of a document processing method based on semantic understanding, including the following steps: Step S301: The terminal extracts structured data from the initial non-standard format document to be processed; Step S302: Perform semantic analysis on the structured data based on the large model to determine the document type of the initial document to be processed.

[0066] Step S303: Determine the standard template based on the document type. There is a preset correspondence between document types and standard templates.

[0067] Step S304: Adjust the format of the initial document to be processed based on the standard template to obtain a standardized document to be processed.

[0068] Step S305: Based on the content type and layout structure of the document to be processed, generate the initial structure tags of the document to be processed.

[0069] Step S306: Based on the mapping relationship between the initial structure tags and the target structure tags, replace the initial structure tags in the document to be processed with the target structure tags.

[0070] Step S307: Perform language logic proofreading on the text content based on the large model to obtain the proofreading results.

[0071] Step S308: Optimize the text content based on the proofreading results to obtain the optimized text content.

[0072] Step S309: Fill the document position corresponding to the target structure tag with the optimized text content and other structured data other than the text content in the document to be processed.

[0073] Step S310: Adjust the format of the document to be processed based on the content type and layout structure identified by the target structure tag, and generate the target document.

[0074] Based on the same technological concept Figure 4 An exemplary schematic diagram illustrates the structure of a semantic understanding-based document processing apparatus provided in an embodiment of this application, such as... Figure 4 As shown, the device specifically includes: The processing module 401 is used to generate initial structure tags for the document to be processed based on its content type and layout structure. The initial structure tags are used to identify the content type and layout structure of the document to be processed. The content type includes digital information presented in different media formats, and the layout structure includes the positional relationship of content blocks, page structure, and layout structure. The processing module 401 is also used to replace the initial structure tags in the document to be processed with target structure tags based on the mapping relationship between the initial structure tags and the target structure tags. The target structure tags are used to identify the content type and layout structure of the target document. The fill module 402 is used to fill the structured data into the document position corresponding to the target structure tag. The structured data is content data extracted from the document to be processed and processed in a structured manner. The structured data includes at least one or more of the following: text content, tables, and images. The processing module 401 is also used to adjust the format of the document to be processed based on the content type and layout structure identified by the target structure tag, and generate the target document.

[0075] Optionally, before generating the initial structure tags of the document to be processed based on its content type and layout structure, the processing module 401 is further used to: Extracting structured data from initial documents in non-standard formats; Semantic analysis of structured data based on a large model is used to determine the document type of the initial document to be processed; The standard template is determined based on the document type, and there is a preset correspondence between document types and standard templates; The format of the initial document to be processed is adjusted based on the standard template to obtain a standardized document to be processed.

[0076] Optionally, the content type of the document to be processed includes at least one or more of the following: text content, table content, and image content. Text content includes at least one or more of the following: paragraph, heading, font, and font size. Table content includes at least one or more of the following: table, table header, and cell. Image content includes at least one or more of the following: image and image description.

[0077] Optionally, the structured data of the document to be processed is filled before the document position corresponding to the target structure tag. The processing module 401 is also used to: perform language logic proofreading on the text content based on the large model to obtain proofreading results. Language logic proofreading includes at least one or more of grammar checking and logic optimization. Grammar checking is used to identify grammatical errors, punctuation errors and spelling errors. Logic optimization is used to analyze the logical structure of paragraphs. Logic optimization is also used to determine whether the expression of sentences conforms to preset standards. Based on the proofreading results, the text content is optimized to obtain optimized text content. The above-mentioned filling module 402 is specifically used to fill the document position corresponding to the target structure tag with the structured data of the document to be processed.

[0078] Optionally, the target document's format includes at least one or more of the following: Hypertext Markup Language (HTML), Cascading Style Sheets (CSS), Portable Document (PDF), and Word. Optionally, the processing module 401 is further configured to: If an error occurs during document processing, an alarm message will be generated.

[0079] Optionally, the above-mentioned processing module 401 is also used for: Monitor document processing progress in real time and output progress prompts. The progress prompts should include at least the percentage of document processing progress, which is the percentage of the document content data that has been processed out of the total content data of the entire document to be processed.

[0080] Based on the same technical concept, embodiments of this application also provide an electronic device. Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0081] At least one processor 501 and a memory 502 connected to at least one processor 501. In this embodiment, the specific connection medium between the processor 501 and the memory 502 is not limited. Figure 5 The example shown is the connection between processor 501 and memory 502 via bus 500. Bus 500 is... Figure 5 The connections between other components are indicated by thick lines and are for illustrative purposes only, not as limiting information. The Bus 500 can be divided into address bus, data bus, control bus, etc., for ease of representation. Figure 5 The term 501 is represented by a single thick line, but this does not imply that there is only one bus or one type of bus. Alternatively, the processor 501 can also be called a controller; there is no restriction on the name.

[0082] In this embodiment, memory 502 stores instructions executable by at least one processor 501. By executing the instructions stored in memory 502, at least one processor 501 can perform a document processing method as described above. Processor 501 can implement... Figure 4 The functions of each module in the device shown.

[0083] The processor 501 is the control center of the device. It can connect to various parts of the control device through various interfaces and lines. By running or executing instructions stored in memory 502 and calling data stored in memory 502, the processor can perform various functions and process data, thereby monitoring the device as a whole.

[0084] In one possible design, processor 501 may include one or more processing units. Processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, driver interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 501. In some embodiments, processor 501 and memory 502 may be implemented on the same chip; in some embodiments, they may also be implemented on separate chips.

[0085] Processor 501 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the semantic understanding-based document processing method disclosed in the embodiments of this application can be directly manifested as execution by a hardware processor, or as execution by a combination of hardware and software modules within the processor.

[0086] Memory 502, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 502 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 502 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 502 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0087] By designing and programming the processor 501, the code corresponding to a document processing method described in the foregoing embodiments can be embedded into the chip, enabling the chip to execute it during runtime. Figure 2 The illustrated embodiment represents a document processing method based on semantic understanding. How to design and program the processor 501 is a technique well-known to those skilled in the art and will not be described further here.

[0088] It should be noted that the communication electronic device provided in this application embodiment can implement all the method steps implemented in the above method embodiment and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.

[0089] This application also provides a computer-readable storage medium storing computer-executable instructions for causing a computer to perform a semantic understanding-based document processing method as described above.

[0090] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0091] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable document processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable document processing device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the function specified in one or more boxes.

[0092] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable document processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0093] These computer program instructions can also be loaded onto a computer or other programmable document processing device, causing a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0094] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A document processing method based on semantic understanding, characterized in that, The method includes: Based on the content type and layout structure of the document to be processed, an initial structure tag is generated for the document to be processed. The initial structure tag is used to identify the content type and layout structure of the document to be processed. The content type includes digital information presented in different carrier forms, and the layout structure includes the positional relationship of content blocks, page structure, and typesetting structure. Based on the mapping relationship between the initial structure tags and the target structure tags, the initial structure tags in the document to be processed are replaced with the target structure tags, whereby the target structure tags are used to identify the content type and layout structure of the target document; The structured data is filled into the document position corresponding to the target structure tag. The structured data is content data extracted from the document to be processed and has undergone structured processing. The structured data includes at least one or more of the following: text content, tables, and images. The format of the document to be processed is adjusted based on the content type and layout structure identified by the target structure tag, and the target document is generated.

2. The method according to claim 1, characterized in that, Before generating the initial structure tags of the document to be processed based on its content type and layout structure, the method further includes: Extract the structured data from the initial unprocessed document in a non-standard format; Based on the large model, semantic analysis is performed on the structured data to determine the document type of the initial document to be processed; The standard template is determined based on the document type, and there is a preset correspondence between the document type and the standard template; The format of the initial document to be processed is adjusted based on the standard template to obtain the document in a standardized format.

3. The method according to claim 1, characterized in that, The content type of the document to be processed includes at least one or more of the following: text content, table content, and image content. The text content includes at least one or more of the following: paragraph, title, font, and font size. The table content includes at least one or more of the following: table, table header, and cell. The image content includes at least one or more of the following: image and image description.

4. The method according to claim 1, characterized in that, The method further includes filling the structured data of the document to be processed before the document position corresponding to the target structure tag, wherein the structured data of the document to be processed is filled before the target structure tag. Based on the large model, the text content is linguistically and logically proofread to obtain proofreading results. The linguistically and logically proofreading includes at least one or more of grammar checking and logic optimization. The grammar checking is used to identify grammatical errors, punctuation errors, and spelling errors. The logic optimization is used to analyze the logical structure of paragraphs. The logic optimization is also used to determine whether the expression of sentences conforms to preset standards. Based on the proofreading results, the text content is optimized to obtain the optimized text content; The step of filling the structured data of the document to be processed into the document position corresponding to the target structure tag specifically includes: The optimized text content, along with other structured data besides the text content in the document to be processed, is filled into the document position corresponding to the target structure tag.

5. The method according to any one of claims 1 to 4, characterized in that, The target document is in at least one or more of the following formats: Hypertext Markup Language (HTML), Cascading Style Sheets (CSS), Portable Document (PDF), and Microsoft Word.

6. The method according to any one of claims 1 to 4, characterized in that, The method further includes: If an error occurs during document processing, an alarm message will be generated.

7. The method according to any one of claims 1 to 4, characterized in that, The method further includes: The document processing progress is monitored in real time, and progress prompts are output. The progress prompts include at least the percentage of document processing progress, which is the percentage of the document content data that has been processed out of the total content data of the entire document to be processed.

8. A document processing device based on semantic understanding, characterized in that, The device includes: The processing module is used to generate initial structure tags for the document to be processed based on its content type and layout structure. The initial structure tags are used to identify the content type and layout structure of the document to be processed. The content type includes digital information presented in different carrier forms, and the layout structure includes the positional relationship of content blocks, page structure, and layout structure. The processing module is further configured to replace the initial structure tags in the document to be processed with the target structure tags based on the mapping relationship between the initial structure tags and the target structure tags, wherein the target structure tags are used to identify the content type and layout structure of the target document; The fill module is used to fill the document position corresponding to the target structure tag with structured data. The structured data is content data extracted from the document to be processed and processed in a structured manner. The structured data includes at least one or more of the following: text content, tables, and images. The processing module is further configured to adjust the format of the document to be processed based on the content type and layout structure identified by the target structure tag, and generate the target document.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program for causing the computer to perform the method of any one of claims 1-7.