Document processing method and apparatus, content generation method and apparatus, and electronic device
By analyzing and splicing document object partitions and using semantic and geometric information to generate semantic labels, the efficiency and accuracy of identifying and restoring document objects and layout layouts in the prior art are solved, and efficient and accurate document processing is achieved.
Patent Information
- Application Number
- PCT/CN2024/122794
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-09-30
- Publication Date
- 2025-05-30
AI Technical Summary
It is difficult for the prior art to effectively identify and restore document objects and layouts in preset format documents, especially when processing documents in PDF formats such as PDFs, traditional methods are inefficient and have low accuracy.
By parsing the non-text object partition and text object partition in the target document, paragraph splicing is performed on non-text objects and text objects based on the object's semantic information and geometric information to generate semantic labels of paragraphs, so as to accurately identify and restore the document object and layout layout.
It realizes accurate identification of document objects in the target document and accurate restoration of layout layout, improving processing efficiency and accuracy.
Smart Images

Figure CN2024122794_30052025_PF_FP_ABST
Abstract
Description
Document processing, content generation method, device and electronic equipment
[0001] This application claims priority to Chinese Patent Application No. 202311570217.1 filed on November 22, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] Embodiments of the present disclosure relate to a document processing and content generation method, apparatus, and electronic device. Background Art
[0003] In many application scenarios, a large amount of information needs to be collected and structured to facilitate subsequent analysis and processing.
[0004] Information can be recorded in documents, such as Word documents. Some documents contain layout information, which can be used to structure the document.
[0005] Summary of the Invention
[0006] Embodiments of the present disclosure provide a document processing and content generation method, apparatus, and electronic device.
[0007] In a first aspect, an embodiment of the present disclosure provides a document processing method, the method comprising: in response to receiving a first instruction for a target document in a preset format, parsing non-text object partitions and text object partitions in the target document to obtain at least one non-text object partition and at least one text object partition; performing paragraph splicing on the non-text objects and text objects in the non-text object partitions and the text object partitions based on the semantic information and geometric information of the objects to obtain multiple paragraphs; and generating semantic tags for the paragraphs based on the position information and semantic information of the paragraphs.
[0008] In a second aspect, an embodiment of the present disclosure provides a content generation method, which includes: receiving a third content generation instruction; generating third content corresponding to the third content generation instruction based on the content generation instruction and a target paragraph from a plurality of paragraphs extracted from a target document in a preset format; wherein the target document includes text objects and non-text objects, and the plurality of paragraphs are obtained based on the following steps: parsing at least one non-text object partition and at least one text object partition from the target document; and performing paragraph splicing on the non-text objects and text objects in the non-text object partition and the text object partition based on the semantic information and geometric information of the objects to obtain a plurality of paragraphs.
[0009] In a third aspect, an embodiment of the present disclosure provides a document processing device, which includes: a parsing unit, configured to, in response to receiving a first instruction for a target document in a preset format, parse non-text object partitions and text object partitions in the target document to obtain at least one non-text object partition and at least one text object partition; a splicing unit, configured to perform paragraph splicing on the non-text objects and text objects in the non-text object partitions and the text object partitions based on the semantic information and geometric information of the objects to obtain multiple paragraphs; and a first generating unit, configured to generate a semantic tag for the paragraph based on the position information and semantic information of the paragraph.
[0010] In a fourth aspect, an embodiment of the present disclosure provides a content generation device, which includes: a receiving unit, configured to receive a third content generation instruction; a second generation unit, configured to generate a third content corresponding to the third content generation instruction based on the content generation instruction and a target paragraph from a plurality of paragraphs extracted from a target document in a preset format; wherein the target document includes text objects and non-text objects, and the plurality of paragraphs are obtained based on the following steps: parsing at least one non-text object partition and at least one text object partition from the target document; and performing paragraph splicing on the non-text objects and text objects in the non-text object partition and the text object partition based on the semantic information and geometric information of the objects to obtain a plurality of paragraphs.
[0011] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
[0012] The memory stores computer-executable instructions;
[0013] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the method described in the first aspect and various possible designs of the first aspect.
[0014] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the method described in the first aspect and various possible designs of the first aspect are implemented.
[0015] In a seventh aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the method described in the first aspect and various possible designs of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0017] FIG1 is a flowchart of a document processing method provided by the present disclosure;
[0018] FIG2A is a schematic diagram of an application scenario;
[0019] FIG2B is a schematic diagram of an application scenario;
[0020] FIG2C is a schematic diagram of an application scenario;
[0021] FIG3 is a schematic flowchart of parsing a first image partition of a non-vector graphic object provided by the present disclosure;
[0022] FIG4 is a schematic flowchart of parsing a second image partition of a vector graphics object provided by the present disclosure;
[0023] FIG5 is a schematic flow chart of parsing a first table partition of a wireframe table provided by the present disclosure;
[0024] FIG6 is a schematic flow chart of parsing a second table partition of a semi-wired frame table provided by the present disclosure;
[0025] FIG7 is a schematic flow chart of parsing text object partitions provided by the present disclosure;
[0026] FIG8 is a flow chart of a content generation method provided by the present disclosure;
[0027] FIG9 is a schematic structural block diagram of a document processing device;
[0028] FIG10 is a schematic structural block diagram of a content generating device; and
[0029] FIG11 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0031] In some application scenarios, partial content can be extracted from a document based on the document's layout information, and the partial content and the corresponding layout information in the document can be processed for subsequent use. For example, when using a language model to generate content based on instructions, it is usually necessary to extract partial content (such as certain paragraphs) from some documents based on the layout information, send the extracted partial content to the language model, and the language model will analyze and process the partial content to obtain the corresponding content corresponding to the instruction.
[0032] Layout information includes object information. Document objects include the basic units that make up a document, including: title, paragraph title, natural paragraph, header and footer, footnote, picture, table, formula, etc.
[0033] Some document formats, such as Portable Document Format (PDF), only store information such as character fonts, coordinates, size, and color, but do not contain document object information. It is necessary to identify these document objects and restore the PDF layout.
[0034] In order to identify the document objects of documents in the above formats and restore the page layout, in some implementations, a certain type of document object (such as a type of object in text, picture or table) can be parsed and the structure restored, such as parsing and restoring the table in the PDF document, parsing and restoring the picture, etc., but these implementations cannot completely parse and restore all these document objects.
[0035] Other implementations convert these document formats into images, then use deep learning to identify objects and restore the document layout. While these methods can fully identify all document objects and their layout, they are less efficient and the recognition results are less accurate, resulting in less accurate restoration results.
[0036] The method provided by the present invention parses the non-text object partitions and text object partitions obtained from the target document, and then performs paragraph splicing on the non-text objects and text objects in the non-text object partitions and text object partitions based on the object semantic information and geometric information to obtain multiple paragraphs; and generates semantic tags for the paragraphs based on the position information and semantic information of the paragraphs, so as to accurately identify the document objects in the target document and restore the document objects and page layout in the target document.
[0037] The document processing and content generation method, device and electronic device provided by the present embodiment, in response to receiving a first instruction for a target document of a preset format, parses the non-text object partition and the text object partition in the target document to obtain at least one non-text object partition and at least one text object partition; based on the semantic information and geometric information of the object, the non-text objects and text objects in the non-text object partition and the text object partition are respectively spliced into paragraphs to obtain multiple paragraphs; based on the position information and semantic information of the paragraph, the semantic tag of the paragraph is generated, which can accurately identify the document object in the target document and can more accurately restore the layout of the objects in the document. Please refer to Figure 1, which shows a schematic flowchart of the document processing method provided by the present disclosure. As shown in Figure 1, the method includes the following steps:
[0038] S101: In response to receiving a first instruction for a target document in a preset format, parsing non-text object partitions and text object partitions in the target document to obtain at least one non-text object partition and at least one text object partition.
[0039] In this embodiment, the execution subject of the document processing method may be a client running in a terminal device, or a server providing services to the client.
[0040] The first instruction may be a user-initiated instruction or an electronic device-initiated instruction. In other words, the document processing method may be triggered by a user or by an electronic device. The electronic device may send the first instruction to the execution subject according to program instructions.
[0041] The first instruction is used to instruct the parsing of objects in the target document and restore the layout of the target document.
[0042] After receiving the first instruction, the execution subject may use various document parsing methods to parse the non-text object partitions and text object partitions in the target document.
[0043] The text object partition here refers to a partition that contains only text, numbers, letters, etc. The above-mentioned non-text object partitions include image partitions, table partitions, and formula partitions.
[0044] That is, by partitioning the target document according to text objects and non-text objects, non-text object partitions and text object partitions can be obtained.
[0045] Specifically, text objects and non-text objects in the target document may be identified, and then the target document may be divided into at least one non-text object partition and at least one text object partition, wherein the text object partition and the non-text object partition are independent of each other.
[0046] S102: Based on the semantic information and geometric information of the objects, paragraph splicing is performed on the non-text objects and text objects in the non-text object partition and the text object partition respectively to obtain multiple paragraphs.
[0047] The geometric information includes the location information, size information and distance information between different objects.
[0048] After obtaining the text object partition and the non-text object partition, for each text object partition, the text blocks in the text object partition can be segmented according to the semantic information and geometric information of the text objects (such as text blocks) in the text object partition to obtain at least one text paragraph. The position information of the text block may include the page number and the position information in the page corresponding to the page number. The position information of the text block in the page may include the coordinate information of the text block in the page. The coordinate information of the above-mentioned text block in the page may include the coordinate information of the text block in the preset coordinate system. In some embodiments, the above-mentioned preset coordinate system is, for example, a coordinate system constructed with a reference point in the page as the origin, a straight line passing through the origin and parallel to the short side of the page as the X-axis, and a straight line passing through the origin and parallel to the long side of the page as the Y-axis. The above-mentioned reference point can be any specified point in the page. The above-mentioned specified point is, for example, the point corresponding to the lower left corner of the page, the point corresponding to the upper left corner, etc.
[0049] The semantic information of the text object may include semantic information obtained by performing semantic recognition on the content of the text object. The semantic recognition of the text object may be performed using various semantic recognition methods.
[0050] For a non-text object partition, paragraph splicing may be performed on the non-text objects in the non-text object partition according to semantic information and geometric information of the non-text objects in the non-text object partition to obtain at least one non-text object paragraph.
[0051] In some application scenarios, each non-text object partition may include multiple local partitions, and different local partitions may have semantic associations. For example, the size ratio, connectivity, proximity, and relative distance between different local partitions can be considered semantic information of the non-text object partition.
[0052] In some application scenarios, a non-text object partition may include multiple local partitions, and different local partitions may include text and other content. The semantic information between different local partitions in the non-text object partition can be determined based on the above text.
[0053] For example, a non-text object partition corresponding to a table object may include multiple cells. Multiple cells may be filled with text, etc. Semantic information corresponding to the cells may be determined based on the text in the multiple cells. For example, cell A containing "weather" may have corresponding semantic information. Another cell B containing weather information may be located in the same column or row as cell A.
[0054] Semantic associations between cells can also be determined based on their size ratios, connectivity, and adjacency. For example, if cell 1 is connected to cell 2, it means that cells 1 and 2 are spatially connected and belong to the same table. For another example, if the size of cell 3 is determined to be twice the size of cell 4, then cells 4 and 3 belong to the same non-text object partition.
[0055] The above-mentioned position information may include coordinate information, and the non-text objects in different non-text object partitions may be segmented according to the above-mentioned semantic information and geometric information, thereby obtaining multiple non-text object paragraphs.
[0056] Multiple paragraphs can be obtained through the above step S102.
[0057] S103: Generate a semantic tag for the paragraph based on the location information and semantic information of the paragraph.
[0058] Semantic tags may be generated for paragraphs in each page according to the order of different pages of the target document from front to back.
[0059] Please refer to Figures 2A to 2C. Figure 2A is a schematic diagram of an application scenario; Figure 2B is a schematic diagram of an application scenario; and Figure 2C is a schematic diagram of an application scenario.
[0060] As shown in FIG2A , a page 201 of a target document can be subjected to text object recognition to obtain a result, which includes multiple text blocks 202 .
[0061] The page may be partitioned by text objects and non-text objects to obtain multiple partitions, such as multiple text object partitions 2031 , 2032 , 2033 , 2034 and 2036 and a non-text object (table) partition 2035 as shown in FIG. 2B .
[0062] For each partition, paragraph concatenation is performed to obtain multiple paragraphs. Referring to FIG. 2C , the obtained multiple paragraphs include text paragraphs 210 , 211 , 212 , 213 , 214 , 215 , 216 , 217 , 218 , 220 , and 221 , and a table paragraph 219 .
[0063] Specifically, for each page, the page layout of the target document can be determined based on the geometric information of the multiple partitions in the page. The geometric information includes the position, size, and distance between different partitions.
[0064] The above page layout forms include single-column layout and double-column layout. The above double-column layout means that the page is divided into two columns in the horizontal direction.
[0065] If two partitions are detected to have the same row, the widths of the two partitions with the same row are checked to see if they are similar. If so, the multiple partitions on the page are divided into partitions on the left side of the page and partitions on the right side of the page based on their position information. The cumulative area of the partitions on the left side of the page is further calculated to see if it is approximately the same as the cumulative area of the partitions on the right side of the page. If so, the page is determined to be in a double-column layout.
[0066] If no sections with the same row are detected on the page, the page layout is considered to be single-column. If two sections are detected to have the same row, but other sections do not, and the widths of the two sections are significantly different, the page layout is considered to be single-column.
[0067] For different partitions in a page, the order of the partitions in the page can be determined from top to bottom and from left to right.
[0068] For a page with a single column layout, the order of different sections in the page can be determined from top to bottom. Furthermore, for the target document, the order of different sections in each page can be determined.
[0069] For a double-column page, the order of different sections in the page can be determined from left to right and from top to bottom.
[0070] The paragraphs in each partition are analyzed in order, and semantic labels for different paragraphs in each partition are generated.
[0071] The semantic tags may include the order in which the paragraphs appear in the document and / or the document objects corresponding to the paragraphs.
[0072] For single-column or double-column pages, the order of each section is analyzed. Specifically, for each section, the order of the paragraphs within the section is determined based on their position within the section. Semantic tags are generated for each paragraph within the section based on this order.
[0073] Specifically, for the first partition containing the first line of the first page of the document, if the partition is a text object partition and the partition includes a paragraph, the semantic information of the paragraph is determined. If it is determined that the paragraph expresses complete semantics, the paragraph can be regarded as the document title of the document. The semantic tag added to the paragraph can include a paragraph number indicating that it is the first paragraph and a title identifier indicating that the paragraph is the document title.
[0074] For a second partition that appears after the first partition, if the partition contains at least one paragraph, the semantics of the paragraph are analyzed. If the semantics of the at least one paragraph include at least one person's name and / or unit name, the paragraph is determined to contain author information, and a semantic tag with the author information is added to the paragraph.
[0075] For example, if the third partition after the second partition is a text object partition and includes a paragraph, the semantics of the paragraph are analyzed. If the paragraph includes "abstract" and the semantics of the paragraph are a general description of the document, the paragraph can be considered an abstract and a semantic tag of abstract is added to the paragraph.
[0076] For the fourth partition after the third partition, if it is a text object partition and contains at least one paragraph, the semantic information of each paragraph can be further analyzed. Based on the semantic recognition results, a semantic label such as "Main text paragraph 1" or "Main text paragraph 2" can be assigned to each paragraph.
[0077] If an image partition appears after the fourth partition, determine the location of the image partition in the document and whether it is the first image partition. If it is the first image partition and it contains only one image, the semantic label for that image is Image 1. If the image partition contains multiple images, the order of the images can be determined from left to right and from top to bottom, and a corresponding semantic label can be assigned to each image.
[0078] For the subsequent image partitions, the order of the images in the subsequent image partitions can be determined, and then the order of the images in the subsequent image partitions is continued with the order of the images in the previous image partitions.
[0079] If the last section of a page contains only a number, the number can be considered the page number. The paragraph corresponding to the number can be considered the footer.
[0080] By analogy, semantic tags can be generated for paragraphs in different partitions of the document.
[0081] In this embodiment, in response to receiving a first instruction for a target document in a preset format, non-text object partitions and text object partitions in the target document are parsed to obtain at least one non-text object partition and at least one text object partition; paragraphs are spliced for the non-text object partitions and text object partitions based on the semantic information and geometric information of the objects to obtain multiple paragraphs; and semantic tags for the paragraphs are generated based on the positional information and semantic information of the paragraphs. By first dividing the target document into text object partitions and non-text object partitions, then splicing each partition into paragraphs, and finally generating paragraph semantics, the document objects of the target document can be more accurately identified, thereby more accurately restoring the layout information of the objects in the target document.
[0082] In some embodiments, the non-text object includes a non-vector image object. Referring to FIG. 3 , step S101 includes the following steps to parse the first image partition corresponding to the non-vector image object in the target document:
[0083] S301: For a non-vector image, first position information corresponding to the non-vector image is obtained based on an underlying protocol of a target document, where the first position information is position information of a first image partition.
[0084] S302: intercepting a non-vector image in the target document based on the first position information to obtain a non-vector image corresponding to the first image partition.
[0085] The aforementioned non-vector images include images in formats such as PNG and JPG. For these non-vector images, a preset interface can be used to detect the underlying protocol of the target document and extract information about the non-vector image inserted into the target document from the underlying protocol. This non-vector image information includes location information. Based on this location information, a rectangular frame enclosing the non-vector image can be determined. This rectangular frame can be considered the first image partition. The location information of this rectangular frame can be considered the location information of the first image partition. The location information of the first image can include the coordinates of the upper left and lower right vertices of the rectangular frame in a preset coordinate system.
[0086] After determining the location information of the first image partition, the images included in the partition can be obtained. Specifically, a screenshot can be taken in the target document based on the location information of the first image partition to obtain the images included in the first image partition.
[0087] According to the above method, all first image partitions of the target document and the images inserted into each first image partition can be obtained.
[0088] In these embodiments, the non-vector image document object is recognized and restored by dividing the non-vector image into first image partitions and acquiring the image in the first image partitions.
[0089] In some embodiments, the non-text object includes a vector graphic object. Referring to FIG. 4 , step S101 includes the following steps to parse the first image partition corresponding to the non-vector graphic object in the target document:
[0090] S401: Hide the text content in the target document.
[0091] Vector graphics are scalable, lossless image formats. They are generated by combining multiple objects, each of which is recorded using mathematical functions. Unlike bitmaps, which record the information of every point on the screen, vector graphics record algorithms for shape and color. When a vector graphics image is opened, the document application performs the functions corresponding to the objects in the vector graphics and displays the results (the shape and color of the image). Therefore, the image's location information cannot be obtained from the underlying protocol.
[0092] A preset interface can be used to detect the underlying protocol of the target document, extract the display information of the text content of each page from the underlying protocol, and modify the display information of the text content of each page to hide the text content of each page.
[0093] It is understandable that after hiding the text content of each page, the text content in the original vector image will also be hidden. After the text content in the page is hidden, the vector image in the page only includes points, line frames, etc.
[0094] S402: In a target document with hidden text content, identifying second position information corresponding to a vector image and a first vector image without text content; the second position information is used as position information of a second image partition.
[0095] For each page, after hiding the text content, the page is converted into an image through format conversion. The image corresponding to the page is then converted into a grayscale image, and the grayscale image is binarized. For the binarized page image, a projection segmentation algorithm (XY-cut algorithm) is used to identify the first vector image in the page binarization image, determine the rectangular frame surrounding the first vector image, and determine the second position information corresponding to the vector image using the projection segmentation algorithm. The second position information may include the coordinates of the rectangular frame in the preset coordinate system corresponding to the page. The rectangular frame can be regarded as the second image partition corresponding to the first vector image.
[0096] S403: Determine text content in the first vector image according to the second position information.
[0097] S404: Merge the first vector image and the text content to form a vector image corresponding to the second image partition.
[0098] After determining the second position information of the second image partition, the text content corresponding to the first vector image can be obtained from the target document according to the second position information. Specifically, a screenshot can be taken in the target document according to the second position information of the second image partition, and the text content can be obtained from the screenshot.
[0099] The above text content can be integrated into the first vector image to obtain a vector image corresponding to the second image partition.
[0100] By analogy, all second image partitions in the page and the vector images corresponding to each second partition can be obtained.
[0101] In these embodiments, the page containing the vector image is converted into a page image, the location information of the vector image without text content and the corresponding second image partition is identified in the page image, and the text content in the vector image is obtained by taking a screenshot of the original page. The text content is then merged with the second image without text content. The second image partition and the vector image corresponding to the second image are then obtained, thereby achieving the recognition and restoration of the vector image.
[0102] In some embodiments, the non-text object includes a wireframe table object. Referring to FIG. 5 , step S101 includes parsing the first table partition corresponding to the wireframe table object in the target document based on the following steps:
[0103] S501: Acquire third position information of a rectangular frame surrounding a wireframe table and information of each first cell in the wireframe table, wherein the third position information is position information of a first table partition.
[0104] S502: Merging a plurality of first cells according to information of adjacent cells of each first cell to obtain a plurality of merged second cells.
[0105] S503: Determine the size of the smallest cell among the plurality of second cells, and determine a proportional relationship between the other second cells and the smallest cell.
[0106] S504: Generate a blank wireframe table corresponding to the first table partition according to the proportional relationship.
[0107] S505: Merge the text content rectangular frame corresponding to the third position information in the target document with the blank wired frame table to obtain the wired frame table corresponding to the first table partition.
[0108] The interface provided by the wireframe table extraction tool can be used to extract the location of the wireframe table in each page of the target document, determine the minimum rectangular frame that includes the table location, and simultaneously extract the information (location and content) of each first cell in the wireframe table.
[0109] The first table partitions corresponding to the wireframe tables in each page may be parsed page by page.
[0110] For each page, the smallest rectangular box enclosing a wired box in the page can be regarded as a first table partition. The position information of the smallest rectangular box (including the vertices of the smallest rectangular box in the preset coordinate system of the page) can be regarded as the third position information of the first table partition.
[0111] For each first cell in the extracted wired frame, the adjacent cells in the four directions of up, down, left and right of the first cell can be found based on the position information of the first cell and the position information of other cells, and the position information of the adjacent cells can be stored to obtain the adjacency matrix of the first cell.
[0112] For the above adjacency matrix, check whether there are blank first cells above and to the left of each first cell that can be merged. If so, merge them. After merging multiple first cells, multiple second cells are obtained. The position and content information of each second cell is obtained. The width and height of each second cell are calculated to determine the minimum cell. The minimum cell is used as the reference cell.
[0113] For other second cells, the width and height of each other second cell are divided by the width and minimum height of the minimum cell respectively to obtain the number of columns and rows occupied by each other second cell.
[0114] According to the number of rows and columns of each second cell, the coordinates of the second cell in the wired frame are calculated (the above coordinates include the starting row, ending row, starting column, and ending column). Thus, the restored information of each cell in the wired frame (cells without text content) is obtained.
[0115] Based on the third position information, a text content rectangular frame is extracted from the target document. The text content corresponding to each cell in the blank wireframe table is extracted based on the position information of each cell in the text content rectangular frame. Each text content is filled into the corresponding blank second cell. This then results in a first table partition corresponding to the wireframe and the restoration of the table content in the first table partition.
[0116] Through the above method, each wireframe table in the target document can be identified and restored.
[0117] In some embodiments, the non-text object includes a half-line frame table object. Referring to FIG. 6 , step S101 includes parsing the second table partition corresponding to the half-line frame table object in the target document based on the following steps:
[0118] S601: Based on the projection segmentation algorithm and the image recognition algorithm, identify the frame lines used to define the semi-wireframe table.
[0119] The semi-line frame table may be a table whose table area is defined by two parallel frame lines.
[0120] For a frame line in a non-vector image format, an image recognition algorithm may be used to identify a frame line for defining a half-frame. The frame line may be a line segment.
[0121] S602: Determine a pair of frame lines for defining a half-frame according to the positions and sizes of the identified frame lines.
[0122] For example, two frame lines that are closest to each other are determined based on the distances between two endpoints of different frame lines, and the two frame lines are similar in size. The two frame lines can be regarded as a frame line pair.
[0123] S603: Using the fourth position information determined by the frame line pair as the position information of the second table partition.
[0124] The smallest rectangular frame surrounding the half-line frame is used as the second table partition. The position information of the smallest rectangular frame is used as the position information of the second table partition (fourth position information). The fourth position information includes the coordinates of the vertices of the smallest rectangular frame in the page preset coordinate system.
[0125] S604: Acquire the half-line frame table content corresponding to the fourth position information from the target document.
[0126] S605: Generate a half-line frame table corresponding to the second table partition according to the half-line frame table content and the frame line pairs.
[0127] A blank half-line frame table is generated according to the position and size of the frame line pair. A screenshot can be taken in the target document according to the fourth position information, and the text content in the half-line frame table is extracted line by line from the screenshot.
[0128] The above text content is merged with the blank half-line frame to obtain the half-line frame table corresponding to the second table partition.
[0129] Through the above process, the semi-wireframe table document object can be identified from the target document and restored.
[0130] In some embodiments, the non-text object of the target document includes a formula object, and step S101 includes parsing the formula partition corresponding to the formula object in the target document based on the following steps:
[0131] Use regular expressions and preset formula fonts to identify formula content, splice the formula content, and obtain the formula in the target document.
[0132] In these embodiments, for a formula partition, special symbols in the formula can be identified in the partition using regular expressions combined with special formula fonts. After identifying the formula symbols, the formula content is concatenated to obtain the formula in the target document. In this way, all formulas in the target document can be identified, achieving the recognition and restoration of formulas in the target document.
[0133] In some embodiments, referring to FIG. 7 , the above step S101 includes the following steps to parse out the text object partitions:
[0134] S701: Acquire multiple text blocks corresponding to the text content of the target document outside the non-text object partition.
[0135] The text content in the target document can be parsed to obtain multiple text blocks. Specifically, a text object parsing tool corresponding to a preset format can be used to parse the text content of the target document to obtain multiple text blocks. Each text block can correspond to position information and the text content of the text block.
[0136] Each text block may include one or more characters. Referring to FIG2A , the target document is parsed for text objects to obtain multiple text blocks.
[0137] S702: Partitioning multiple text blocks based on a projection segmentation algorithm to obtain at least one text object partition.
[0138] The text objects can be partitioned page by page. For each page, each text block is projected horizontally and vertically onto the page. The position of each text block in the page coordinate system and the rectangular box surrounding the text block are determined. The distance between the rectangular boxes corresponding to each text block is then calculated.
[0139] A first distance threshold can be pre-set to determine whether the distance between the rectangular boxes corresponding to the upper and lower adjacent text blocks is greater than the first distance threshold. If the distance is greater than the first distance threshold, a line is drawn to separate the context blocks. A second distance threshold can be pre-set to determine whether the distance between the left and right adjacent text blocks is greater than the second distance threshold. If the distance is greater than the second distance threshold, a line is drawn to separate the left and right adjacent text blocks. Repeated vertical and horizontal determination and division can be performed to obtain at least one text object partition.
[0140] It is understandable that the first distance threshold and the second distance threshold can be set according to specific application scenarios and are not limited here.
[0141] In some embodiments, the method further comprises:
[0142] Determining a layout of the target document based on geometric information of the non-text object partitions and / or text object partitions determined in the target document page, and determining an order of the partitions based on the layout;
[0143] For each text object partition, multiple text blocks are sorted according to position information of the text blocks.
[0144] After the text object partitions are obtained, the order of the multiple text objects is determined according to the position information of the text object partitions.
[0145] In some embodiments, the above step S102 includes:
[0146] For each text object partition, multiple text blocks in the text object partition are segmented according to position information and / or semantic information of the text blocks to obtain multiple text paragraphs.
[0147] In some application scenarios, for each text object partition, the text object partition may include multiple text blocks. It is possible to determine whether the multiple text blocks are in the same row based on the position information of the rectangular boxes of the multiple text blocks. If they are in the same row, these text blocks are spliced to obtain a row of text blocks. If there are adjacent rows of text blocks below, the distance between the adjacent rows and the left-right alignment information are determined based on the position information of the rectangular boxes corresponding to the text blocks in the adjacent rows. If the line spacing is less than a certain distance threshold, determine the coordinates of the leftmost text block of the adjacent row and whether it is to the left of the leftmost coordinate corresponding to the row or whether it is aligned with the leftmost coordinate corresponding to the row.
[0148] If the text block is to the left of the leftmost coordinate of the row or aligned with the leftmost coordinate of the row, the adjacent row is divided into the same text paragraph as the row. Otherwise, a new paragraph is created for the adjacent row text block. Similarly, the text blocks in the text object partition can be spliced into multiple text paragraphs.
[0149] In some application scenarios, the semantic information of different text block combinations is analyzed, and the text blocks representing a complete semantic content are spliced together to obtain a text paragraph.
[0150] In some embodiments, for each text object partition, multiple text blocks in the text object partition are segmented according to the position information and / or semantic information of the text blocks to obtain multiple text paragraphs, including:
[0151] For adjacent paragraphs in two different text object partitions, whether to splice the adjacent paragraphs is determined based on at least two paragraphs corresponding to each of the two text object partitions, where the at least two paragraphs corresponding to each of the two text object partitions include the adjacent paragraphs.
[0152] The two different text object partitions may be two text object partitions that are adjacent across columns, or two text object partitions that are adjacent across pages.
[0153] In two adjacent text object partitions, the last paragraph of the preceding text object partition and the first paragraph of the succeeding text object partition are adjacent paragraphs. Whether the two adjacent paragraphs need to be spliced is determined based on the last at least two paragraphs in the preceding text object partition and the first at least two paragraphs in the succeeding text object partition, thereby determining whether to splice the two adjacent paragraphs in different partitions.
[0154] For example, for the last paragraph S in the previous text object partition Q1 K , get the second to last paragraph S in the text object partition Q1 K-1 , and the third to last paragraph S K-2 . Determine the paragraphs S in the text object partition Q1 K-1、 S K-2 The leftmost coordinate and the rightmost coordinate. For the first paragraph W1 in the next text object partition Q2, the second paragraph W2 and the third paragraph W3 in the text object partition Q2 are obtained. 2、 W3 leftmost and rightmost coordinates. If paragraph S K The leftmost coordinate of the last line is the same as paragraph S K-1 、S K-2 The leftmost coordinates are aligned, and the leftmost coordinates of the first paragraph W1, the second paragraph W2 and the third paragraph W3 in the text object partition Q2 are aligned. k It needs to be spliced with paragraph W1, and then determine paragraph S k If the semantics of the two paragraphs are consistent with those of paragraph W1, the two paragraphs are concatenated into one paragraph.
[0155] In some embodiments, the above step S101 includes:
[0156] The non-text object partitions and the text object partitions of the target document are parsed in the order of first the non-text object partitions and then the text object partitions.
[0157] In these embodiments, non-text objects in the target document may be identified first, and then non-text object partitions corresponding to the non-text objects may be determined.
[0158] After parsing out the non-text object partition, the text object partition including the text object is parsed.
[0159] Since non-text objects may include text content, parsing the non-text object partitions first and then parsing the text object partitions can avoid incorrectly dividing the text content in the non-text objects into text object partitions, thereby obtaining more accurate text object partitioning and non-text object partitioning results.
[0160] In some embodiments, parsing the non-text object partitions and the text object partitions of the target document in the order of first the non-text object partitions and then the text object partitions includes:
[0161] The image partition corresponding to the image, the table partition corresponding to the table, and / or the formula partition corresponding to the formula in the target document are parsed in sequence.
[0162] In these embodiments, if the target document includes pictures, tables, and formulas, the non-text object partitions in the target document may be parsed in the order of first picture partitioning, then table partitioning, and finally formula partitioning.
[0163] In these embodiments, by first identifying the image, it is possible to avoid subsequent misrecognition of tables and formulas due to the content in the image. By first identifying the image and table, and then performing formula recognition, it is possible to avoid misrecognition of formulas due to the content of the image or table. This further improves the accuracy of non-text object recognition, thereby improving the accuracy of non-text object partitioning.
[0164] In some embodiments, the method further comprises the steps of:
[0165] First, according to the received first content generation instruction, a target paragraph corresponding to the content generation instruction is determined from a plurality of paragraphs.
[0166] Secondly, task information is generated according to the first content generation instruction and the target paragraph.
[0167] Finally, the first model is called so that the first model generates first content corresponding to the first content generation instruction according to the task information.
[0168] After obtaining the semantic tags of each paragraph of the target document, the semantic tags of each paragraph and the corresponding text content can be associated and stored, and a vector index can be established.
[0169] The first content generation instruction may be an instruction to generate various contents, such as an instruction to generate a reply to a question in the first content generation instruction, or an instruction to generate a summary of a document, or an instruction to rewrite the content of a document.
[0170] After receiving the first content generation instruction, the first content generation instruction may be matched with multiple paragraphs in the target document, and the target paragraph may be determined according to the matching result.
[0171] The first content generation instruction and the target paragraph content can be used to generate task information. For example, the first content generation instruction and the target paragraph content can be written into a prompt template to generate task information. The target content can include non-text objects, such as tables, images, or formulas.
[0172] The first model may be called through a preset interface, and the first model processes the target paragraph according to the task information to obtain the first content matching the first content generation instruction.
[0173] The first model may be a language model, which may be a model pre-trained using massive training corpus and having preliminary reasoning capabilities.
[0174] In these embodiments, the first content is generated by using a target paragraph in the target document that matches the first content generation instruction. Since the target paragraph is a paragraph among multiple paragraphs obtained by accurately identifying and restoring objects in the target document, the determined target paragraph has a relatively high accuracy and does not contain redundant information. Therefore, the first content that matches the first content generation instruction can be obtained relatively quickly and accurately.
[0175] In some embodiments, the method further comprises the steps of:
[0176] First, a second content generation instruction is received;
[0177] Secondly, calling a second model to generate second content according to a second content instruction by the second model, wherein the second content is generated by a second target paragraph that matches the second content generation instruction and is determined by the second model from a plurality of paragraphs based on the second content instruction;
[0178] The second content generation instruction may be an instruction to generate various contents, such as an instruction to generate a reply to a question in the second content generation instruction, or an instruction to generate a summary of a document, or an instruction to rewrite the content of a document.
[0179] The second content generation instruction may be sent to a second model. The second model may be a language model. The language model may be a model that has been pre-trained using a large amount of training corpus and has preliminary reasoning capabilities.
[0180] The second model can obtain multiple paragraphs corresponding to the target document based on the second content generation instruction. The model can then determine a target paragraph from the multiple paragraphs based on the matching results with the second content instruction. The model can then generate content based on the target paragraph according to the second content instruction to generate the second content.
[0181] For example, if the target document contains a picture A of animal A, a table B describing the food it eats, and a paragraph C describing animal A's daily life, and the second content generation instruction could be "please summarize the appearance characteristics of animal A," the second model can determine that picture A is the corresponding target paragraph based on the second content generation instruction and generate a description of animal A's appearance based on picture A.
[0182] In these embodiments, by sending the second content generation instruction to the second model, and having the second model obtain the target paragraph in the target document according to the second content generation instruction to generate the second content, more second content than the second content generation instruction can be obtained.
[0183] Please refer to Figure 8, which shows a schematic flow chart of the content generation method provided by the present disclosure. As shown in Figure 8, the method includes the following steps:
[0184] S801: Receive a third content generation instruction.
[0185] The execution subject of the above content generation method can be a client running in a terminal device, or a server that provides services to the client.
[0186] The third content generation instruction may be an instruction for generating various contents, such as an instruction for generating a reply to a question in the third content generation instruction, or an instruction for generating a summary of a document, or an instruction for rewriting the content of a document.
[0187] The user can send the third content instruction to the execution subject through the instruction input window in the client.
[0188] In some application scenarios, the client interface may display the content of the target document. The target document may be a document in a preset format as in the embodiments shown in FIG1 and FIG3 , which will not be described in detail here.
[0189] S802: Generate third content corresponding to the third content generation instruction based on the third content generation instruction and a target paragraph from a plurality of paragraphs extracted from a target document in a preset format; wherein the target document includes text objects and non-text objects, and the plurality of paragraphs are obtained based on the following steps: parsing at least one non-text object partition and at least one text object partition from the target document; and performing paragraph splicing on the non-text objects and text objects in the non-text object partition and the text object partition based on the semantic information and geometric information of the objects, respectively, to obtain a plurality of paragraphs.
[0190] After receiving the third content generation instruction, the target paragraph may be determined from the multiple paragraphs corresponding to the target document, for example, based on a degree of matching with the third content generation instruction.
[0191] In some embodiments, multiple paragraphs in the target document may be pre-parsed according to the methods of the embodiments shown in FIG. 1 and FIG. 3 to FIG. 7 .
[0192] In some other embodiments, after receiving the third content instruction, the target document indicated by the third content instruction may be parsed according to the methods of the embodiments shown in FIG. 1 and FIG. 3 to FIG. 7 to parse out multiple paragraphs with semantic tags.
[0193] The execution entity may generate third content matching the third content instruction according to the target paragraph.
[0194] In some embodiments, step S802 includes: sending a third content generation instruction to a third model, so that the third model determines a target paragraph matching the third content from multiple paragraphs of the target document according to the third instruction, and generates the third content according to the target paragraph.
[0195] In these embodiments, the third content generation instruction may be sent to a third model, which may be a language model.
[0196] The third model can obtain multiple paragraphs corresponding to the target document based on the third content generation instruction. For example, the target document can be parsed according to the method of the embodiments shown in Figures 1 and 3 to 7 to obtain multiple paragraphs with semantic tags. A target paragraph is then determined from the multiple paragraphs based on the degree of match with the third content instruction. Content generation is then performed based on the target paragraph to generate the third content.
[0197] In some embodiments, step S802 includes the following steps:
[0198] First, a target paragraph matching the third content generation instruction is determined from a plurality of paragraphs extracted from a target document in a preset format.
[0199] Secondly, a content generation request including a third content generation instruction and a target paragraph is sent to the third model, and the third model generates the third content according to the target paragraph.
[0200] The target document may be parsed according to the methods of the embodiments shown in FIG. 1 and FIG. 3 to FIG. 7 to obtain multiple paragraphs with semantic tags.
[0201] After obtaining the semantic tags of each paragraph of the target document, the semantic tags of each paragraph and the corresponding text content can be associated and stored, and a vector index can be established.
[0202] After receiving the third content generation instruction, the first content generation instruction may be matched with multiple paragraphs from the target document, and the target paragraph may be determined according to the matching result.
[0203] A content generation request may be generated according to the third content generation instruction and the content of the target paragraph, and the content generation request is sent to the third model, which generates the third content according to the content of the target paragraph.
[0204] In some embodiments, the method further comprises:
[0205] Display third content.
[0206] In these embodiments, after the third content is obtained, the third content may be displayed for the user to read.
[0207] In this embodiment, after receiving the third content generation instruction, the third content is generated based on the third content generation instruction and the target paragraph in the target document that matches the first content generation instruction. Since the target paragraph is a paragraph among multiple paragraphs obtained by accurately identifying and restoring objects in the target document, the accuracy of the determined target paragraph is relatively high and does not contain redundant information. Therefore, the third content that matches the third content generation instruction can be obtained relatively quickly and accurately.
[0208] Corresponding to the document processing method of the embodiment shown in FIG1 above, FIG9 is a schematic structural diagram of a document processing device provided by an embodiment of the present disclosure. For ease of explanation, only the parts related to the embodiment of the present disclosure are shown. Referring to FIG9, the document processing device 90 includes: a parsing unit 901, a first splicing unit 902, and a first generating unit 903.
[0209] The parsing unit 901 is configured to, in response to receiving a first instruction for a target document in a preset format, parse the non-text object partitions and the text object partitions in the target document to obtain at least one non-text object partition and at least one text object partition;
[0210] A splicing unit 902 is configured to perform paragraph splicing on the non-text objects and text objects in the non-text object partition and the text object partition based on the semantic information and geometric information of the objects, respectively, to obtain multiple paragraphs;
[0211] The first generating unit 903 is configured to generate a semantic tag of the paragraph according to the position information and semantic information of the paragraph.
[0212] In some embodiments, the non-text object includes a non-vector image object, and the parsing unit 901 is further configured to parse the target document to obtain a first image partition corresponding to the non-vector image object based on the following steps:
[0213] For a non-vector image, obtaining first position information corresponding to the non-vector image based on an underlying protocol of the target document, where the first position information is position information of a first image partition;
[0214] A non-vector image is intercepted in the target document based on the first position information to obtain a non-vector image corresponding to the first image partition.
[0215] In some embodiments, the non-text object includes a vector graphic object, and the parsing unit 901 is further configured to parse the target document to obtain a second image partition corresponding to the vector graphic object based on the following steps:
[0216] Hide the text content in the target document;
[0217] In the target document with hidden text content, identifying second position information corresponding to the vector image and the first vector image without text content; the second position information is used as position information of the second image partition;
[0218] determining text content in the first vector image according to the second position information;
[0219] The first vector image and the text content are merged to form a vector image corresponding to the second image partition.
[0220] In some embodiments, the non-text object includes a wireframe table object, and the parsing unit 901 is further configured to parse the target document to obtain a first table partition corresponding to the wireframe table object based on the following steps:
[0221] Obtaining third position information of a rectangular frame surrounding the wireframe table and information of each first cell in the wireframe table, wherein the third position information is position information of a partition of the first table;
[0222] Merging the plurality of first cells according to information of adjacent cells of each first cell to obtain a plurality of merged second cells;
[0223] determining a size of a minimum cell among the plurality of second cells, and determining a proportional relationship between the other cells and the minimum cell;
[0224] Generate a blank wireframe table corresponding to the first table partition according to the proportional relationship;
[0225] The text content rectangular frame corresponding to the third position information in the target document is merged with the blank wired frame table to obtain the wired frame table corresponding to the first table partition.
[0226] In some embodiments, the non-text object includes a half-line frame table object, and the parsing unit 901 is further configured to parse the second table partition corresponding to the half-line frame table object in the target document based on the following steps:
[0227] Based on the projection segmentation algorithm and the image recognition algorithm, the frame lines used to define the semi-wireframe table are identified;
[0228] Determining a pair of frame lines for defining a half-frame according to the position and size of the identified frame lines;
[0229] Using the fourth position information determined by the frame line pair as the position information of the second table partition;
[0230] Obtaining the semi-line frame table content corresponding to the fourth position information from the target document;
[0231] The content of the half-line frame table is written into the half-line frame table to obtain the half-line frame table corresponding to the second table partition.
[0232] In some embodiments, the non-text object includes a formula object, and the parsing unit 901 is further configured to parse the formula partition corresponding to the formula object in the target document based on the following steps:
[0233] Use regular expressions and preset formula fonts to identify formula content, splice the formula content, and obtain the formula in the target document.
[0234] In some embodiments, the parsing unit 901 is further configured to:
[0235] Acquire multiple text blocks corresponding to text content of the target document outside the non-text object partition;
[0236] Based on the projection segmentation algorithm, multiple text blocks are partitioned to obtain at least one text object partition.
[0237] In some embodiments, the parsing unit 901 is further configured to:
[0238] The layout form of the target document is determined according to the geometric information of the multiple partitions determined in the target document page, and the order of the partitions is determined according to the layout form.
[0239] In some embodiments, the splicing unit 902 is further configured to:
[0240] For each text object partition, multiple text blocks in the text object partition are segmented according to position information and / or semantic information of the text blocks to obtain multiple text paragraphs.
[0241] In some embodiments, the splicing unit 902 is further configured to:
[0242] For adjacent paragraphs in two different text object partitions, whether to splice the adjacent paragraphs is determined based on at least two paragraphs corresponding to each of the two text object partitions, where the at least two paragraphs corresponding to each of the two text object partitions include the adjacent paragraphs.
[0243] In some embodiments, the parsing unit 901 is further configured to:
[0244] The non-text object partitions and the text object partitions of the target document are parsed in the order of first the non-text object partitions and then the text object partitions.
[0245] In some embodiments, the parsing unit 901 is further configured to:
[0246] The image partition corresponding to the image, the table partition corresponding to the table, and the formula partition corresponding to the formula in the target document are parsed in sequence.
[0247] In some embodiments, the apparatus 90 further includes a first content generation unit (not shown). The first content generation unit is configured to:
[0248] generating task information based on the first content generation instruction and the target paragraph;
[0249] The first model is called to generate first content corresponding to the first content generation instruction according to the task information.
[0250] In some embodiments, the apparatus 90 further includes a second content generation unit (not shown in the figure). The second content generation unit is configured to:
[0251] the received second content generation instruction;
[0252] The second model is called to generate second content according to the second content instruction by the second model, wherein the second content is generated by a second target paragraph that matches the second content generation instruction and is determined by the second model from a plurality of paragraphs based on the second content instruction.
[0253] Corresponding to the content generation method of the embodiment shown in FIG7 above, FIG10 is a schematic structural diagram of the content generation device provided by the embodiment of the present disclosure. For the sake of convenience, only the parts related to the embodiment of the present disclosure are shown. Referring to FIG10, the content generation device 100 includes: a receiving unit 1001 and a second generating unit 1002.
[0254] The receiving unit 1001 is configured to receive a third content generation instruction;
[0255] The second generation unit 1002 is configured to generate third content corresponding to the third content generation instruction based on the content generation instruction and a target paragraph from a plurality of paragraphs extracted from a target document in a preset format; wherein the target document includes text objects and non-text objects, and the plurality of paragraphs are obtained based on the following steps: parsing at least one non-text object partition and at least one text object partition from the target document; and performing paragraph splicing on the non-text objects and text objects in the non-text object partition and the text object partition based on the semantic information and geometric information of the objects, respectively, to obtain a plurality of paragraphs.
[0256] In some embodiments, the second generating unit 1002 is further configured to:
[0257] The third content generation instruction is sent to the third model, so that the third model determines a target paragraph matching the third content from multiple paragraphs in the target document according to the third instruction, and generates the third content according to the target paragraph.
[0258] In some embodiments, the second generating unit 1002 is further configured to:
[0259] determining a target paragraph matching the third content generation instruction from a plurality of paragraphs extracted from a target document in a preset format;
[0260] A content generation request including a third content generation instruction and a target paragraph is sent to a third model, and the third model generates third content according to the target paragraph.
[0261] In some embodiments, the apparatus 100 further includes a display unit (not shown). The display unit is configured to:
[0262] Display third content.
[0263] In order to implement the above embodiment, the present disclosure further provides an electronic device.
[0264] Referring to FIG11 , a schematic diagram of the structure of an electronic device 1100 suitable for implementing an embodiment of the present disclosure is shown. The electronic device 1100 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG11 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0265] As shown in FIG11 , the electronic device 1100 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage device 1108 into a random access memory (RAM) 1103. Various programs and data required for the operation of the electronic device 1100 are also stored in the RAM 1103. The processing device 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0266] Typically, the following devices may be connected to the I / O interface 1105: an input device 1106 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1107 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1108 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1109. The communication device 1109 may allow the electronic device 1100 to communicate with other devices wirelessly or by wire to exchange data. Although FIG11 shows an electronic device 1100 having various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0267] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 1109, or installed from the storage device 1108, or installed from the ROM 1102. When the computer program is executed by the processing device 1101, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0268] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0269] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0270] The computer-readable medium carries one or more programs. When the one or more programs (computer-executable instructions) are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0271] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0272] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0273] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.
[0274] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0275] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0276] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0277] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0278] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A document processing method, comprising: In response to receiving a first instruction for a target document in a preset format, parsing non-text object partitions and text object partitions in the target document to obtain at least one non-text object partition and at least one text object partition; Based on the semantic information and geometric information of the objects, paragraph splicing is performed on the non-text objects and text objects in the non-text object partition and the text object partition respectively to obtain multiple paragraphs; Generate a semantic tag for the paragraph based on the location information and semantic information of the paragraph.
2. The method according to claim 1, wherein: The non-text object includes a non-vector image object, and the parsing of the non-text object partition and the text object partition in the target document to obtain at least one non-text object partition and at least one text object partition includes parsing a first image partition corresponding to the non-vector image object in the target document based on the following steps, wherein the first image partition belongs to the non-text object partition: For a non-vector picture, obtaining first position information corresponding to the non-vector picture based on an underlying protocol of the target document, wherein the first position information is position information of the first picture partition; The non-vector picture is intercepted in the target document based on the first position information to obtain the non-vector picture corresponding to the first picture partition.
3. The method according to claim 1, wherein: The non-text object includes a vector image object, and the parsing of the non-text object partitions and the text object partitions in the target document to obtain at least one non-text object partition and at least one text object partition includes parsing a second image partition corresponding to the vector image object in the target document based on the following steps: Hiding text content in the target document; In the target document in which the text content is hidden, second position information corresponding to the vector image and a first vector image without text content are identified; the second position information is used as position information of a second image partition; determining text content in the first vector image according to the second position information; The first vector image and the text content are merged to form a vector image corresponding to the second image partition.
4. The method according to claim 1, wherein: The non-text object includes a wireframe table object, and the parsing of the non-text object partitions and the text object partitions in the target document to obtain at least one non-text object partition and at least one text object partition includes parsing a first table partition corresponding to the wireframe table object in the target document based on the following steps: Acquire third position information of a rectangular frame surrounding the wireframe table, and information of each first cell in the wireframe table, wherein the third position information is position information of a first table partition; Merge multiple first cells according to the information of the adjacent cells of each first cell to obtain multiple merged second cells; Determine the size of the smallest cell among the plurality of second cells, and determine a proportional relationship between other cells and the smallest cell; Generate a blank wireframe table corresponding to the first table partition according to the proportional relationship; The text content rectangular frame corresponding to the third position information in the target document is merged with the blank wired frame table to obtain a wired frame table corresponding to the first table partition.
5. The method according to claim 1, wherein: The non-text object includes a semi-wireframe table object, and the parsing of the non-text object partitions and the text object partitions in the target document to obtain at least one non-text object partition and at least one text object partition includes parsing a second table partition corresponding to the semi-wireframe table object in the target document based on the following steps: Based on the projection segmentation algorithm and the image recognition algorithm, the frame lines used to define the semi-wireframe table are identified; Determine a pair of frame lines for defining a half-frame according to the position and size of the identified frame lines; Using the fourth position information determined by the frame line pair as the position information of the second table partition; Acquire the semi-line frame table content corresponding to the fourth position information from the target document; The content of the half-line frame table is written into the half-line frame table to obtain the half-line frame table corresponding to the second table partition.
6. The method according to claim 1, wherein: The non-text object includes a formula object, and the parsing of the non-text object partitions and the text object partitions in the target document to obtain at least one non-text object partition and at least one text object partition includes parsing the formula partition corresponding to the formula object in the target document based on the following steps: Use regular expressions and preset formula fonts to identify formula content, concatenate the formula content, and obtain the formula in the target document.
7. The method according to claim 1, wherein: The step of parsing the non-text object partitions and the text object partitions in the target document to obtain at least one non-text object partition and at least one text object partition includes: Acquire multiple text blocks corresponding to the text content of the target document outside the non-text object partition; Based on the projection segmentation algorithm, multiple text blocks are partitioned to obtain at least one text object partition.
8. The method according to claim 7, further comprising: The layout of the target document is determined according to the geometric information of the multiple partitions determined in the target document page, and the order of the partitions is determined according to the layout.
9. The method according to claim 1, wherein: The method of performing paragraph splicing on the non-text object partition and the text object in the text object partition based on the semantic information and geometric information of the object, respectively, obtains multiple paragraphs, including: For each text object partition, multiple text blocks in the text object partition are segmented according to position information and / or semantic information of the text blocks to obtain multiple text paragraphs.
10. The method according to claim 9, wherein: For each text object partition, the multiple text blocks in the text object partition are segmented according to the position information and / or semantic information of the text blocks to obtain multiple text paragraphs, and further includes: For adjacent paragraphs in two different text object partitions, the corresponding Whether to splice the adjacent paragraph is determined based on at least two paragraphs corresponding to each of the two text object partitions, and the at least two paragraphs corresponding to each of the two text object partitions include the adjacent paragraph.
11. The method according to claim 1, wherein: The parsing of at least one non-text object partition and at least one text object partition in the target document comprises: The non-text object partitions and the text object partitions of the target document are parsed in the order of first the non-text object partitions and then the text object partitions.
12. The method according to claim 11, wherein: The non-text object partition and the text object partition of the target document are parsed in the order of first the non-text object partition and then the text object partition, including parsing the following non-text object partitions: The image partition corresponding to the image, the table partition corresponding to the table, and / or the formula partition corresponding to the formula in the target document are parsed in sequence.
13. The method according to claim 1, further comprising: According to the received first content generation instruction, determining a target paragraph corresponding to the content generation instruction from the plurality of paragraphs; generating task information according to the first content generation instruction and the target paragraph; The first model is called so that the first model generates first content corresponding to the first content generation instruction according to the task information.
14. The method according to claim 1, further comprising: The received second content generation instruction; The second model is called to generate second content according to the second content instruction by the second model, wherein the second content is generated by a second target paragraph that matches the second content generation instruction and is determined by the second model from the plurality of paragraphs based on the second content instruction.
15. A content generation method, comprising: receiving a third content generation instruction; generating, based on the third content generation instruction and a target paragraph from a plurality of paragraphs extracted from a target document in a preset format, a third content corresponding to the third content generation instruction; Wherein, the target document includes text objects and non-text objects, and the multiple paragraphs are obtained based on the following steps: parsing at least one non-text object partition and at least one text object partition from the target document; based on the semantic information and geometric information of the objects, paragraph splicing of the non-text objects and text objects in the non-text object partition and the text object partition is performed to obtain multiple paragraphs.
16. The method according to claim 15, wherein: The step of generating the third content corresponding to the third content generating instruction based on the content generating instruction and a target paragraph extracted from a plurality of paragraphs in a target document of a preset format comprises: The third content generation instruction is sent to a third model, so that the third model determines a target paragraph matching the third content from a plurality of paragraphs in the target document according to the third instruction, and generates the third content according to the target paragraph.
17. The method according to claim 15, wherein: The step of generating the third content corresponding to the third content generating instruction based on the content generating instruction and a target paragraph extracted from a plurality of paragraphs in a target document of a preset format comprises: Determining a target paragraph matching the third content generation instruction from a plurality of paragraphs extracted from a target document in a preset format; A content generation request including the third content generation instruction and the target paragraph is sent to a third model, and the third model generates the third content according to the target paragraph.
18. The method according to any one of claims 15 to 17, further comprising: The third content is displayed.
19. A document processing device, comprising: A parsing unit, configured to parse non-text object partitions and text object partitions in the target document in response to receiving a first instruction for a target document in a preset format, to obtain at least one non-text object partition and at least one text object partition; A splicing unit is configured to perform paragraph splicing on the non-text objects and text objects in the non-text object partition and the text object partition respectively based on the semantic information and geometric information of the objects to obtain a plurality of paragraphs; as well as The first generating unit is configured to generate a semantic tag of the paragraph according to the position information and semantic information of the paragraph.
20. A content generation device, comprising: A receiving unit, configured to receive a third content generation instruction; as well as The second generation unit is configured to generate a third content corresponding to the third content generation instruction based on the content generation instruction and a target paragraph from a plurality of paragraphs extracted from a target document in a preset format; wherein the target document includes text objects and non-text objects, and the plurality of paragraphs are obtained based on the following steps: parsing at least one non-text object partition and at least one text object partition from the target document; and performing paragraph splicing on the non-text objects and text objects in the non-text object partition and the text object partition respectively based on the semantic information and geometric information of the objects to obtain a plurality of paragraphs.
21. An electronic device comprising: Processor and memory; Wherein, the memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the document processing method according to any one of claims 1 to 14 or the content generating method according to any one of claims 15 to 18.
22. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer-executable instructions. When the processor executes the computer-executable instructions, the document processing method according to any one of claims 1 to 14 or the content generation method according to any one of claims 15 to 18 is implemented.
Citation Information
Patent Citations
Document processing method, electronic equipment and storage medium
CN114118011A
Method and system for extracting information of PDF (Portable Document Format) document in security and futures scene
CN114821612A
Document processing method and device, electronic equipment and storage medium
CN115759039A
Document processing method and device, content generation method and device and electronic equipment
CN117610549A
A heuristic method for analyzing content of an electronic document
US20200364452A1
Cited By
Hybrid PDF (Portable Document Format) document analysis and knowledge fragment construction method and device
CN121527785A
Certificate report generation method and device based on multi-modal large model
CN122366390A