Research report generation method and device, electronic equipment and computer readable medium
Through the method of segment-by-stage split processing and chart interception, the problem of excessive computing resource occupancy in research report form generation is solved, and the effects of resource conservation and data accuracy are achieved.
Patent Information
- Application Number
- CN202510144886.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-10
AI Technical Summary
During the process of generating research report forms, when there is a lot of document content, full text analysis and image analysis are required, resulting in excessive computing resources occupied.
Processing fixed documents by segmentation, splitting them into small paragraphs, and extracting chart coordinates for chart intercepting, and integrating document span checksum content to generate integrated document data sequences.
Reduce the use of computing resources, avoid data loss, and ensure the accuracy of generated document forms, avoid the waste of printing resources.
Smart Images

Figure CN120106028A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of document recognition technology, and in particular to a research report generation method, device, electronic device and computer-readable medium. Background Art
[0002] The generation of research reports often requires more accurate document recognition technology. At present, when generating forms, the usual method is to parse the content of the entire document, identify the document semantics and image semantics, and extract the document content.
[0003] However, in practice, it is found that when the above method is used to generate research reports, the following technical problems often occur:
[0004] The document contains a lot of content. If all the text and image parsing are performed, a lot of computing resources will be required.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention
[0006] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.
[0007] Some embodiments of the present disclosure propose research report generation methods, devices, electronic devices and computer-readable media to solve one or more of the technical problems mentioned in the above background technology section.
[0008] In a first aspect, some embodiments of the present disclosure provide a method for generating a research report, the method comprising: obtaining a fixed document from a database; splitting the fixed document segment by segment according to the page number sequence of the fixed document to generate a split document data sequence set, wherein each split document data in the split document data sequence set corresponds to a section of content in the fixed document; extracting chart features from the fixed document to obtain a chart feature information sequence, wherein each chart feature information in the chart feature information sequence includes a chart coordinate group; performing chart interception on the fixed document according to the chart feature information sequence to obtain a document chart sequence; performing document cross-page verification on each split document data sequence in the split document data sequence set to generate a document cross-page identification sequence, and for the split document data sequence corresponding to the document cross-page identification representing the cross-page of document content in the document cross-page identification sequence, performing document content cross-page integration on the split document data sequence to obtain an integrated document data sequence; constructing a document form corresponding to the fixed document based on the document chart sequence and each integrated document data sequence for storage.
[0009] In a second aspect, some embodiments of the present disclosure provide a research report generation device, which includes: an acquisition unit, configured to acquire a fixed document from a database; a splitting processing unit, configured to split the fixed document segment by segment according to the page number sequence of the fixed document to generate a split document data sequence set, wherein each split document data in the split document data sequence set corresponds to a section of content in the fixed document; a feature extraction unit, configured to extract chart features from the fixed document to obtain a chart feature information sequence, wherein each chart feature information in the chart feature information sequence includes a chart coordinate group; a chart interception unit, configured The fixed document is intercepted according to the chart feature information sequence to obtain a document chart sequence; the document cross-page verification and integration unit is configured to perform document cross-page verification on each split document data sequence in the split document data sequence set to generate a document cross-page identification sequence, and for the split document data sequence corresponding to the document cross-page identification representing the document content cross-page in the document cross-page identification sequence, perform document content cross-page integration on the split document data sequence to obtain an integrated document data sequence; the construction unit is configured to construct a document form corresponding to the fixed document based on the document chart sequence and each integrated document data sequence for storage.
[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the above-mentioned first aspect.
[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the above-mentioned first aspect is implemented.
[0012] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the research report generation method of some embodiments of the present disclosure, the occupation of computing resources can be reduced. Specifically, the reason why a large amount of computing resources needs to be occupied is that the document content is large, and if all text parsing and image parsing are performed, a large amount of computing resources need to be occupied. Based on this, the research report generation method of some embodiments of the present disclosure, first of all, considering that splitting the document into lines for recognition or semantic recognition consumes more computing resources, the fixed document is split into small paragraphs by segment splitting processing. Thus, compared with splitting into finer lines, the occupation of computing resources can be reduced. Secondly, by extracting the coordinates of the chart, it can be used to locate the position of the chart in the document. Thus, it can be used to intercept the chart. In addition, considering that the text of the document spans pages, the text of the same paragraph may be divided into different paragraphs. Therefore, through the document cross-page check, it can be identified whether there is a situation where the text spans pages. Then, through the cross-page integration of the document content, the same paragraph of text across pages can be integrated into the same paragraph. Thus, not only can the loss of data be avoided, but also the occupation of computing resources can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0014] Figure 1 is a flow chart of some embodiments of the method for generating a research report form according to the present disclosure;
[0015] Figure 2 It is a schematic diagram of a document content extraction process according to some embodiments of the research report form generation method disclosed in the present invention;
[0016] Figure 3 It is a structural schematic diagram of some embodiments of the research report form generating device according to the present disclosure;
[0017] Figure 4 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0019] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0020] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0021] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0024] Figure 1 The process 100 of some embodiments of the method for generating a research report according to the present disclosure is shown. The method for generating a research report includes the following steps:
[0025] Step 101, obtaining a fixed document from a database.
[0026] In some embodiments, the execution subject of the research report generation method can obtain a fixed document from a database by wired or wireless means. The database can be a pre-set unit for storing documents uploaded by users. Secondly, the fixed document can be a document with a fixed structure format. For example, a PDF (Powder Diffraction File) document.
[0027] It should be noted that the above-mentioned wireless connection methods may include but are not limited to 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future.
[0028] Step 102, splitting the fixed document into sections according to the page number sequence of the fixed document to generate a set of document data sequences after splitting.
[0029] In some embodiments, the execution entity may split the fixed document segment by segment according to the page number sequence of the fixed document to generate a split document data sequence set, wherein each split document data in the split document data sequence set may correspond to a segment of content in the fixed document.
[0030] In some optional implementations of some embodiments, the execution subject performs segmentation processing on the fixed document according to the page number sequence of the fixed document to generate a segmented document data sequence set, which may include the following steps:
[0031] The first step is to identify the page numbers of each page in the fixed document. The page numbers of each page in the fixed document can be identified by using a pre-set PDF page number recognition function.
[0032] As an example, you can use the VBA (Visual Basic for Applications) algorithm to read the binary content in the PDF file to identify the page number. The specific steps are as follows: First, open the PDF file in binary mode. Then, read the file content and use a regular expression to match the number after " / Count". Finally, extract the last match and remove " / Count" to get the page number.
[0033] The second step is to split the fixed document page by page according to the order of the page numbers to obtain a document split page sequence.
[0034] The third step is to identify the paragraph information in each document split page in the above-mentioned document split page sequence, and to split the document split page into paragraphs according to the paragraph information to generate a set of split document data sequences. Among them, the indentation value of the first word in each line can be identified. If the indentation is 2 characters, the line is determined to be the starting line of a paragraph. Secondly, the text content from one starting line to the next starting line can be determined as a paragraph of text. Finally, the paragraph of text can be determined as paragraph information according to the paragraph order. Thus, the position of each paragraph can be determined. Afterwards, the paragraph splitting can be performed according to the paragraph position in the paragraph information, that is, the text of each paragraph is determined as the split document data. Here, each split document data sequence can correspond to a page of text in the above-mentioned fixed document.
[0035] As an example, you can use libraries such as pdfplumber and fitz to extract the text of each paragraph in a document split.
[0036] In practice, PyMuPDF is comprehensive in functionality. It can not only extract PDF content, but also edit, delete, and add DF content. However, the extracted content format only supports text and inserted images. The PDFMiner3k algorithm has extremely high extraction efficiency. Through a specific delayed parsing mechanism, it can maximize the efficiency and effect of parsing, but it has almost no ability to parse images and tables. In addition, the processing of cross-page content and tables in PDF is also a common problem in extraction.
[0037] In the process of adopting technical solutions to solve the problems mentioned in the background technology, the following technical problem 2 is often accompanied: there are often many cross-page texts and cross-page charts in the document. If formatted one-to-one extraction is performed, it is easy to split the same paragraph of text or the same chart into multiple ones, which not only leads to document semantic errors and missing, but also document typesetting errors. Therefore, when printing a form, the wrong form is printed, resulting in a waste of printing resources. In response to the above technical problem 2, the inventor decided to adopt the following solution.
[0038] Step 103: extract chart features from the fixed document to obtain a chart feature information sequence.
[0039] In some embodiments, the execution entity may extract chart features from the fixed document to obtain a chart feature information sequence.
[0040] In some optional implementations of some embodiments, the execution subject extracts chart features from the fixed document to obtain a chart feature information sequence, which may include the following steps:
[0041] The first step is to extract chart labels from each of the split document data in the above split document data sequence set to generate a picture label sequence and a table label sequence. Each picture label in the above picture label sequence may include: picture sequence number label coordinates and picture source label coordinates. Each table label in the above table label sequence may include: table sequence number label coordinates and table source coordinates. Here, the PDF parsing algorithm can be used to extract chart labels from each of the split document data in the above split document data sequence set to generate a picture label sequence and a table label sequence.
[0042] As an example, the PDF parsing algorithm may include but is not limited to at least one of the following: PDFPageInterpreter interpreter, PDFParser parsing library.
[0043] In the second step, for each image tag in the above image tag sequence, perform the following image extraction steps:
[0044] Step 1: perform label matching on the above picture label and other picture labels in the above picture label sequence to generate a picture matching label. The label matching is performed by the following steps:
[0045] First, determine the coordinate interval between the coordinates of the image sequence number tag and the coordinates of the image source tag in the above image tag. The coordinate interval may be the interval between the minimum vertical coordinate and the maximum vertical coordinate between the coordinates of the image sequence number tag and the coordinates of the image source tag. Here, the vertical coordinate may be the vertical coordinate of the document.
[0046] Then, in response to the existence of image sequence number tag coordinates and / or image source tag coordinates in the above coordinate interval in other image tags, it is determined to match the above image tag. Wherein, the existence of image sequence number tag coordinates and / or image source tag coordinates in the above coordinate interval indicates that the image corresponding to the image sequence number tag coordinates and / or image source tag coordinates and the image corresponding to the above image tag are in the same section in the document.
[0047] Step 2, in response to the absence of other picture tags matching the above picture tags, the picture position is located according to the picture sequence number tag coordinates and the picture source tag coordinates included in the above picture tags to obtain a picture position coordinate group. The picture coordinates can be located on the area between the picture sequence number tag coordinates and the picture source tag coordinates by the above pdf parsing algorithm to obtain a picture position coordinate group.
[0048] Step 3, in response to the existence of other picture tags matching the above picture tag, determine that the pictures corresponding to the above picture tag and the other picture tags are in the same paragraph, and locate the picture position according to the picture sequence number tag coordinates and the picture source tag coordinates included in the above picture tag and the other picture tags, and obtain a picture position coordinate group. Among them, the minimum horizontal coordinate and the minimum vertical coordinate in the picture sequence number tag coordinates and the picture source tag coordinates can be used as the coordinates of the upper left corner of the picture, and the maximum horizontal coordinate and the maximum vertical coordinate can be used as the coordinates of the lower right corner of the picture. Secondly, the coordinates of the upper left corner of the picture and the coordinates of the lower right corner of the picture can be determined as the picture position coordinate group.
[0049] Step 4, according to the position of each image position coordinate in the above-mentioned image position coordinate group, the above-mentioned fixed document is subjected to image cropping processing to generate a document image. Among them, the image cropping function can be used to perform image cropping processing on the above-mentioned fixed document according to the position of each image position coordinate in the above-mentioned image position coordinate group to generate a document image. Here, the size of the PDF-to-image and the size of the PDF page itself can be proportionally processed. In addition, during the image cropping process, an offset can be added to the image position coordinates to avoid cropping to the document boundary.
[0050] In the third step, the table tags lacking the table number tag coordinates or the table source coordinates in the table tag sequence are determined as cross-page table tags. Here, since the document is identified by page, the lack of tag coordinates can indicate the situation that the table spans pages.
[0051] In the fourth step, for each cross-page table tag, according to the preset tag source order, a cross-page table tag corresponding to the cross-page table tag is determined as the table tag to be spliced. The cross-page table tag and the table tag to be spliced correspond to the same cross-page table.
[0052] The fifth step is to crop the fixed document table area according to the table number label coordinates and the table source coordinates included in each table label in the table label sequence to generate a cropped document table sequence. The fixed document table area can be cropped according to the upper left corner coordinates and the lower right corner coordinates of the table number label coordinates and the table source coordinates to generate a cropped document table sequence.
[0053] The sixth step is to splice the cropped document tables corresponding to the same cross-page table in the above cropped document table sequence based on each cross-page table tag and the corresponding table tag to be spliced to obtain a spliced document table. Among them, each cross-page table tag and the cropped document table matched with the corresponding splicing table tag can be spliced to obtain a spliced document table. Here, the splicing can be top-to-bottom splicing, that is, the cropped document table with the table sequence number tag is on the top, and the cropped document table with the table source coordinates is spliced at the bottom.
[0054] In the seventh step, each document image and the corresponding image label, as well as each spliced document table and the corresponding table label are respectively determined as chart feature information to obtain a chart feature information sequence. Among them, each document image and the corresponding image label can be determined as chart feature information. The spliced document table and the corresponding table label can also be determined as a chart feature label, and the document page numbers corresponding to the image document and the spliced document table are combined into a chart feature information sequence.
[0055] Step 104, performing chart interception on the fixed document according to the chart feature information sequence to obtain a document chart sequence.
[0056] In some embodiments, the execution entity may perform chart interception on the fixed document according to the chart feature information sequence to obtain a document chart sequence.
[0057] In some optional implementations of some embodiments, the execution subject performs chart interception on the fixed document according to the chart feature information sequence to obtain a document chart sequence, which may include the following steps: for each chart feature information in the chart feature information sequence, perform the following steps:
[0058] The first step is to determine the document page corresponding to the above-mentioned chart characteristic information in the above-mentioned fixed document as the target document page. The target document page can be identified by a page number.
[0059] The second step is to perform a chart interception on the target document page according to the chart coordinates of the chart coordinate group included in the chart feature information to obtain a document chart. Among them, the minimum horizontal coordinate and the minimum vertical coordinate in the chart coordinates can be used as the coordinates of the upper left corner of the chart, and the maximum horizontal coordinate and the maximum vertical coordinate can be used as the coordinates of the lower right corner of the chart. Secondly, the coordinates of the upper left corner of the chart and the coordinates of the lower right corner of the chart can be determined as a chart position coordinate group. Finally, the chart is intercepted according to the chart position coordinates to obtain a document chart.
[0060] Step 105: Perform a document cross-page check on each split document data sequence in the above-mentioned split document data sequence set to generate a document cross-page identification sequence, and for the split document data sequences corresponding to the document cross-page identification representing the document content cross-page in the above-mentioned document cross-page identification sequence, perform document content cross-page integration on the above-mentioned split document data sequences to obtain an integrated document data sequence.
[0061] In some embodiments, the above-mentioned execution entity can perform a document cross-page check on each split document data sequence in the above-mentioned split document data sequence set to generate a document cross-page identification sequence, and for the split document data sequences corresponding to the document cross-page identification representing the document content cross-page in the above-mentioned document cross-page identification sequence, perform document content cross-page integration on the above-mentioned split document data sequences to obtain an integrated document data sequence.
[0062] In some optional implementations of some embodiments, the execution subject performs a document cross-page check on each split document data sequence in the split document data sequence set to generate a document cross-page identification sequence, and for the split document data sequences corresponding to the document cross-page identification representing the document content cross-page in the document cross-page identification sequence, performs document content cross-page integration on the split document data sequences to obtain the integrated document data sequence, which may include the following steps:
[0063] For each split document data sequence in the above split document data sequence set, perform the following steps:
[0064] The first step is to perform text segment recognition on the last split document data in the split document data sequence to determine whether the last split document data is a cross-page text. The PDF parsing algorithm can be used to identify whether there is a period at the end of the last split document data.
[0065] The second step is to generate a document cross-page mark corresponding to the last split document data in response to the above. If there is no period, generate a document cross-page mark corresponding to the last split document data. Here, the page number of the split document data sequence can be used as the document cross-page mark.
[0066] For the split document data sequence corresponding to the document page span identifier representing the page span of the document content in the above document page span identifier sequence, the following steps are performed:
[0067] In the first step, the last split document data in the split document data sequence is merged with the first split document data in the next page split document data sequence adjacent to the split document data sequence to obtain merged document data. The data fusion can be to splice the texts in the last split document data and the first split document data in the next page split document data sequence adjacent to the split document data sequence in chronological order into new text data to obtain merged document data.
[0068] The second step is to perform title text deduplication processing on the above fused document data to obtain deduplication document data, wherein the split document data corresponding to the above fused document data is removed.
[0069] The third step is to determine the above deduplicated document data and other split document data in the above split document data sequence as an integrated document data sequence.
[0070] As an example, the document content extraction process for a fixed document can be as follows: Figure 2 As shown, the current PDF parsing method is mainly used to parse the research report file page by page, and each page of the research report file is divided into blocks. Each block contains many lines, and the lines are divided into different small fragments. Although this method obtains the details of every text in the research report, it is still too cumbersome to extract paragraphs. After the text extraction optimization operation specialized for the research report, the above details can be shielded, and when used, the paragraph corresponding to the block can be directly obtained. The rules include blank lines between paragraphs and paragraphs next to the title, and the positioning and trimming of charts.
[0071] Step 106: construct a document form corresponding to the fixed document based on the document chart sequence and each integrated document data sequence for storage.
[0072] In some embodiments, the execution entity may construct a document form corresponding to the fixed document based on the document chart sequence and each integrated document data sequence for storage.
[0073] In some optional implementations of some embodiments, the execution subject constructs a document form corresponding to the fixed document based on the document chart sequence and each integrated document data sequence, which may include the following steps:
[0074] According to a preset document layout model, each document chart in the above document chart sequence and each integrated document data in the integrated document data sequence are laid out to generate a document form. The document layout model may be a preset template. Here, document layout may be rearranged to generate a non-fixed document, i.e., a document form.
[0075] Optionally, in response to receiving a print instruction for the document form, the execution entity sends the document form to a printing device to print the form file. The print instruction is used to instruct the printing device to print the target content in the document form, and the target content may include but is not limited to at least one of the following: a document image, a document table, and a document structure diagram.
[0076] The above steps 103 to 106 and their related contents, as an inventive point of an embodiment of the present disclosure, solve the second technical problem mentioned in the background technology: "There are often many cross-page texts and cross-page charts in a document. If a formatted one-to-one extraction is performed, it is easy to divide the same paragraph of text or the same chart into multiple ones, which not only leads to document semantic errors and missing, but also document typesetting errors. Therefore, when printing a form, an incorrect form is printed, resulting in a waste of printing resources." The factors that lead to the waste of printing resources are often as follows: There are often many cross-page texts and cross-page charts in a document. If a formatted one-to-one extraction is performed, it is easy to divide the same paragraph of text or the same chart into multiple ones, which not only leads to document semantic errors and missing, but also document typesetting errors. Therefore, when printing a form, an incorrect form is printed. In order to achieve this effect, first, by comparing and identifying the paginated document before and after, it can be used to determine whether there is a text segmentation. Therefore, it can be used to splice and merge the segmented text to avoid the separation of the same paragraph of text. Next, consider the situation where image or chart recognition requires a lot of computing resources. Therefore, by identifying the coordinates of the serial number label in the upper left corner of the picture and table and the coordinates of the source label in the lower left corner, the position of the picture and table can be located. Here, the situation where multiple pictures or forms are in the same row is also taken into account. Therefore, by distinguishing the positions of different coordinates, it is judged whether the picture and the table are in the same section. In this way, the complete picture and table can be accurately captured. Avoid missing problems. Thus, the accuracy of the generated document form can be ensured. In addition, the waste of printing resources is avoided.
[0077] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: through the research report generation method of some embodiments of the present disclosure, the occupation of computing resources can be reduced. Specifically, the reason why a large amount of computing resources needs to be occupied is that the document content is large, and if all text parsing and image parsing are performed, a large amount of computing resources need to be occupied. Based on this, the research report generation method of some embodiments of the present disclosure, first of all, considering that splitting the document into lines for recognition or semantic recognition consumes more computing resources, the fixed document is split into small paragraphs by segment splitting processing. Thus, compared with splitting into finer lines, the occupation of computing resources can be reduced. Secondly, by extracting the coordinates of the chart, it can be used to locate the position of the chart in the document. Thus, it can be used to intercept the chart. In addition, considering that the text of the document spans pages, the text of the same paragraph may be divided into different paragraphs. Therefore, through the document cross-page check, it can be identified whether there is a situation where the text spans pages. Then, through the cross-page integration of the document content, the same paragraph of text across pages can be integrated into the same paragraph. Thus, not only can the loss of data be avoided, but also the occupation of computing resources can be reduced.
[0078] Further references Figure 3 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of devices of some embodiments of a research report form generation method, and these device embodiments are Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.
[0079] like Figure 3As shown, the apparatus 300 of some embodiments of the research report form generation method of some embodiments includes: an acquisition unit 301, a splitting processing unit 302, a feature extraction unit 303, a chart interception unit 304, a document cross-page verification and integration unit 305 and a construction unit 306. Among them, the acquisition unit 301 is configured to acquire a fixed document from a database; the splitting processing unit 302 is configured to split the fixed document segment by segment according to the page number sequence of the fixed document to generate a split document data sequence set, wherein each split document data in the split document data sequence set corresponds to a section of content in the fixed document; the feature extraction unit 303 is configured to extract chart features from the fixed document to obtain a chart feature information sequence, wherein each chart feature information in the chart feature information sequence includes a chart coordinate group; the chart interception unit 304 is configured to extract the chart feature information from the fixed document according to the chart feature information sequence. The fixed-type document is intercepted by a chart to obtain a document chart sequence; the document cross-page verification and integration unit 305 is configured to perform a document cross-page verification on each split document data sequence in the split document data sequence set to generate a document cross-page identification sequence, and for the split document data sequence corresponding to the document cross-page identification representing the document content cross-page in the document cross-page identification sequence, perform document content cross-page integration on the split document data sequence to obtain an integrated document data sequence; the construction unit 306 is configured to construct a document form corresponding to the fixed-type document based on the document chart sequence and each integrated document data sequence for storage.
[0080] It is understood that the units described in the device 300 are similar to those described in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 300 and the units contained therein, and will not be described in detail here.
[0081] Reference below Figure 4 , which shows a structural schematic diagram of an electronic device (eg, a computing device) 400 suitable for implementing some embodiments of the present disclosure. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0082] like Figure 4As shown, the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory 402 or a program loaded from a storage device 408 into a random access memory 403. Various programs and data required for the operation of the electronic device 400 are also stored in the random access memory 403. The processing device 401, the read-only memory 402, and the random access memory 403 are connected to each other via a bus 404. An input / output interface 405 is also connected to the bus 404.
[0083] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 The electronic device 400 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 4 Each block shown in the figure may represent one device, or may represent multiple devices as required.
[0084] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the read-only memory 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.
[0085] It should be noted that the computer-readable medium recorded in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0086] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0087] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains a fixed document from a database; splits the fixed document segment by segment according to the page number sequence of the fixed document to generate a split document data sequence set, wherein each split document data in the split document data sequence set corresponds to a section of content in the fixed document; extracts chart features from the fixed document to obtain a chart feature information sequence, wherein each chart feature information in the chart feature information sequence includes a chart coordinate group; performs chart interception on the fixed document according to the chart feature information sequence to obtain a document chart sequence; performs document cross-page verification on each split document data sequence in the split document data sequence set to generate a document cross-page identification sequence, and for the split document data sequence corresponding to the document cross-page identification representing the cross-page of document content in the document cross-page identification sequence, performs document content cross-page integration on the split document data sequence to obtain an integrated document data sequence; and constructs a document form corresponding to the fixed document based on the document chart sequence and each integrated document data sequence for storage.
[0088] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0089] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0090] The units described in some embodiments of the present disclosure may be implemented by software or hardware. The units described may also be set in a processor, for example, it may be described as: a processor includes an acquisition unit, a splitting processing unit, a feature extraction unit, a chart interception unit, a document cross-page verification and integration unit and a construction unit. The names of these units do not constitute a limitation on the units themselves in some cases. For example, the acquisition unit may also be described as a "unit for acquiring fixed documents from a database."
[0091] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0092] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.
Claims
1. A method for generating a research report form, comprising: Get fixed documents from the database; According to the page number sequence of the fixed document, the fixed document is segmented and processed to generate a segmented document data sequence set, wherein each segmented document data in the segmented document data sequence set corresponds to a segment of content in the fixed document; Extracting chart features from the fixed document to obtain a chart feature information sequence, wherein each chart feature information in the chart feature information sequence includes a chart coordinate group; According to the chart feature information sequence, the fixed document is subjected to chart interception to obtain a document chart sequence; Performing a document cross-page check on each split document data sequence in the split document data sequence set to generate a document cross-page identification sequence, and for the split document data sequences corresponding to the document cross-page identification representing the document content cross-page in the document cross-page identification sequence, performing document content cross-page integration on the split document data sequences to obtain an integrated document data sequence; Based on the document chart sequence and each integrated document data sequence, a document form corresponding to the fixed document is constructed for storage.
2. The method according to claim 1, wherein: The method further comprises: In response to receiving a print instruction for the document form, the document form is sent to a printing device to print the form file, wherein the print instruction is used to instruct the printing device to print the target content in the document form, and the target content includes at least one of the following: document image, document table, document structure diagram.
3. The method according to claim 1, wherein: The step of splitting the fixed document segment by segment according to the page number sequence of the fixed document to generate a split document data sequence set includes: Identifying page numbers of each page in the fixed document; According to the order of the page numbers, the fixed document is split page by page to obtain a document split page sequence; The paragraph information in each document split page in the document split page sequence is identified, and the document split page is segmented into paragraphs according to the paragraph information to generate a segmented document data sequence set.
4. The method according to claim 1, wherein: The step of performing chart interception on the fixed document according to the chart feature information sequence to obtain a document chart sequence includes: for each chart feature information in the chart feature information sequence, performing the following steps: Determine a document page corresponding to the chart feature information in the fixed document as a target document page; According to each chart coordinate of the chart coordinate group included in the chart feature information, the target document page is subjected to chart interception to obtain a document chart.
5. The method according to claim 1, wherein: The step of performing a document cross-page check on each split document data sequence in the split document data sequence set to generate a document cross-page identification sequence, and performing document content cross-page integration on the split document data sequences corresponding to the document cross-page identification representing the document content cross-page in the document cross-page identification sequence to obtain an integrated document data sequence, including: For each split document data sequence in the split document data sequence set, the following steps are performed: Performing text segment recognition on the last split document data in the split document data sequence to determine whether the last split document data is a cross-page text; In response, a document cross-page identifier corresponding to the last split document data is generated: For the split document data sequence corresponding to the document cross-page identifier representing the cross-page of the document content in the document cross-page identifier sequence, the following steps are performed: Merge the last split document data in the split document data sequence with the first split document data in the next page of the split document data sequence adjacent to the split document data sequence to obtain merged document data; Perform title text deduplication processing on the merged document data to obtain deduplicated document data; The deduplicated document data and other split document data in the split document data sequence are determined as an integrated document data sequence.
6. The method according to claim 5, wherein: The step of constructing a document form corresponding to the fixed document based on the document chart sequence and each integrated document data sequence includes: According to a preset document layout model, document layout is performed on each document chart in the document chart sequence and each integrated document data in the integrated document data sequence to generate a document form.
7. A research report form generating device, comprising: An acquisition unit configured to acquire a fixed document from a database; A splitting processing unit is configured to split the fixed document segment by segment according to the page number sequence of the fixed document to generate a split document data sequence set, wherein each split document data in the split document data sequence set corresponds to a segment of content in the fixed document; A feature extraction unit is configured to extract chart features from the fixed document to obtain a chart feature information sequence, wherein each chart feature information in the chart feature information sequence includes a chart coordinate group; A chart interception unit is configured to intercept the chart of the fixed document according to the chart feature information sequence to obtain a document chart sequence; a document cross-page verification and integration unit, configured to perform a document cross-page verification on each split document data sequence in the split document data sequence set to generate a document cross-page identification sequence, and for the split document data sequences corresponding to the document cross-page identification representing the document content cross-page in the document cross-page identification sequence, perform document content cross-page integration on the split document data sequences to obtain an integrated document data sequence; The construction unit is configured to construct a document form corresponding to the fixed document based on the document chart sequence and each integrated document data sequence for storage.
8. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Text processing method and device
CN113362026A
Document layout identification method and related device
CN115223182A
Digital delivery system based on digital twinborn body
CN115496452A
Document key element identification method and device, equipment and medium
CN117765544A
Relative fuzziness for fast reduction of false positives and false negatives in computational text searches
US20230401274A1