A method and apparatus for processing electronic document merging comments
Patent Information
- Application Number
- CN202310884084.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-18
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-07-18
AI Technical Summary
[0005]本发明要解决的技术问题在于,针对现有技术缺陷,本发明提供一种电子文档合并批注的处理方法及装置,以解决现有的自动化办公技术还无法自动实现多个文档的多个批注记录的合并的问题
[0049]本发明实现了基于python-docx合并docx文档,通过选择文件控件获取待合并文件,并在待合并文件中抽取得到对应的批注文档,可根据所有抽取得到的批注文档生成新批注文档,进而,通过对新批注文档中的文档进行重新编号及命名,可以将待合并文档中的批注编号更新为与注文档中的批注编号一致,从而得到合并内容及批注后的电子文档;本发明实现了对多个电子文档按照预定格式合并,同时在合并时会处理文档批注,保证合并后文档批注在正确的位置显示。
Smart Images

Figure CN116911263B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automated office technology, and in particular to a method and apparatus for merging annotations in electronic documents. Background Technology
[0002] With the development of the information society, people's demand for automated office work is becoming more and more widespread. At present, most work is completed by team collaboration. In order to ensure the quality of everyone's work, the whole process cannot be separated from the review or evaluation of electronic documents. Document review is done by annotating or revising the documents.
[0003] Because team members are divided into groups, each responsible for a specific chapter in the manual for their respective project, there is an urgent need to merge the chapters of the manual for the same project into a single published version. During this division of labor, project members only need to update their assigned chapters before publishing or updating the merged document. However, existing automated office technologies cannot automatically merge multiple annotation records from multiple documents.
[0004] Therefore, existing technologies still need improvement. Summary of the Invention
[0005] The technical problem to be solved by the present invention is that, in view of the defects of the prior art, the present invention provides a method and apparatus for merging annotations in electronic documents, so as to solve the problem that existing automated office technologies cannot automatically merge multiple annotation records of multiple documents.
[0006] The technical solution adopted by this invention to solve the technical problem is as follows:
[0007] In a first aspect, the present invention provides a method for merging annotations in electronic documents, including:
[0008] Obtain the path of the files to be merged by selecting the file control;
[0009] The files to be merged are obtained according to the path of the files to be merged, the files to be merged are merged, and the corresponding annotation documents are extracted from the files to be merged.
[0010] A new annotation document is generated based on all the extracted annotation documents. The annotation numbers in the new annotation document are renumbered, and the images and videos referenced in the content of the new annotation document are renumbered and renamed.
[0011] Update the comment numbers in the document to be merged to match the comment numbers in the new commented document;
[0012] Add the processed new annotated document to the merged file to obtain the merged content and the annotated electronic document.
[0013] In one implementation, the step of obtaining the files to be merged based on the path of the files to be merged, and merging the files to be merged, includes:
[0014] Based on the path of the files to be merged, retrieve individual files to be merged in a loop, and determine whether each retrieved file is a docx document;
[0015] If the extracted single file to be merged is not a docx document, then the corresponding document will be converted into the docx document;
[0016] Merge all docx documents.
[0017] In one implementation, extracting the corresponding annotation document from the files to be merged includes:
[0018] Based on the methods for manipulating docx files in python-docx, determine in turn whether each docx document contains comments;
[0019] If the current docx document contains comments, then the current docx document will be decompressed to obtain multiple files with comments and multiple folders.
[0020] In one implementation, after unzipping the annotated docx document, multiple annotated files and multiple folders are obtained, including: the word / comments.xml document, the word / _rels / comments.xml.rels folder, the word / media folder, the word / document.xml document, and the word / _rels / document.xml.rels file;
[0021] The word / _rels / comments.xml.rels folder contains references to images and videos in the comments.xml annotations;
[0022] The word / media folder contains all the images for the word document.
[0023] The word / document.xml document contains Word content, and the Word content includes annotation URL references;
[0024] The word / _rels / document.xml.rels folder contains reference information for files in document.xml.
[0025] In one implementation, the step of generating a new annotation document based on all extracted annotation documents and renumbering the annotation numbers in the new annotation document includes:
[0026] Copy the content of the word / comments.xml document to the main comments.xml document, and set the tag values in the current comments.xml document to increment sequentially;
[0027] Modify the historical comment reference values in the word / document.xml document after unzipping a single docx document to the comment reference values in the current comments.xml document, and modify the start tag value and standard tag value in the current document.xml document to be consistent with the tag values in the total comments.xml document.
[0028] In one implementation, the step of renumbering the annotations in the new annotation document and renumbering and renaming the images and videos referenced in the content of the new annotation document includes:
[0029] Determine if a specified file exists in the comments;
[0030] If the specified file exists, copy the contents of the word / _rels / comments.xml.rels file to the main comments.xml.rels file, and set the relationship tag values in the current comments.xml.rels file to increment sequentially;
[0031] Create a new file named comments.xml.rels in the current project directory. Extract the content of word / _rels / comments.xml.rels from the document to be merged and decompressed into the newly created comments.xml.rels file. Then, renumber the attribute Id values of the relationship tags in the newly created comments.xml.rels file according to the order.
[0032] Using the target relationship tag value in the newly created comments.xml.rels file as the filename, search for the corresponding file in the word / media folder. Once the corresponding file is found, copy it to the current system file directory, rename the corresponding file according to the current filename and order, and write the renamed filename back to the target relationship tag in the newly created comments.xml.rels file.
[0033] In one implementation, adding the processed new annotated document to the merged file to obtain the merged content and the annotated electronic document includes:
[0034] Restore the folder and files extracted from a single docx document to the original single docx document;
[0035] The docx file template is parsed to obtain the format of the docx document, ensuring that the files are merged according to the template format and guaranteeing the consistency of the format after merging;
[0036] According to the document format, merge individual docx documents one by one. After merging all the documents to be merged, the merged docx document is generated.
[0037] Unzip the merged docx document to obtain the merged unzipped files and folders;
[0038] Move the main comments.xml file to the word / directory of the merged unzipped folder; move the main comments.xml.rels file to the word / _rels / directory of the merged unzipped folder; move the files in the file folder to the word / media / directory of the merged unzipped folder.
[0039] After merging, the extracted files and folders are converted into a merged docx document.
[0040] Secondly, the present invention provides a processing apparatus for merging annotations in electronic documents, comprising:
[0041] The path acquisition module is used to obtain the path of the files to be merged by selecting a file control.
[0042] The annotation document extraction module is used to obtain the file to be merged according to the path of the file to be merged, merge the file to be merged, and extract the corresponding annotation document from the file to be merged.
[0043] The annotation document processing module is used to generate a new annotation document based on all the extracted annotation documents, renumber the annotation numbers in the new annotation document, and renumber and rename the images and videos referenced in the content of the new annotation document.
[0044] The numbering update module is used to update the annotation numbers in the document to be merged to match the annotation numbers in the new annotation document.
[0045] The file merging module is used to add the processed new annotated document to the merged file, resulting in a merged electronic document with annotations.
[0046] Thirdly, the present invention provides a terminal, comprising: a processor and a memory, wherein the memory stores a processing program for merging annotations in electronic documents, and the processing program for merging annotations in electronic documents, when executed by the processor, is used to implement the processing method for merging annotations in electronic documents as described in the first aspect.
[0047] Fourthly, the present invention also provides a medium, which is a computer-readable storage medium, storing a processing program for merging annotations in electronic documents, wherein the processing program for merging annotations in electronic documents, when executed by a processor, is used to implement the operation of the processing method for merging annotations in electronic documents as described in the first aspect.
[0048] The present invention, by employing the above technical solution, has the following effects:
[0049] This invention implements the merging of docx documents based on python-docx. It obtains the files to be merged by selecting a file control, extracts the corresponding annotation documents from these files, and generates a new annotation document based on all extracted annotation documents. Furthermore, by renumbering and renaming the documents in the new annotation document, the annotation numbers in the files to be merged are updated to match the annotation numbers in the new annotation document, thus obtaining the merged content and the annotated electronic document. This invention enables the merging of multiple electronic documents according to a predetermined format, and simultaneously processes document annotations during the merging process to ensure that the annotations are displayed in the correct positions after merging. Attached Figure Description
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0051] Figure 1 This is a flowchart of a method for merging annotations in electronic documents in one implementation of the present invention.
[0052] Figure 2 This is a schematic diagram of the files and folders obtained by decompressing a docx file in one implementation of the present invention.
[0053] Figure 3 This is a flowchart illustrating the overall scheme of one implementation of the present invention.
[0054] Figure 4 This is a functional schematic diagram of the terminal in one implementation of the present invention.
[0055] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0057] Exemplary methods
[0058] The demand for automated office work is expanding, and most tasks are now completed collaboratively in teams. After a project is finished, the various chapters and manuals distributed throughout the project need to be merged into a single project manual according to a specific order and format. To ensure the quality of everyone's work, the entire process relies heavily on electronic document review. Furthermore, to avoid affecting the document content, reviewers typically don't directly modify the document during review; instead, they insert comments. Word automatically assigns unique numbers and names to each comment. While many document merging software programs are available, most ignore comments or the files contained within comments (e.g., images, audio, video files). Therefore, these programs cannot automatically merge multiple comment records from multiple documents.
[0059] To address the aforementioned issues, this invention provides a method for merging annotations in electronic documents. This method, based on python-docx, merges docx documents. It obtains the files to be merged by selecting a file control and extracts the corresponding annotation documents from these files. A new annotation document is generated based on all extracted annotation documents. Furthermore, by renumbering and renaming the documents in the new annotation document, the annotation numbers in the files to be merged are updated to match those in the original annotation document, resulting in a merged electronic document with annotations. Therefore, this invention enables the merging of multiple electronic documents according to a predetermined format, while simultaneously processing document annotations during the merging process to ensure that the annotations are displayed in the correct positions after merging.
[0060] like Figure 1 As shown, this embodiment of the invention provides a method for merging annotations in electronic documents, including the following steps:
[0061] Step S100: Obtain the path of the files to be merged by selecting the file control.
[0062] In this embodiment, the method for merging annotations in electronic documents is applied to a terminal, which includes, but is not limited to, devices such as computers.
[0063] In this embodiment, various types of documents are first converted into docx documents, and then docx documents are merged based on python-docx. python-docx is a Python library for creating and modifying Microsoft Word documents, which provides a full set of Word operations and is the most commonly used Word tool. docx is the file extension of Microsoft Word. The main content of a docx file is saved in XML format, and a docx file is essentially a ZIP file.
[0064] In this embodiment, the python-docx Python library is used. This python-docx library does not support comment merging. Comments in docx format files are document notes or annotations added to a separate comment window for review. Comments are inserted when the reviewer only comments on the document and does not directly modify it, as comments do not affect the document's content. Comments are hidden text, and Word automatically assigns each comment a unique number and name.
[0065] To address the issue that python-docx does not support comment merging, this invention provides a method for merging electronic documents and processing comments during the merging process. This method can merge electronic documents containing various comments, while ensuring that the comments are displayed in the same position in the document after merging as before, thus maintaining consistency between the document and the comment content before and after the merge.
[0066] In this embodiment, for electronic documents that need to be merged, the path of the file to be merged can be obtained by selecting a file control (e.g., FineReport), so that the electronic documents to be merged can be retrieved from the database according to the obtained file path.
[0067] like Figure 1 As shown, in one implementation of this invention, the method for merging annotations in electronic documents further includes the following steps:
[0068] Step S200: Obtain the files to be merged according to the path of the files to be merged, merge the files to be merged, and extract the corresponding annotation documents from the files to be merged.
[0069] In this embodiment, the files to be merged can be obtained based on their paths. For multiple documents to be merged, each document is first converted into a docx document, and then the python-docx library is used to merge the docx documents. However, python-docx currently does not support annotation features. Therefore, to ensure that the annotations in each docx document are combined and displayed correctly in the merged document, this embodiment decompresses the docx documents, resulting in many files and folders, from which the corresponding annotation documents can be extracted.
[0070] Specifically, in one implementation of this embodiment, step S200 includes the following steps:
[0071] Step S201: Extract individual files to be merged in a loop according to the file path to be merged, and determine whether the extracted individual file to be merged is a docx document;
[0072] Step S202: If the extracted single file to be merged is not a docx document, then convert the corresponding document into the docx document;
[0073] Step S203: Merge all docx documents.
[0074] In this embodiment, multiple document addresses to be merged are obtained, and a single document is retrieved in a loop. If the document is not a docx document, it is converted into a docx document so that it can be decompressed according to the python-docx method for operating the docx file library.
[0075] Specifically, in one implementation of this embodiment, step S200 further includes the following steps:
[0076] Step S204: Based on the method of python-docx to operate the docx file library, determine in turn whether each docx document contains comments;
[0077] Step S205: If the current docx document contains comments, then decompress the current docx document to obtain multiple files with comments and multiple folders.
[0078] In this embodiment, based on the method of python-docx operating the docx file library, it is determined whether the document contains comments; if it contains comments, the docx document is first decompressed to obtain many files and folders; after decompressing the docx file, many files and folders will be obtained (e.g., Figure 2 As shown in the figure, by analyzing and comparing the differences between commented and uncommented files and their correlations, the comments and documents are renumbered and renamed.
[0079] Furthermore, after unzipping the annotated docx document, you will get multiple annotated files and multiple folders, including: the word / comments.xml document, the word / _rels / comments.xml.rels folder, the word / media folder, the word / document.xml document, and the word / _rels / document.xml.rels file; among them, the word / _rels / comments.xml.rels folder contains the reference descriptions of images and videos in the comments.xml annotations; the word / media folder contains all the images in the word document; the word / document.xml document contains the word content, and the word content contains the annotation URL references; the word / _rels / document.xml.rels folder contains the file reference information in document.xml.
[0080] like Figure 1 As shown, in one implementation of this invention, the method for merging annotations in electronic documents further includes the following steps:
[0081] Step S300: Generate a new annotation document based on all the extracted annotation documents, renumber the annotations in the new annotation document, and renumber and rename the images and videos referenced in the content of the new annotation document.
[0082] In this embodiment, for the newly generated annotation document, it is necessary to renumber the annotations in the new annotation document and renumber and rename the documents in the new annotation document to ensure that the annotations in the merged document are displayed in the correct position.
[0083] The following rules apply to the numbering and naming of new commented documents:
[0084] In general, comments.xml.rels references which files, and document.xml calls the files referenced in document.xml.rels. For example, comments.xml.rels includes a reference to the comments.xml file. document.xml displays the annotation reference address. comments.xml.rels also references which files, and comments.xml calls the files. For example, if comments.xml.rels references an image and assigns it a number, the comments.xml annotation content contains an image and references that image number.
[0085] Specific approach:
[0086] 1. Copy word / comments.xml to the main / comments.xml file and renumber the text.
[0087] 2. Modify the annotation reference address in word / document.xml.
[0088] 3. Similarly, copy word / _rels / comments.xml.rels to the main comments.xml.rels file and renumber the references in order.
[0089] 4. Modify the previous reference number in word / comments.xml to the new number in comments.xml.rels.
[0090] 5. In comments.xml.rels, reference the corresponding files (e.g., images), then copy the files to the main file directory, rename the files in the new order, and then modify the new file names back into comments.xml.rels.
[0091] 6. Unzip the merged docx document and move all files from the main comments.xml, main comments.xml.rels, and the file directory to the corresponding location in the unzipped folder of the merged docx document.
[0092] 7. Restore the extracted folder to a docx document.
[0093] Specifically, in one implementation of this embodiment, step S300 includes the following steps:
[0094] Step S301: Copy the content of the word / comments.xml document to the main comments.xml document, and set the tag values in the current comments.xml document to increment sequentially;
[0095] Step S302: Modify the historical comment reference values in the word / document.xml document after decompressing a single docx document to the comment reference values in the current comments.xml document, and modify the start tag value and standard tag value in the current document.xml document to be consistent with the tag values in the total comments.xml document.
[0096] In this embodiment, the content of the word / comments.xml file is first copied to the main comments.xml file, and the tags in comments.xml are set.<comment w:id> The values increase sequentially.
[0097] Then, modify the previous comment references in the word / document.xml file after unzipping a single docx file.<comment w:id> The value is currently in the total comments.xml.<comment w:id> Values, and settings for the tags in document.xml.<w:commentRangeStart w:id> Value (i.e., the starting tag value) and tag<w:CommentReference w:id> The values (i.e., standard tag values) are used to make these two tag values consistent with the total values in comments.xml.<comment w:id> The values are consistent.
[0098] Specifically, in one implementation of this embodiment, step S300 further includes the following steps:
[0099] Step S303: Determine whether the specified file exists in the comment;
[0100] Step S304: If the specified file exists, copy the contents of the word / _rels / comments.xml.rels file to the main comments.xml.rels file, and set the relationship tag values in the current comments.xml.rels file to increment sequentially;
[0101] Step S305: Create a new file named comments.xml.rels in the current project directory, extract the content of word / _rels / comments.xml.rels from the document to be merged and decompressed into the newly created comments.xml.rels file, and renumber the attribute Id values of the relationship tags in the newly created comments.xml.rels file according to the order.
[0102] Step S306: Using the target relationship tag value in the newly created comments.xml.rels file as the filename, search for the corresponding file in the word / media folder. After finding the corresponding file, copy the corresponding file to the current system file directory, rename the corresponding file according to the current filename and order, and write the renamed filename back to the target relationship tag in the newly created comments.xml.rels file.
[0103] In this embodiment, the tag is set in document.xml.<w:commentRangeStart w:id> Values and Labels<w:CommentReference w:id> After setting the value, it's necessary to check if the comment contains any files (such as images, video files, audio files, etc.). If it does, copy the contents of the word / _rels / comments.xml.rels file to the main comments.xml.rels file, and set the tag in comments.xml.rels.<Relationship Id> The values increase sequentially.
[0104] Furthermore, modify the tags in the main comments.xml file.<a:blip r:embed> The value (i.e., the embedded tag value) is the total value of the tag in comments.xml.rels.<Relationship Id> Values (i.e., relationship label values) are used to ensure a one-to-one correspondence.
[0105] Furthermore, the tags in the main comments.xml.rels file will be...<Relationship Target> The value (i.e., the target relation tag value) is set to the filename. The program searches for the file in the word / media folder, copies it to the current system's file directory, and renames the file according to the current filename and its order. The renamed filename is then written back to the comments.xml.rels file under the `<relation>` tag.<Relationship Target> (i.e., the target relationship label).
[0106] like Figure 1 As shown, in one implementation of this invention, the method for merging annotations in electronic documents further includes the following steps:
[0107] Step S400: Update the annotation numbers in the document to be merged to match the annotation numbers in the new annotated document.
[0108] In this embodiment, for the annotation numbers in the documents to be merged, based on the annotation numbers in the new annotation document of step S300 above, as well as the document number and name, the annotation numbers in the documents to be merged are updated to be consistent with the annotation numbers in the new annotation document, so as to ensure that the annotations in the merged documents are displayed in the correct positions.
[0109] like Figure 1 As shown, in one implementation of this invention, the method for merging annotations in electronic documents further includes the following steps:
[0110] Step S500: Add the processed new annotated document to the merged file to obtain the merged content and the annotated electronic document.
[0111] In this embodiment, after renumbering and renaming the new annotated documents, it is necessary to restore the single docx decompressed folder and file to the original single docx document. Then, according to the docx file template format, the document and annotation content are merged to obtain the merged annotated electronic document.
[0112] Specifically, in one implementation of this embodiment, step S500 includes the following steps:
[0113] Step S501: Restore the folder and files extracted from a single docx document to the original single docx document;
[0114] Step S502: Parse the docx file template to obtain the format of the docx document, and ensure that the files are merged according to the template format to ensure the consistency of the format after merging;
[0115] Step S503: According to the document format, merge individual docx documents one by one. After merging all the documents to be merged, a merged docx document is generated.
[0116] Step S504: Decompress the merged docx document to obtain the merged decompressed files and folders;
[0117] Step S505: Move the main comments.xml to the word / directory of the merged and unzipped folder; move the main comments.xml.rels to the word / _rels / directory of the merged and unzipped folder; move the files in the file folder to the word / media / directory of the merged and unzipped folder.
[0118] Step S506: After merging, the unzipped files and folders are converted into a merged docx document.
[0119] In this embodiment, the single docx decompressed folder and files are first restored to the original single docx document, and the docx file template is parsed to obtain the format of the docx document, so as to ensure the consistency of the document format after merging the files according to the template format.
[0120] Then, the process begins by merging individual documents according to their format. After repeating this process, all documents to be merged are merged to generate a merged docx document.
[0121] Finally, the merged docx document is decompressed to obtain the merged decompressed files and folders; and by moving the documents, the merged decompressed files and folders are converted back into the merged docx document.
[0122] Specifically: Move the main comments.xml file to the word / directory of the merged unzipped folder (if comments.xml already exists in the word / directory, replace it with the main comments.xml file); move the main comments.xml.rels file to the word / _rels / directory of the merged unzipped folder (if comments.xml.rels already exists in the word / _rels / directory, replace it with the main comments.xml.rels file); move the files in the file folder to the word / media / directory of the merged unzipped folder; after moving the documents, the merged unzipped files and folders are converted into a merged docx document.
[0123] The technical solution in this embodiment is illustrated below through a practical application scenario:
[0124] like Figure 3 As shown, in one implementation of this embodiment, the method for merging annotations in electronic documents includes:
[0125] 1. Obtain the path of the files to be merged based on the file selection control.
[0126] 2. Based on the path of the files to be merged, determine the document type. If it is not a docx document, convert it to a docx document to be merged.
[0127] 3. Determine whether the documents to be merged contain comments.
[0128] 4. If there are comments, first call ZipFile() to convert the documents to be merged into zip files, and then call extractall() to decompress the zip files.
[0129] 5. Create a new file named comments.xml in the current project directory. Extract the content of word / comments.xml from the document to be merged and unzipped into the new comments.xml file, and add the tags within it. <w:comment>Copy the values of the w:id attribute in the middle to the newly created comments.xml file, renumbering them in order.
[0130] 6. The `<script>` tag in the word / document.xml file of the document to be merged and unzipped. <w:commentrangestart> 、 <w:commentrangeend> 、 <w:commentreference>The value of the attribute w:id in the middle is updated to match the value of the tag in the newly created comments.xml file. <w:comment>The value of the attribute w:id is consistent.
[0131] 7. Create a new file named comments.xml.rels in the current project directory. Extract the content of word / _rels / comments.xml.rels from the document to be merged and unzipped into the newly created comments.xml.rels file, and add the tags to it. <relationship>The values of the attribute Id are renumbered sequentially in the newly created comments.xml.rels file.
[0132] 8. Update the tags in the newly created comments.xml file.<a:blip r:embed> The value of the tag in the newly created comments.xml.rels file.<Relationship Id> Values that ensure a one-to-one correspondence.
[0133] 9. Based on the tags in the newly created comments.xml.rels file<Relationship Target> The value is the filename. Locate the file in the word / media folder, copy it to the current system's file directory, and rename the file according to the current filename and its order. Then, write the renamed filename back to the newly created comments.xml.rels file using the `<comments.xml.rels>` tag.<Relationship Target> middle.
[0134] 10. Unzip the docx folder and files to restore the original docx document to be merged.
[0135] 11. Based on the file paths to be merged, merge all documents to obtain the merged document.
[0136] 12. For the merged docx document, first call ZipFile() to convert the merged file into a zip file, and then call extractall() to decompress the zip file.
[0137] 13. Create a new comments.xml file in the merged and unzipped folder word / directory in the mobile project (if a comments.xml file already exists in the word / directory, simply replace it with the main comments.xml file).
[0138] 14. Create a new comments.xml.rels file in the merged and unzipped folder word / _rels / in the mobile project (if comments.xml.rels already exists in word / _rels / , simply replace it).
[0139] 15. Move the files in the "file" folder to the merged, unzipped folder "word / media / ".
[0140] 16. After merging and decompressing, the files and folders are converted into docx documents and the target document is obtained.
[0141] This embodiment achieves the following technical effects through the above technical solution:
[0142] This embodiment implements merging docx documents based on python-docx. It obtains the files to be merged by selecting a file control and extracts the corresponding annotation documents from them. A new annotation document is generated based on all extracted annotation documents. Then, by renumbering and renaming the documents in the new annotation document, the annotation numbers in the files to be merged are updated to match the annotation numbers in the new annotation document, thus obtaining the merged content and the annotated electronic document. This embodiment achieves the merging of multiple electronic documents according to a predetermined format, and simultaneously processes document annotations during the merging process to ensure that the annotations are displayed in the correct positions after merging.
[0143] Exemplary device
[0144] Based on the above embodiments, the present invention also provides a processing apparatus for merging annotations in electronic documents, comprising:
[0145] The path acquisition module is used to obtain the path of the files to be merged by selecting a file control.
[0146] The annotation document extraction module is used to obtain the file to be merged according to the path of the file to be merged, merge the file to be merged, and extract the corresponding annotation document from the file to be merged.
[0147] The annotation document processing module is used to generate a new annotation document based on all the extracted annotation documents, renumber the annotation numbers in the new annotation document, and renumber and rename the images and videos referenced in the content of the new annotation document.
[0148] The numbering update module is used to update the annotation numbers in the document to be merged to match the annotation numbers in the new annotation document.
[0149] The file merging module is used to add the processed new annotated document to the merged file, resulting in a merged electronic document with annotations.
[0150] Based on the above embodiments, the present invention also provides a terminal, the principle block diagram of which can be as follows: Figure 4 As shown.
[0151] The terminal includes: a processor, a memory, an interface, a display screen, and a communication module connected via a system bus; wherein, the processor of the terminal provides computing and control capabilities; the memory of the terminal includes a storage medium and internal memory; the storage medium stores the operating system and computer programs; the internal memory provides an environment for the operation of the operating system and computer programs in the storage medium; the interface is used to connect to external devices, such as mobile terminals and computers; the display screen is used to display relevant information; and the communication module is used to communicate with a cloud server or mobile terminal.
[0152] This computer program is executed by the processor to implement the processing method of merging annotations in electronic documents.
[0153] It will be understood by those skilled in the art that Figure 4 The block diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal to which the present invention is applied. A specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0154] In one embodiment, a terminal is provided, comprising: a processor and a memory, the memory storing a processing program for merging annotations in electronic documents, the processing program for merging annotations in electronic documents being executed by the processor to implement the above-described processing method for merging annotations in electronic documents.
[0155] In one embodiment, a storage medium is provided, wherein the storage medium stores a processing program for merging annotations in electronic documents, which, when executed by a processor, is used to implement the above-described processing method for merging annotations in electronic documents.
[0156] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory.
[0157] In summary, this invention provides a method and apparatus for merging annotations in electronic documents. The method includes: obtaining the path of files to be merged by selecting a file control; obtaining the files to be merged according to the file path, merging the files to be merged, and extracting corresponding annotation documents from the files to be merged; generating a new annotation document based on all extracted annotation documents, renumbering the annotations in the new annotation document, and renumbering and renaming the images and videos referenced in the content of the new annotation document; updating the annotation numbers in the files to be merged to match the annotation numbers in the new annotation document; and adding the processed new annotation document to the merged file to obtain the merged content and annotated electronic document. This invention enables the merging of multiple electronic documents according to a predetermined format, and simultaneously processes document annotations during the merging process to ensure that the annotations are displayed in the correct positions after merging.
[0158] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.< / relationship> < / w:comment> < / w:commentreference> < / w:commentrangeend> < / w:commentrangestart> < / w:comment>
Claims
1. A method for merging annotations in electronic documents, characterized in that, include: Obtain the path of the files to be merged by selecting the file control; The files to be merged are obtained according to the path of the files to be merged, the files to be merged are merged, and the corresponding annotation documents are extracted from the files to be merged. A new annotation document is generated based on all the extracted annotation documents. The annotation numbers in the new annotation document are renumbered, and the images and videos referenced in the content of the new annotation document are renumbered and renamed. Update the comment numbers in the files to be merged to match the comment numbers in the new comment document; Add the processed new annotated document to the merged file to obtain the merged content and the annotated electronic document. The step of generating a new annotation document based on all extracted annotation documents, and renumbering the annotations in the new annotation document, includes: Copy the content of the word / comments.xml document to the main comments.xml document, and set the tag values in the current comments.xml document to increment sequentially; Modify the historical comment reference values in the word / document.xml document after unzipping a single docx document to the comment reference values in the current comments.xml document, and modify the start tag value and standard tag value in the current document.xml document to be consistent with the tag values in the total comments.xml document; The renumbering and renaming of images and videos referenced in the content of the new annotated document includes: Determine if a specified file exists in the comments; If the specified file exists, copy the contents of the word / _rels / comments.xml.rels file to the main comments.xml.rels file, and set the relationship tag values in the current comments.xml.rels file to increment sequentially; Create a new file named comments.xml.rels in the current project directory. Extract the content of word / _rels / comments.xml.rels from the document to be merged and decompressed into the newly created comments.xml.rels file. Then, renumber the attribute Id values of the relationship tags in the newly created comments.xml.rels file according to the order. Using the target relationship tag value in the newly created comments.xml.rels file as the filename, search for the corresponding file in the word / media folder. Once the corresponding file is found, copy it to the current system file directory, rename the corresponding file according to the current filename and order, and write the renamed filename back to the target relationship tag in the newly created comments.xml.rels file.
2. The method for merging annotations in electronic documents according to claim 1, characterized in that, The step of obtaining the files to be merged based on the path of the files to be merged, and merging the files to be merged, includes: Based on the path of the files to be merged, retrieve individual files to be merged in a loop, and determine whether each retrieved file is a docx document; If the extracted single file to be merged is not a docx document, then the corresponding document will be converted into the docx document; Merge all docx documents.
3. The method for merging annotations in electronic documents according to claim 1, characterized in that, The corresponding annotation documents extracted from the files to be merged include: Based on the methods for manipulating docx files in python-docx, determine in turn whether each docx document contains comments; If the current docx document contains comments, then the current docx document will be decompressed to obtain multiple files with comments and multiple folders.
4. The method for merging annotations in electronic documents according to claim 3, characterized in that, After unzipping the annotated docx document, you will get multiple annotated files and multiple folders, including: word / comments.xml document, word / _rels / comments.xml.rels folder, word / media folder, word / document.xml document, and word / _rels / document.xml.rels file; The word / _rels / comments.xml.rels folder contains references to images and videos in the comments.xml annotations; The word / media folder contains all the images for the word document. The word / document.xml document contains Word content, and the Word content includes annotation URL references; The word / _rels / document.xml.rels folder contains reference information for files in document.xml.
5. The method for merging annotations in electronic documents according to claim 1, characterized in that, The process involves adding the processed new annotated document to the merged file, resulting in a merged electronic document with annotated content, including: Restore the folder and files extracted from a single docx document to the original single docx document; The docx file template is parsed to obtain the format of the docx document, ensuring that the files are merged according to the template format and guaranteeing the consistency of the format after merging; According to the document format, merge individual docx documents one by one. After merging all the documents to be merged, the merged docx document is generated. Unzip the merged docx document to obtain the merged unzipped files and folders; Move the original comments.xml file to the word / directory of the merged unzipped folder; move the original comments.xml.rels file to the word / _rels / directory of the merged unzipped folder; move the files in the file folder to the word / media / directory of the merged unzipped folder. After merging, the extracted files and folders are converted into a merged docx document.
6. A processing apparatus for merging annotations in electronic documents, used to implement the processing method for merging annotations in electronic documents as described in any one of claims 1-5, characterized in that, include: The path acquisition module is used to obtain the path of the files to be merged by selecting a file control. The annotation document extraction module is used to obtain the file to be merged according to the path of the file to be merged, merge the file to be merged, and extract the corresponding annotation document from the file to be merged. The annotation document processing module is used to generate a new annotation document based on all the extracted annotation documents, renumber the annotation numbers in the new annotation document, and renumber and rename the images and videos referenced in the content of the new annotation document. The numbering update module is used to update the annotation numbers in the document to be merged to match the annotation numbers in the new annotation document. The file merging module is used to add the processed new annotated document to the merged file, resulting in a merged electronic document with annotations.
7. A terminal, characterized in that, include: The processor and memory, wherein the memory stores a processing program for merging annotations in electronic documents, and the processing program for merging annotations in electronic documents, when executed by the processor, is used to implement the operation of the processing method for merging annotations in electronic documents as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a processing program for merging annotations in electronic documents. When executed by a processor, the processing program for merging annotations in electronic documents is used to implement the operation of the processing method for merging annotations in electronic documents as described in any one of claims 1-5.
Citation Information
Patent Citations
Method and device for exporting multi-document annotations
CN102819543A
Annotation method and system for automatically realizing fine granularity and diversification of docx files
CN110968999A