Document processing, optimizing and polishing method and device based on large model
By parsing and processing the XML structure of the .docx document, extracting and splicing text content for the big model to polish, and replacing the original text content while keeping the document structure and style unchanged, the problem of poor polishing effects in Chinese and English in the prior art and the inability to handle complex formats is solved, and a more efficient document polishing effect is achieved.
Patent Information
- Application Number
- CN202510214791.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art cannot effectively optimize English expression and maintain the consistency of professional terms in English polishing, and the large model cannot receive and automatically identify formulas, pictures, and hyperlink formats in academic writing of thesis, and cannot directly process, optimize, and polish English thesis documents.
By analyzing the underlying XML structure of the .docx document, extracting the text content and judging the integrity, splicing the continuous text and submitting it to the big model to polish it, generating new text content, and replacing it back to the original text content while keeping the document structure and style unchanged.
It improves context understanding and consistency of professional terms, greatly improves the polishing effect, and supports direct upload of paper documents for polishing, solving the problem that large models cannot recognize and retain complex format content.
Smart Images

Figure CN120146028A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of document processing, and specifically relates to a method and device for document processing, optimization, and polishing based on a large model. Background Art
[0002] Currently, document processing generally relies on deep learning technology. Through a trained deep learning model, English text in a document is converted into word vectors, and then by calculating the distance between word vectors, the best polished English text is predicted.
[0003] Problems and Disadvantages of the Existing Technology: The deep learning model's ability to understand the context of text is far inferior to that of a large model. Therefore, in the field of English polishing, it cannot optimize English expressions well and maintain the consistency of professional terms; However, existing large text models cannot receive and automatically recognize common formula, picture, and hyperlink formats in academic paper writing, nor can they generate word documents. Therefore, it is difficult to directly process, optimize, and polish English paper documents using large model technology.
[0004] Content of the Application The purpose of this application is to solve the problems existing in the prior art, and provide a method and device for document processing, optimization, and polishing based on a large model.
[0005] To solve the technical problems, the technical solution of this application is: A method for document processing, optimization, and polishing based on a large model, including the following steps: Step 1: Document Upload: The user uploads the.docx document to be polished through a Web page; Step 2: Save the Document: Save the uploaded.docx document to be polished; Step 3: Parse the.docx Document Structure: Use the lxml library to parse the underlying XML structure of the.docx document to be polished to obtain the document.xml document; Step 4: Text Content Extraction and Polishing; Traverse the XML structure tree in the document.xml document, parse the w:r tags in the paragraphs of the XML structure tree in order, extract the text content and judge its integrity, splice the continuous text and submit it to the large model for polishing to generate new text content; Step 5: Restore the Document and Replace the Text: Replace the polished new text content back to the corresponding positions in the original text content while keeping the document structure and style unchanged; Step 6: Document Difference Comparison: Compare the polished new text content with the original text content, mark the text differences as revision items, and generate and save a comparison document; Step 7: The user downloads the comparison document to complete the process of document processing, optimization, and polishing.
[0006] Preferably, step 1 is specifically as follows: Through the Flask backend framework, implement a document upload interface. Users call the document upload interface through the Web page to upload the.docx document to be polished; Step 2 is specifically as follows: In the document upload interface, the received.docx document to be polished will be saved to both the local document server and the cloud object storage system at the same time.
[0007] Preferably, step 3 is specifically as follows: Unzip the.docx document to be polished through zipfile and obtain the document.xml document, which contains all the content and format information in the.docx document to be polished.
[0008] Preferably, step 4 is specifically as follows: Step 4-1: Construct text dictionary d1: Traverse the XML structure tree in the document.xml document. The text in the XML structure tree consists of paragraphs, and paragraphs consist of multiple runs. Retrieve according to the w:t tag, store all the XML tags of the paragraphs in an array, extract the text content of the paragraphs, and store it in the text dictionary d1. There are multiple text elements in the text dictionary d1, and after splicing multiple text elements, it is a complete sentence; at the same time, extract the non-text part, add the corresponding index, and store it in the index dictionary g1. Integrity judgment rule: Judge whether each spliced text element constitutes a complete English sentence, and determine whether the text is a complete sentence by checking the end punctuation. If it is a complete sentence after splicing, splice the relevant multiple text elements together to form a complete sentence. Process the runs in the paragraph: Parse and process the w:r tags in each paragraph in the order from front to back. For consecutive runs, store their text content in a list for temporary storage; when encountering non-text parts or reaching the end of the paragraph, splice the text content of all consecutive runs in the list and submit it to the large model for polishing. Call the large model for polishing: Submit the text that has passed the integrity judgment and is spliced with the prompt word to the large model for polishing to generate new text content.
[0009] Preferably, the non-text part in step 4-1 is in the format of pictures, formulas or hyperlinks.
[0010] Preferably, in step 4-2, when the end punctuation is a period ".", an exclamation mark "!" or a question mark "?", it is a complete sentence.
[0011] Preferably, the large model in step 4-4 is the Qwen-2.5 Max large model, and the prompt word is the input text provided for the user when interacting with the large model.
[0012] Preferably, step 5 is specifically as follows: Replace the polished new text content back to the corresponding position in the original text content and follow the original format of run; meanwhile, insert the non-text parts extracted from the index dictionary g1 into the polished new text content in sequence, while keeping the document structure and style unchanged, and restore the structure of the original document.
[0013] Preferably, step 6 is specifically as follows: Compare the polished new text content with the original text content, mark the text differences as revision items, generate a difference comparison document containing the revised content, and save the difference comparison document to the document server and object storage, and generate a download link for the generated difference comparison document; Step 7 is specifically as follows: Send the download link of the difference comparison document to the user page, and the user obtains the difference comparison document by clicking the download link to complete the process of document processing, optimization, and polishing.
[0014] Preferably, a device for document processing, optimization, and polishing based on a large model includes a user module, a front end, a back end, and a large model module; The front end is used for the user to upload the.docx document to be polished and download the revised difference comparison document through the user module. Meanwhile, the front end sends the uploaded.docx document to be polished to the back end; The back end saves the uploaded.docx document to be polished to the local file server and cloud object storage system, parses the uploaded.docx document to be polished, and then submits the parsed text content to the large model module; at the same time, the back end receives the new text content generated by the large model module, replaces the new text content back to the corresponding position in the original text content of the back end, and conducts document difference comparison to generate a difference comparison document. The back end sends the download link of the difference comparison document to the user page of the user module, and the user downloads the difference comparison document to complete the process of document processing, optimization, and polishing; The large model module polishes the parsed text content to form new text content.
[0015] Compared with the prior art, the advantages of this application are as follows: (1) Through a series of document processing and optimization, this application applies more advanced large model technology to document polishing. Compared with common deep learning models, it improves the context understanding ability and the consistency of professional terms, and greatly improves the polishing effect; (2) This application supports directly uploading a thesis document and polishing the document, solving the problem that ordinary large models cannot directly recognize and generate documents; (3) This application can better retain the original thesis format in the newly generated thesis document, solving the problem that large models cannot recognize and retain complex format content such as pictures, formulas, or hyperlink formats in the document. Brief Description of the Drawings
[0016] Figure 1 、Schematic flowchart of a method for document processing, optimization, and polishing based on a large model in this application; Figure 2 、Schematic structural diagram of a device for document processing, optimization, and polishing based on a large model in this application. Detailed Embodiments
[0017] The following describes the specific embodiments of this application in conjunction with the embodiments: It should be noted that the structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the implementation conditions of this application. Any modification of the structure, change of the proportional relationship, or adjustment of the size should still fall within the scope covered by the technical content disclosed in this application without affecting the effects that this application can produce and the purposes that can be achieved.
[0018] This application discloses a method for document processing, optimization, and polishing based on a large model, including the following steps: Step 1: Document upload: The user uploads the.docx document to be polished through the Web page; Step 2: Save the document: Save the uploaded.docx document to be polished; Step 3: Parse the.docx document structure: Use the lxml library to parse the underlying XML structure of the.docx document to be polished to obtain the document.xml document; Step 4: Text content extraction and polishing; Traverse the XML structure tree in the document.xml document, parse the w:r tags in the paragraphs of the XML structure tree in order, extract the text content and judge its integrity, splice the continuous text and submit it to the large model for polishing to generate new text content; Step 5: Restore the document and replace the text: Replace the polished new text content back to the corresponding positions in the original text content while keeping the document structure and style unchanged; Step 6: Document difference comparison: Compare the polished new text content with the original text content, mark the text differences as revision items, and generate and save a comparison document; Step 7: The user downloads the comparison document to complete the process of document processing, optimization, and polishing.
[0019] Preferably, the specific content of step 1 is: Through the Flask backend framework, implement a document upload interface, and the user calls the document upload interface through the Web page to upload the.docx document to be polished; The specific steps of step 2 are as follows: In the document upload interface, the.docx document to be polished received will be saved to both the local document server and the cloud object storage system (such as Alibaba Cloud OSS, Tencent Cloud COS, etc.) simultaneously to ensure file backup and access.
[0020] Preferably, the specific steps of step 3 are as follows: Unzip the.docx document to be polished through zipfile and obtain the document.xml document, which contains all the content and format information in the.docx document to be polished.
[0021] .docx document is essentially a compressed file containing multiple XML files and resources. Unzip the file through zipfile and obtain the document.xml document, which contains all the content and format information in the.docx document, such as text, formulas, pictures, hyperlinks, etc.
[0022] Preferably, the specific steps of step 4 are as follows: Step 4-1: Construct the text dictionary d1: Traverse the XML structure tree in the document.xml document. The text in the XML structure tree consists of paragraphs, and each paragraph consists of multiple runs. Retrieve according to the w:t tag, store all the XML tags of the paragraphs in an array, extract the text content of the paragraphs, and store it in the text dictionary d1. There are multiple text elements in the text dictionary d1, and after splicing the multiple text elements, it is a complete sentence; at the same time, extract the non-text part, add the corresponding index, and store it in the index dictionary g1; For example, if a certain paragraph is “Based on these indices, we selected two categories as the optimal model division.”, then the text dictionary d1 is {1: “Based on these indices,”, 2: “we selected two categories”, 3: “as the optimal model division.”}. There are 3 text elements in the text dictionary d1, and after splicing the 3 text elements, it is a complete sentence.
[0023] Traverse the XML structure tree in the document.xml document (the text in the XML structure tree consists of paragraphs, each paragraph consists of multiple runs, and the format of each run is different. For example, in a sentence "The model trajectory is depicted in Figure 1", the formats of [The model trajectory], [is depicted], and [in Figure 1] are different, so they are divided into three runs), and retrieve according to the w:t tag (in the XML structure tree, the text information is after the w:t tag, and the format information is after other tags). Store all the XML tags of the paragraphs in an array. Extract the text content (including format and attributes) of the paragraphs and store it in the text dictionary d1. For non-text parts such as pictures and formulas, add corresponding indexes and store them in the special format index dictionary g1.
[0024] Step 4-2: Integrity judgment rule: Judge whether each spliced text element forms a complete English sentence, and determine whether the text is a complete sentence by checking the end punctuation. If the spliced text is a complete sentence, splice the relevant multiple text elements together to form a complete sentence. Step 4-3: Process the runs in the paragraph: Parse and process the w:r tags in each paragraph in the order from front to back. For consecutive runs, store their text content in a list for temporary storage; when encountering non-text parts or reaching the end of the paragraph, splice the text content of all consecutive runs in the list and submit it to the large model for polishing. Parse and process the w:r tags in each paragraph in the order from front to back. For consecutive ordinary runs (such as run1, run2, run3), store their text content in a list for temporary storage; when encountering non-text runs (such as runs containing pictures, formulas, or tables) or reaching the end of the paragraph, splice the text content of all consecutive runs in the list and submit it to the large model for polishing.
[0025] Step 4-4: Call the large model for polishing: Submit the text that has passed the integrity judgment and is spliced with the prompt words to the large model for polishing to generate new text content.
[0026] Preferably, the non-text part in step 4-1 is in the form of a picture, formula, or hyperlink.
[0027] Preferably, when the end punctuation in step 4-2 is a period ".", an exclamation mark "!", or a question mark "?", it is a complete sentence.
[0028] Preferably, in step 4-4, the large model is the Qwen-2.5 Max large model, and the prompt is the input text provided to the user when interacting with the large model. Through carefully designed prompts, the model can be guided to generate outputs of specific types, topics, or formats. The design of the prompts directly affects the accuracy, relevance, and creativity of the content generated by the model.
[0029] Preferably, step 5 is specifically as follows: Replace the polished new text content back to the corresponding position in the original text content and follow the original run format; at the same time, insert the non-text parts extracted from the index dictionary g1 into the polished new text content in order, while keeping the document structure and style unchanged and restoring the structure of the original document.
[0030] Preferably, step 6 is specifically as follows: Compare the polished new text content with the original text content, mark the text differences as revision items, generate a difference comparison document containing the revised content, and save the difference comparison document to the document server and object storage, and generate a download link for the generated difference comparison document; the revision items will be highlighted for the user to view the changes.
[0031] Step 7 is specifically as follows: Send the download link of the difference comparison document to the user page, and the user obtains the difference comparison document by clicking the download link to complete the process of document processing, optimization, and polishing.
[0032] Preferably, a device for document processing, optimization, and polishing based on a large model is characterized in that it includes a user module, a front end, a back end, and a large model module; The front end is used for the user to upload the.docx document to be polished and download the revised difference comparison document through the user module. At the same time, the front end sends the uploaded.docx document to be polished to the back end; The back end saves the uploaded.docx document to be polished to the local file server and cloud object storage system, parses the uploaded.docx document to be polished, and then submits the parsed text content to the large model module; at the same time, the back end receives the new text content generated by the large model module, replaces the new text content back to the corresponding position in the original text content of the back end, and performs document difference comparison to generate a difference comparison document. The back end sends the download link of the difference comparison document to the user page of the user module, and the user downloads the difference comparison document to complete the process of document processing, optimization, and polishing; The large model module polishes the parsed text content to form new text content.
[0033] Embodiment 1 As Figure 1 shown, using the method of this application for document processing, optimization, and polishing only takes 5 minutes; The specific process is as follows: Step 1: The user uploads a file; the user calls the file upload interface through the web page to upload the file to be polished.
[0034] Step 2: Save the uploaded document; the backend receives the document, saves it to the local and remote file servers, and generates a download link for the original document.
[0035] Step 3: Parse the.docx file structure; use the lxml library to parse the underlying XML structure of the.docx file.
[0036] Step 4: Extract and polish the text content; parse the w:r tags of the paragraphs in the XML structure tree in sequence, extract the text content and judge its integrity, splice the continuous text and submit it to the large model for polishing. The generated new content replaces the original text while keeping the document structure and style unchanged.
[0037] Step 5: Restore the document and replace the text; restore the polished text and the non-text parts extracted by the g1 index to the original document in sequence to ensure the integrity of the structure, format, and style.
[0038] Step 6: Compare the differences between the documents; compare the polished document with the original text, mark the differences as revision items, and generate a visual document for comparing the differences.
[0039] Step 7: The user downloads the comparison file, sends the download link of the difference comparison file to the user's page, and the user can click the link to download the revised document to complete the entire processing flow.
[0040] Comparative Example 1 1. Using the existing large model to polish the paper in Example 1 takes a long time and is complicated to operate. It is necessary to repeatedly copy the paragraphs in the paper to the web page, wait for the large model to generate the polished paragraphs, and then copy them back to the paper document. The above operation is repeated 17 times, taking a total of 20 minutes.
[0041] Comparative Example 2 2. Using Competitor A to polish the paper in Example 1, after uploading the paper document to Competitor A, complex formats such as pictures / equations / hyperlinks are lost. After downloading the polished document to the local, it is necessary to re-insert the pictures / equations / hyperlinks of the original paper into the polished document, taking a total of 10 minutes.
[0042] The foreign English native expert group comprehensively evaluated the polishing effects of Example 1, Comparative Example 1, and Comparative Example 2 on a 100-point scale from multiple dimensions such as grammar and spelling, consistency, paper format, writing style, language fluency, and accuracy of proper nouns. The score of Example 1 is 93 points, the score of Comparative Example 1 is 89 points, and the score of Comparative Example 2 is 82 points. The effect of Example 1 is better than that of Comparative Example 1 and Comparative Example 2.
[0043] This application discloses a method for document processing, optimization, and polishing based on a large model. First, the uploaded document is parsed, then a text dictionary d1 and an index dictionary g1 are constructed. The text elements in the text dictionary d1 are concatenated, and whether the text is a complete sentence is determined by checking the end punctuation, and then it is submitted to the large model for polishing to generate new text content. The polished new text content is replaced back to the corresponding position in the original text, and at the same time, the non-text parts extracted from the index dictionary g1 are inserted into the polished new text content in order to restore the structure of the original document. Through a series of document processing and optimization, this application applies more advanced large model technology to document polishing. Compared with common deep learning models, it improves the ability to understand context and the consistency of professional terms, and greatly improves the polishing effect.
[0044] This application can first parse the.docx document to be polished uploaded by the user, use the lxml library to parse the underlying XML structure of the.docx document to be polished to obtain the document.xml document, and then traverse the XML structure tree in the document.xml document, parse the w:r tags of the paragraphs in the XML structure tree in order, extract the text content and judge its integrity, and submit the concatenated continuous text to the large model for polishing. Through the previous parsing, this application can support directly uploading a paper document and polishing the document, solving the problem that ordinary large models cannot directly recognize and generate documents.
[0045] This application parses the document and constructs a text dictionary d1 and an index dictionary g1. Images, formulas, or hyperlink formats are stored in the index dictionary g1, and only the text content in the text dictionary d1 is polished by the large model. This application can better retain the original paper format in the newly generated paper document, solving the problem that large models cannot recognize and retain complex format content such as images, formulas, or hyperlink formats in the document.
[0046] The above has made a detailed description of the preferred implementation manner of this application. However, this application is not limited to the above implementation manner. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the purpose of this application.
[0047] Many other changes and modifications can be made without departing from the concept and scope of this application. It should be understood that this application is not limited to a specific implementation manner, and the scope of this application is defined by the appended claims.
Claims
1. A document processing, optimization and polishing method based on a large model, characterized in that: The following steps are involved: Step 1: Document upload: The user uploads the .docx document to be polished through the web page; Step 2: Save the document: save the uploaded .docx document to be polished; Step 3: Parse the .docx document structure: Use the lxml library to parse the underlying XML structure of the .docx document to be polished and obtain the document.xml document; Step 4: Text content extraction and polishing: traverse the XML structure tree in the document.xml document, parse the w:r tags of the paragraphs in the XML structure tree in order, extract the text content and judge the completeness, splice the continuous text and submit it to the large model for polishing to generate new text content; Step 5: Restore the document and replace the text: Replace the polished new text content with the corresponding position in the original text content, while keeping the document structure and style unchanged; Step 6: Document difference comparison: Compare the polished new text content with the original text content, mark the text differences as revisions, generate and save the comparison document; Step 7: The user downloads the comparison document and completes the document processing, optimization, and polishing process.
2. According to claim 1, a document processing, optimization and polishing method based on a large model is characterized in that: The step 1 is specifically as follows: a document upload interface is implemented through the Flask backend framework, and the user calls the document upload interface through a web page to upload the .docx document to be polished; the step 2 is specifically as follows: in the document upload interface, the received .docx document to be polished will be saved to the local document server and the cloud object storage system at the same time.
3. According to the large model-based document processing, optimization and polishing method of claim 1, it is characterized in that: The step 3 is specifically as follows: decompress the .docx document to be polished through zipfile and obtain the document.xml document, which contains all the content and format information in the .docx document to be polished.
4. The method for document processing, optimization and polishing based on a large model according to claim 1, characterized in that: The step 4 is specifically as follows: Step 4-1: Construct text dictionary d1: traverse the XML structure tree in the document.xml document. The text in the XML structure tree consists of paragraphs, and paragraphs consist of multiple runs. Search according to the w:t tag, store the XML tags of all paragraphs into an array, extract the text content in the paragraph, and store it in the text dictionary d1. There are multiple text elements in the text dictionary d1, and multiple text elements are concatenated to form a complete sentence; at the same time, extract the non-text part, add the corresponding index, and store it in the index dictionary g1; Step 4-2: Completeness judgment rule: Determine whether each text element after splicing constitutes a complete English sentence, and determine whether the text is a complete sentence by checking the end-of-sentence symbol. If it is a complete sentence after splicing, splice the relevant multiple text elements together to form a complete sentence; Step 4-3: Process the runs in the paragraph: Parse and process the w:r tags in each paragraph in order from the front to the back. For continuous runs, store their text content in a list for temporary storage. When encountering a non-text part or reaching the end of a paragraph, concatenate the text content of all continuous runs in the list and submit it to the big model for polishing. Step 4-4: Call the big model for polishing: Submit the text that has been judged for completeness and spliced with the prompt words to the big model for polishing to generate new text content.
5. The method for document processing, optimization and polishing based on a large model according to claim 4, characterized in that: The non-text part in step 4-1 is in the form of a picture, a formula or a hyperlink.
6. A document processing, optimization and polishing method based on a large model according to claim 4, characterized in that: In step 4-2, when the end symbol of a sentence is a period ".", an exclamation mark "!" or a question mark "?", it is a complete sentence.
7. The method for document processing, optimization and polishing based on a large model according to claim 4, characterized in that: The large model in step 4-4 is the Qwen-2.5 Max large model, and the prompt word is the input text provided to the user when interacting with the large model.
8. The method for document processing, optimization and polishing based on a large model according to claim 4, characterized in that: The step 5 is specifically as follows: replacing the polished new text content back to the corresponding position in the original text content, and using the format of the original run; at the same time, inserting the non-text part extracted from the index dictionary g1 into the polished new text content in sequence, while keeping the document structure and style unchanged, and restoring the structure of the original document.
9. The method for document processing, optimization and polishing based on a large model according to claim 1, characterized in that: The step 6 specifically includes: comparing the polished new text content with the original text content, marking the text differences as revision items, generating a difference comparison document containing the revision content, and saving the difference comparison document to the document server and the object storage, and generating a download link for the generated difference comparison document; The step 7 specifically includes: sending a download link of the difference comparison document to the user page, and the user obtains the difference comparison document by clicking the download link, thereby completing the process of document processing, optimization, and polishing.
10. A document processing, optimization, and polishing device based on a large model, characterized in that: Includes user module, front-end, back-end and large model module; The front end is used for users to upload the .docx document to be polished and download the revised difference comparison document through the user module, and the front end sends the uploaded .docx document to be polished to the back end; The backend saves the uploaded .docx document to be polished to the local file server and the cloud object storage system, parses the uploaded .docx document to be polished, and then submits the parsed text content to the large model module; at the same time, the backend receives the new text content generated by the large model module, replaces the new text content back to the corresponding position in the original text content of the backend, and performs document difference comparison to generate a difference comparison document. The backend sends the download link of the difference comparison document to the user page of the user module, and the user downloads the difference comparison document, completing the document processing, optimization, and polishing process; The large model module polishes the parsed text content to form new text content.