Method and system for realizing paragraph similarity comparison based on large model Word document
By using a large-model-based method for calculating document paragraph similarity, we have overcome the shortcomings of traditional methods in capturing deep semantic information, achieving more accurate document similarity calculation. This method is applicable to fields such as academic research and intellectual property protection, and improves the efficiency and accuracy of natural language processing.
Patent Information
- Application Number
- CN202510992033.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-31
AI Technical Summary
Traditional document similarity calculation methods struggle to capture deep semantic information when dealing with complex texts, resulting in inaccurate document similarity comparisons in fields such as academic research and intellectual property protection.
A large model-based approach is adopted, which involves data preprocessing, document vectorization, and similarity calculation. Pre-trained models such as BERT and GPT are used for text vectorization, and cosine similarity, Jaccard similarity, and edit distance are used to calculate the similarity of document paragraphs.
It improves the accuracy and efficiency of document similarity calculation, especially maintaining high computing speed when processing large-scale text data, and enhances the ability to understand synonyms and sentence structure changes, making it suitable for scenarios such as online document classification and real-time information retrieval.
Smart Images

Figure CN120874805A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, specifically to a method and system for implementing paragraph similarity comparison in large-scale Word documents. Background Technology
[0002] In this era of rapid information growth, the importance of document processing and analysis is increasingly prominent. With the continuous development of big data technology, the number of documents is exploding, making the efficient management and utilization of these documents a pressing issue. This is particularly true in fields such as academic research, intellectual property protection, and information retrieval, where comparing the similarity between documents is especially crucial.
[0003] Traditional document similarity calculation methods, based on keyword matching, while meeting the requirements to some extent, still have significant shortcomings in processing complex text and capturing deep semantic information. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for implementing paragraph similarity comparison in large-scale Word documents, so as to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for implementing paragraph similarity comparison based on a large model of Word documents, comprising the following steps:
[0006] Data preprocessing stage: Extract paragraph text from Word documents, ensuring the text content is complete and accurate, and unify the encoding of all paragraphs to UTF-8 format; perform word segmentation on the extracted text, using a deep learning-based Chinese word segmentation tool for Chinese documents and NLTK and SpaCy natural language processing libraries for English documents; simultaneously, use a stop word list to remove words that contribute little to semantics; clean the text, removing special characters, numbers, punctuation marks, and repeated spaces; for English documents, use stemming or lemmatization techniques to restore words to their basic forms;
[0007] Document vectorization stage: Based on task requirements and computing resources, select a suitable pre-trained large model to convert the preprocessed text paragraphs into high-dimensional vector representations and normalize the text vectors; fine-tune the large model for specific tasks or datasets, including adjusting model parameters and adding task-specific layers.
[0008] Similarity calculation stage: Select an appropriate similarity measurement method; using the selected similarity measurement method, calculate the similarity value between two text vectors. This value is usually a floating-point number between 0 and 1, representing the degree of semantic similarity between the two texts.
[0009] Preferably, in the data preprocessing stage, for Chinese documents, a deep learning-based Chinese word segmentation tool is used for word segmentation, which divides continuous Chinese text into independent word units; for English documents, the NLTK and SpaCy natural language processing libraries are used to segment English text by spaces and punctuation marks, while stop words are removed to reduce noise.
[0010] Preferably, in the document vectorization stage, when selecting a pre-trained large model, if the task requires understanding the text context, the BERT (Bidirectional Encoder Representations from Transformers) model should be selected. This model is a powerful bidirectional encoder that captures deep semantic information of the text. If the task focuses on generating text, the GPT (Generative Pre-trained Transformer) series of models should be selected. This series of models is good at generating text.
[0011] Preferably, in the similarity calculation stage, when using the cosine similarity method, the similarity between two vectors is measured by measuring the cosine value of the angle between them. The closer the cosine value is to 1, the more similar the two vectors are. When using the Jaccard similarity method, it is used to measure the intersection ratio between sets. When using the edit distance method, it calculates the minimum number of editing operations required to convert one string into another. Permitted editing operations include replacing one character with another, inserting a character, and deleting a character.
[0012] Preferably, it also includes the results evaluation and optimization phase and the system implementation and deployment phase:
[0013] Results evaluation and optimization phase: Based on the application scenario requirements, set a similarity threshold to determine whether paragraphs are similar; evaluate the accuracy and efficiency of similarity calculation through comparative experiments and cross-validation methods, including comparison with traditional similarity calculation methods and performance on different datasets; based on the evaluation results, adjust the parameters of the large model, fine-tune strategies or similarity measurement methods to improve performance.
[0014] System Implementation and Deployment Phase: Integrate the above technical solutions into the Word document processing system to achieve automatic calculation and result display of paragraph similarity. This includes developing a user interface, processing user input, calling a large model for text vectorization, calculating similarity and returning results; designing an intuitive and easy-to-use user interface to facilitate users in uploading documents, selecting comparison options, and viewing results; and employing distributed computing and parallel processing strategies to improve processing efficiency for large-scale text data.
[0015] A system for implementing a method of comparing paragraph similarity in large-scale Word documents includes:
[0016] Data preprocessing module: It has the function of extracting paragraph text from Word documents, ensuring the integrity and accuracy of the text content, and unifying the encoding of all paragraphs into UTF-8 format; performing word segmentation on the extracted text. For Chinese documents, use a deep learning-based Chinese word segmentation tool for word segmentation. For English documents, use NLTK and SpaCy natural language processing libraries for word segmentation, and at the same time combine the stop word list to remove words with little semantic contribution; it can clean the text, removing special characters, numbers, punctuation marks, non-text content, and duplicate spaces; for English documents, use stemming or lemmatization techniques to restore words to their basic forms.
[0017] Document vectorization module: According to the task requirements and computing resources, select a suitable pre-trained large model, convert the preprocessed text paragraphs into high-dimensional vector representations, and normalize the text vectors; for a specific task or dataset, fine-tune the large model, including adjusting model parameters and adding specific task layers to obtain a more accurate text vector representation.
[0018] Similarity calculation module: Provide a function to select various similarity measurement methods, including cosine similarity, Jaccard similarity, and edit distance; according to the selected similarity measurement method, calculate the similarity value between two text vectors, which is usually a floating point number between 0 and 1, indicating the semantic similarity degree between the two texts.
[0019] Preferably, in the data preprocessing module: For Chinese document word segmentation, use a deep learning-based Chinese word segmentation tool to split continuous Chinese text into independent lexical units, and combine the stop word list to remove commonly used but meaningless words such as "de" and "le"; for English document word segmentation, use NLTK and SpaCy natural language processing libraries to split English text by spaces and punctuation marks, and at the same time remove stop words to reduce noise; in terms of stemming and lemmatization, use stemming or lemmatization techniques to restore English words to their basic forms to reduce the impact of lexical diversity on similarity calculation.
[0020] Preferably, in the document vectorization module: when selecting a pre-trained large model, if the task requires understanding the textual context, the BERT (Bidirectional Encoder Representations from Transformers) model is chosen. This model is a powerful bidirectional encoder capable of capturing deep semantic information of text. If the task focuses on generating text, the GPT (Generative Pre-trained Transformer) series of models is chosen. This series of models excels at generating text. The fine-tuning process involves adjusting the parameters of the pre-trained large model to enable it to more accurately understand textual semantics and calculate similarity. The fine-tuned model can better adapt to textual data in specific domains, thereby improving the accuracy of similarity calculation.
[0021] Preferably, in the similarity calculation module: when using the cosine similarity method, the similarity between two vectors is measured by the cosine value of the angle between them. The closer the cosine value is to 1, the more similar the two vectors are. When using the Jaccard similarity method, it is used to measure the intersection ratio between sets. When using the edit distance method, it calculates the minimum number of editing operations required to convert one string into another. Permitted editing operations include replacing one character with another, inserting a character, and deleting a character. The larger the edit distance, the more different the two strings are.
[0022] Preferably, it also includes: a result evaluation and optimization module: setting a similarity threshold to determine whether paragraphs are similar according to the application scenario requirements; evaluating the accuracy and efficiency of similarity calculation through comparative experiments and cross-validation methods, including comparison with traditional similarity calculation methods and performance on different datasets; adjusting the parameters of the large model, fine-tuning strategies or similarity measurement methods according to the evaluation results to improve performance. This process may require multiple iterations and experiments to find the optimal model configuration and similarity measurement method.
[0023] System Implementation and Deployment Module: Integrates the above technical solutions into the Word document processing system to achieve automatic calculation and result display of paragraph similarity. This includes developing a user interface, processing user input, calling a large model for text vectorization, calculating similarity and returning results; designing an intuitive and easy-to-use user interface to facilitate users in uploading documents, selecting comparison options, and viewing results; and employing distributed computing and parallel processing strategies to improve processing efficiency for large-scale text data.
[0024] Compared with the prior art, the beneficial effects of the present invention are:
[0025] This invention proposes a method and system for comparing the similarity of paragraphs in Word documents based on a large model. By introducing a pre-trained large model, this method can gain a deeper understanding of the text content and capture the deep semantic information of the text, thereby more accurately measuring the semantic similarity between paragraphs. Compared with traditional methods based on keyword matching or surface features, this method performs better in complex semantic environments.
[0026] The large model possesses powerful parallel processing capabilities, enabling this method to maintain high computational speed when processing large-scale text data. This is of great significance for applications requiring rapid processing of large numbers of documents, such as online document classification and real-time information retrieval.
[0027] By mining the deep semantic features of text, the ability to understand complex semantic phenomena such as synonyms, near-synonyms, and sentence structure variations is improved. This helps to provide more accurate and reliable text analysis results in various natural language processing tasks.
[0028] This proposed method not only overcomes the limitations of existing technologies but also provides new ideas and approaches for research in the field of natural language processing. Furthermore, its broad application prospects, such as document deduplication, plagiarism detection, and intelligent question answering, further promote the development of natural language processing technology and its application expansion across multiple fields. Attached Figure Description
[0029] Figure 1 This is a system block diagram of the present invention;
[0030] Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of the present invention clear and complete, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only some, not all, embodiments of the present invention, and are merely illustrative of the embodiments of the present invention. They are not intended to limit the embodiments of the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] Example 1, please refer to Figure 1 This invention provides a technical solution: a method for comparing the similarity of paragraphs in a large-scale Word document model, comprising the following steps:
[0033] Data preprocessing stage:
[0034] Extract paragraph text from Word documents to ensure the integrity and accuracy of the text content.
[0035] Unify the encoding of all paragraphs into UTF-8 format to reduce problems caused by encoding differences.
[0036] Perform word segmentation on the extracted text, breaking continuous text into individual words or phrases. For Chinese documents, tools such as Jieba can be used, while for English documents, it is split by spaces and punctuation marks.
[0037] Perform word form reduction or stemming on English words to reduce the impact of lexical diversity on similarity calculation.
[0038] 1) Word Segmentation and Stop Word Removal:
[0039] Perform preprocessing operations such as word segmentation, stop word removal, and stemming on Word document paragraphs to improve the accuracy and efficiency of subsequent text vectorization. For the document, character-based or word-based word segmentation methods can be used and filtered in combination with a stop word list.
[0040] Chinese documents: Use a deep learning-based Chinese word segmentation tool (such as Jieba) for word segmentation to split continuous text into independent lexical units. At the same time, combine with a stop word list to remove words that contribute less to semantics, such as commonly used but meaningless words like "de", "le", etc.
[0041] English documents: For English documents, natural language processing libraries such as NLTK and SpaCy can be used for word segmentation. Similarly, stop words also need to be removed to reduce noise.
[0042] 2) Text Cleaning:
[0043] Remove non-text content such as special characters, numbers, punctuation marks, etc. in the text, as well as duplicate spaces, etc. These non-text contents contribute less to the text semantics and may interfere with similarity calculation. Stemming and Word Form Reduction (mainly for English): For English documents, use stemming or word form reduction techniques to reduce the impact of lexical diversity on similarity calculation.
[0044] 3) Stemming and Word Form Reduction (mainly for English)
[0045] For English documents, use stemming or word form reduction techniques to reduce the impact of lexical diversity on similarity calculation. For English documents, stemming or lemmatization techniques can be used to reduce the impact of lexical diversity on similarity calculation. This helps to reduce the impact of lexical diversity on similarity calculation because different forms of the same word are usually semantically similar.
[0046] Document vectorization stage:
[0047] Pre-trained large models (such as Word2Vec, BERT, GPT, etc.) are used to convert pre-processed text paragraphs into high-dimensional vector representations. These vectors can capture the deep semantic information of the text, providing a foundation for similarity calculation.
[0048] Text vectors are normalized to ensure the stability and accuracy of similarity calculations. For specific tasks or datasets, large models can be fine-tuned to improve performance. This includes adjusting model parameters and adding task-specific layers.
[0049] The preprocessed paragraph text is input into a large model for encoding to obtain its high-dimensional vector representation. These vectors capture the deep semantic information of the text and can be used for subsequent similarity calculations.
[0050] 1) Choose a large model
[0051] Choose a suitable pre-trained large model based on task requirements and computational resources. For example, BERT is suitable for understanding textual contextual relationships, while the GPT series excels at text generation.
[0052] 2) Model fine-tuning
[0053] Fine-tuning is the process of retraining a large model on task-specific data to better adapt it to that task. By fine-tuning, the pre-learned knowledge and skills of the large model can be leveraged to improve its performance on a specific task.
[0054] In document similarity comparison modules, fine-tuning typically involves adjusting the parameters of a large, pre-trained model to make it more accurate in understanding text semantics and calculating similarity. For specific tasks or datasets, large models can be fine-tuned to improve performance; this includes adjusting model parameters, adding task-specific layers, and so on.
[0055] The fine-tuned model is better able to adapt to text data in specific domains, thereby improving the accuracy of similarity calculation.
[0056] 3) Text encoding
[0057] The preprocessed text is input into a large model to obtain its high-dimensional vector representation. These vectors capture the deep semantic information of the text and can be used for subsequent similarity calculations.
[0058] Similarity calculation stage:
[0059] The similarity between two vectors is measured by taking the cosine of the angle between them. The cosine of the angle indicates whether the two vectors share the same direction; the closer the cosine is to 1, the closer the angle is to 0 degrees, and the more similar the two vectors are. Cosine similarity is particularly useful for measuring directional consistency between vectors and is a commonly used method in text similarity calculations.
[0060] Extract several keywords from each article, merge them into a set, and calculate the word frequency of each article for the words in this set (relative word frequency can be used to avoid differences in article length).
[0061] Convert the word frequencies calculated for each of the two articles into word frequency vectors.
[0062] Based on the word frequency vectors of the two articles, calculate the cosine similarity between the two vectors. The larger the value, the more similar the two articles are.
[0063] 1) Choosing a similarity measurement method
[0064] Choose an appropriate similarity measurement method based on the application scenario. Commonly used similarity measurement methods include cosine similarity, Jaccard similarity, and edit distance. The minimum number of edit operations required to transform one string into another. The larger the edit distance between two strings, the more different they are. Permitted edit operations include replacing one character with another, inserting a character, and deleting a character.
[0065] 2) Calculate similarity
[0066] Cosine similarity is suitable for measuring the directional consistency between vectors; Jaccard similarity is suitable for measuring the intersection ratio between sets; and edit distance is suitable for measuring character-level differences between texts.
[0067] Calculate similarity: Using the selected similarity metric, calculate the similarity value between the two text vectors.
[0068] Using the selected similarity metric, a similarity value is calculated between two text vectors. This value is typically a floating-point number between 0 and 1, representing the degree of semantic similarity between the two texts.
[0069] Results evaluation and optimization phase:
[0070] Ensure the objectivity and impartiality of the evaluation methods, and avoid the influence of subjective factors and human interference on the evaluation results.
[0071] A comprehensive evaluation using multiple assessment methods can improve the accuracy and reliability of the evaluation results.
[0072] 1) Set threshold
[0073] Depending on the application scenario, a similarity threshold can be set to determine whether paragraphs are similar. For example, in plagiarism detection, a higher threshold can be set to ensure that only highly similar paragraphs are marked as plagiarism; while in document classification, a lower threshold can be set to allow for a certain degree of diversity.
[0074] 2) Performance Evaluation
[0075] The accuracy and efficiency of similarity calculation were evaluated using methods such as comparative experiments and cross-validation. This included comparisons with traditional similarity calculation methods and performance on different datasets.
[0076] 3) Model optimization
[0077] Based on the evaluation results, adjust the parameters of the large model, fine-tune the strategy, or the similarity metric to improve performance. This may require multiple iterations and experiments to find the optimal model configuration and similarity metric.
[0078] System implementation and deployment phase:
[0079] Integrating the above technical solutions into the Word document processing system enables automatic calculation and result display of paragraph similarity. This includes developing a user interface, processing user input, calling a large model for text vectorization, calculating similarity, and returning the results.
[0080] System Integration: Integrate the above technical solutions into the Word document processing system to achieve automatic calculation and result display of paragraph similarity.
[0081] User interface design: Design an intuitive and easy-to-use user interface that allows users to upload documents, select comparison options, and view results.
[0082] Performance optimization: For large-scale text data, strategies such as distributed computing and parallel processing are adopted to improve processing efficiency.
[0083] Example 2, based on Example 1, provides a system for implementing a method of comparing paragraph similarity in a large-scale Word document model, comprising:
[0084] Data Preprocessing Module: It has the function of extracting paragraph text from Word documents, ensuring the integrity and accuracy of the text content, and unifying the encoding of all paragraphs into UTF-8 format; performing word segmentation on the extracted text, using a deep learning-based Chinese word segmentation tool for Chinese documents and NLTK and SpaCy natural language processing libraries for English documents, and at the same time removing words with little semantic contribution in combination with a stop word list; cleaning the text to remove special characters, numbers, punctuation marks, non-text content, and duplicate spaces; for English documents, using stemming or lemmatization techniques to restore words to their basic forms; for Chinese documents, using a deep learning-based Chinese word segmentation tool to segment continuous Chinese text into independent lexical units and removing commonly used but meaningless words such as "de" and "le" in combination with a stop word list; for English documents, using NLTK and SpaCy natural language processing libraries to split English text by spaces and punctuation marks and removing stop words to reduce noise; in terms of stemming and lemmatization, using stemming or lemmatization techniques to restore English words to their basic forms to reduce the impact of lexical diversity on similarity calculation.
[0085] Document Vectorization Module: According to the task requirements and computing resources, select a suitable pre-trained large model to convert the preprocessed text paragraphs into high-dimensional vector representations and normalize the text vectors; for a specific task or dataset, fine-tune the large model, including adjusting model parameters and adding specific task layers to obtain a more accurate text vector representation; when selecting a pre-trained large model, if the task requires understanding the context relationship of the text, then select the BERT: Bidirectional Encoder Representations from Transformers model, which is a powerful bidirectional encoder that can capture the deep semantic information of the text; if the task focuses on text generation, then select the GPT: Generative Pre-trained Transformer series of models, which are good at text generation; the fine-tuning process involves adjusting the parameters of the pre-trained large model to make it better understand the text semantics and calculate similarity, and the fine-tuned model can better adapt to the text data in a specific field, thereby improving the accuracy of similarity calculation.
[0086] The similarity calculation module offers multiple similarity measurement methods, including cosine similarity, Jaccard similarity, and edit distance. Based on the selected method, it calculates the similarity value between two text vectors. This value is typically a floating-point number between 0 and 1, representing the semantic similarity between the two texts. When using cosine similarity, it measures the similarity by the cosine of the angle between the two vectors; the closer the cosine value is to 1, the more similar the two vectors are. When using Jaccard similarity, it measures the intersection ratio between sets. When using edit distance, it calculates the minimum number of edit operations required to transform one string into another. Permitted edit operations include replacing one character with another, inserting a character, and deleting a character; the larger the edit distance, the more different the two strings are.
[0087] It also includes: Result evaluation and optimization module: Based on the application scenario requirements, a similarity threshold is set to determine whether paragraphs are similar; the accuracy and efficiency of similarity calculation are evaluated through comparative experiments and cross-validation methods, including comparison with traditional similarity calculation methods and performance on different datasets; based on the evaluation results, the parameters of the large model, fine-tuning strategies or similarity measurement methods are adjusted to improve performance. This process may require multiple iterations and experiments to find the optimal model configuration and similarity measurement method.
[0088] System Implementation and Deployment Module: Integrates the above technical solutions into the Word document processing system to achieve automatic calculation and result display of paragraph similarity. This includes developing a user interface, processing user input, calling a large model for text vectorization, calculating similarity and returning results; designing an intuitive and easy-to-use user interface to facilitate users in uploading documents, selecting comparison options, and viewing results; and employing distributed computing and parallel processing strategies to improve processing efficiency for large-scale text data.
[0089] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for implementing paragraph similarity comparison based on a large model of Word documents, characterized in that: Includes the following steps: Data preprocessing stage: Extract paragraph text from Word documents, ensuring the text content is complete and accurate, and unify the encoding of all paragraphs to UTF-8 format; perform word segmentation on the extracted text, using a deep learning-based Chinese word segmentation tool for Chinese documents and NLTK and SpaCy natural language processing libraries for English documents; simultaneously, use a stop word list to remove words that contribute little to semantics; clean the text, removing special characters, numbers, punctuation marks, and repeated spaces; for English documents, use stemming or lemmatization techniques to restore words to their basic forms; Document vectorization stage: Based on task requirements and computing resources, select a suitable pre-trained large model to convert the preprocessed text paragraphs into high-dimensional vector representations and normalize the text vectors; fine-tune the large model for specific tasks or datasets, including adjusting model parameters and adding task-specific layers. Similarity calculation stage: Select an appropriate similarity measurement method; using the selected similarity measurement method, calculate the similarity value between two text vectors. This value is usually a floating-point number between 0 and 1, representing the degree of semantic similarity between the two texts.
2. The method for implementing paragraph similarity comparison in a large-scale Word document according to claim 1, characterized in that: In the data preprocessing stage, deep learning-based Chinese word segmentation tools are used for word segmentation of Chinese documents. These tools divide continuous Chinese text into independent word units. For word segmentation of English documents, the NLTK and SpaCy natural language processing libraries are used to segment English text by spaces and punctuation marks, while stop words are removed to reduce noise.
3. The method for implementing paragraph similarity comparison based on a large model of Word documents according to claim 2, characterized in that: In the document vectorization stage, when selecting a pre-trained large model, if the task requires understanding the textual context, then BERT (Bidirectional Encoder Representations from Transformers) is chosen. This model is a powerful bidirectional encoder that captures deep semantic information of the text. If the task focuses on generating text, then GPT (Generative Pre-trained Transformer) is chosen. This series of models excels at generating text.
4. The method for implementing paragraph similarity comparison based on a large model of Word documents according to claim 3, characterized in that: In the similarity calculation stage, when using the cosine similarity method, the similarity between two vectors is measured by the cosine value of the angle between them. The closer the cosine value is to 1, the more similar the two vectors are. When using the Jaccard similarity method, it is used to measure the proportion of intersection between sets. When using the edit distance method, it calculates the minimum number of editing operations required to transform one string into another. Permitted editing operations include replacing one character with another, inserting a character, and deleting a character.
5. The method for implementing paragraph similarity comparison in a large-scale Word document according to claim 4, characterized in that: It also includes the results evaluation and optimization phase, as well as the system implementation and deployment phase: Results evaluation and optimization phase: Based on the application scenario requirements, a similarity threshold is set to determine whether paragraphs are similar; the accuracy and efficiency of similarity calculation are evaluated through comparative experiments and cross-validation methods, including comparison with traditional similarity calculation methods and performance on different datasets; Adjust the parameters, fine-tuning strategies, or similarity measurement methods of the large model according to the evaluation results to improve performance; System implementation and deployment phase: Integrate the above technical solutions into the Word document processing system to achieve automatic calculation and result display of paragraph similarity, including developing a user interface, processing user input, calling the large model for text vectorization, calculating similarity, and returning results; Design an intuitive and user-friendly interface to facilitate users to upload documents, select comparison options, and view results; For large-scale text data, adopt distributed computing and parallel processing strategies to improve processing efficiency.
6. A system for implementing the method of paragraph similarity comparison in a large-scale Word document according to claim 5, characterized in that: Including: Data preprocessing module: It has the function of extracting paragraph text from Word documents to ensure the integrity and accuracy of text content, and unified encoding of all paragraphs into UTF-8 format; Perform word segmentation on the extracted text. For Chinese documents, use a deep learning-based Chinese word segmentation tool for word segmentation, and for English documents, use NLTK and SpaCy natural language processing libraries for word segmentation. At the same time, combine the stop word list to remove words with little semantic contribution; The text can be cleaned to remove special characters, numbers, punctuation marks, non-text content, and duplicate spaces; For English documents, use stemming or lemmatization techniques to restore words to their basic forms. Document vectorization module: Select a suitable pre-trained large model according to task requirements and computing resources, convert the preprocessed text paragraphs into high-dimensional vector representations, and normalize the text vectors; For specific tasks or datasets, fine-tune the large model, including adjusting model parameters and adding specific task layers to obtain a more accurate text vector representation. Similarity calculation module: Provide a function to select various similarity measurement methods, including cosine similarity, Jaccard similarity, and edit distance; Calculate the similarity value between two text vectors according to the selected similarity measurement method. This value is usually a floating point number between 0 and 1, indicating the semantic similarity degree between the two texts.
7. The system according to claim 6, characterized in that: In the data preprocessing module: For Chinese document word segmentation, use a deep learning-based Chinese word segmentation tool to split continuous Chinese text into independent lexical units, and combine the stop word list to remove commonly used but meaningless words such as "的" and "了"; For English document word segmentation, use NLTK and SpaCy natural language processing libraries to split English text by spaces and punctuation marks, and remove stop words to reduce noise; In terms of stemming and lemmatization, use stemming or lemmatization techniques to restore English words to their basic forms to reduce the impact of lexical diversity on similarity calculation.
8. A system according to claim 6, characterized in that: In the document vectorization module: when selecting a pre-trained large model, if the task requires understanding the textual context, the BERT (Bidirectional Encoder Representations from Transformers) model is chosen. This model is a powerful bidirectional encoder capable of capturing deep semantic information of text. If the task focuses on generating text, the GPT (Generative Pre-trained Transformer) series of models is chosen. This series of models excels at generating text. The fine-tuning process involves adjusting the parameters of the pre-trained large model to enable it to more accurately understand textual semantics and calculate similarity. The fine-tuned model can better adapt to textual data in specific domains, thereby improving the accuracy of similarity calculation.
9. A system according to claim 6, characterized in that: In the similarity calculation module: when using the cosine similarity method, the similarity between two vectors is measured by the cosine value of the angle between them. The closer the cosine value is to 1, the more similar the two vectors are. When using the Jaccard similarity method, it is used to measure the proportion of intersection between sets. When using the edit distance method, it calculates the minimum number of editing operations required to transform one string into another. Permitted editing operations include replacing one character with another, inserting a character, and deleting a character. The larger the edit distance, the more different the two strings are.
10. A system according to claim 6, characterized in that: Also includes: Results Evaluation and Optimization Module: Based on application scenario requirements, set similarity thresholds to determine whether paragraphs are similar; evaluate the accuracy and efficiency of similarity calculation through comparative experiments and cross-validation methods, including comparison with traditional similarity calculation methods and performance on different datasets; based on the evaluation results, adjust the parameters of the large model, fine-tune strategies or similarity measurement methods to improve performance. This process may require multiple iterations and experiments to find the optimal model configuration and similarity measurement method. System Implementation and Deployment Module: Integrates the above technical solutions into the Word document processing system to realize the automatic calculation and result display of paragraph similarity, including developing the user interface, processing user input, calling a large model for text vectorization, calculating similarity and returning the results; The user interface is designed to be intuitive and easy to use, allowing users to upload documents, select comparison options, and view results. For large-scale text data, distributed computing and parallel processing strategies are adopted to improve processing efficiency.