Intelligent document duplicate checking system and method based on vector database and large language model

By combining large language model and vector database, using self-supervised learning and dynamic feedback optimization technology, high-dimensional semantic vectors are generated and hierarchical compressed index construction is solved, and the traditional document plagiarism checking method has high error rate and low computational efficiency in complex semantic matching is achieved, and an efficient and accurate document plagiarism checking system is realized.

CN120373283AInactive Publication Date: 2025-07-25BEIJING RONGJIA HECHUANG TECHNOLOGY CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510449121.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing document plagiarism checking technology has high misjudgment rate and low computational efficiency when dealing with complex semantic matching, and it is difficult to adapt to changes in different document types and user needs. Traditional methods cannot effectively capture the deep semantic relationships of text, resulting in unstable plagiarism checking results.

Method used

Combining large language models and vector databases, high-dimensional semantic vectors are generated through self-supervised learning and dynamic feedback optimization, hierarchical compression and index construction are carried out, fine-grained similarity calculation is performed based on outlier recognition, cosine similarity and Mahayana distance, and fine-tuning of the model is used to achieve dynamic index updates.

Benefits of technology

It improves the accuracy and efficiency of document plagiarism checking, reduces computing resource consumption, adapts to changes in different document types and user needs, provides self-learning and continuous optimization capabilities, and significantly improves the response speed and accuracy of the plagiarism checking system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373283A_ABST
    Figure CN120373283A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent document duplicate checking system and method based on a vector database and a large language model, and the method comprises the following steps: S1, collecting document data in various formats, and carrying out the preprocessing of the document data; s2, semantic coding is carried out through a large language model, and a document semantic vector is generated; s3, storing the document semantic vector into a vector database, constructing a vector index and recording historical query data; s4, carrying out preliminary candidate document retrieval, and carrying out approximate nearest neighbor retrieval based on outlier identification; s5, calculating the similarity between the candidate document and the document to be subjected to duplicate checking, and screening a final high-similarity document; s6, generating a duplicate checking report, and recording user operation behaviors; and S7, receiving user feedback, and dynamically updating the document semantic vector and the vector index. According to the method, efficient and accurate intelligent document duplicate checking is realized by utilizing the large language model and the vector database, the semantic matching capability is improved, the duplicate checking efficiency is optimized, and the intelligence and adaptability of a duplicate checking system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent document processing and information retrieval, and particularly to an intelligent document duplicate checking system and method based on a vector database and a large language model. Background Art

[0002] Currently, with the explosive growth of Internet information and the wide application of digital document data, document duplicate checking technology plays an increasingly important role in many fields such as academia, enterprises, law, and media. Traditional document duplicate checking methods are mostly based on keyword matching, string comparison, or algorithms based on statistical features. These methods can detect direct repetition or simple plagiarism of texts to a certain extent, but they have obvious limitations. Since these traditional methods mainly rely on surface features and cannot effectively capture the deep semantic relationships of texts, misjudgment or missed judgment often occurs when dealing with complex situations such as synonymous substitution, word order change, paragraph reconstruction, or language style conversion. In addition, with the rapid growth of the number of documents and the scale of data, the requirements of traditional algorithms for computing efficiency and storage resources are becoming increasingly severe, and it is difficult to achieve fast and efficient duplicate checking functions on large-scale document data sets.

[0003] In recent years, deep learning technology, especially large language models, has made breakthrough progress in the field of natural language processing. Large language models such as the GPT series, BERT, and their derivative models have demonstrated powerful capabilities in semantic understanding, text generation, and context modeling, enabling the deep semantic features of texts to be more accurately expressed through vector representation. At the same time, the development of vector database technology provides an effective means for the storage and rapid retrieval of high-dimensional vector data, making it possible to use duplicate checking methods based on semantic vectors. However, there are still many deficiencies in the existing technology in practical applications. First, although large language models can generate high-quality semantic vectors, in large-scale document duplicate checking tasks, how to efficiently manage and retrieve these high-dimensional vectors is still a challenge. Although existing vector databases have improved in storage and retrieval speed, they often lack effective judgment of local density and abnormal situations when dealing with complex semantic matching problems, resulting in deviations in semantic similarity evaluation. Second, in the process of constructing semantic vectors in the existing technology, the deep combination of self-supervised learning and contrastive learning is usually lacking, making it difficult to fully explore the implicit semantic information between texts, thus affecting the accuracy and robustness of duplicate checking. In addition, due to the diversity of document types, language styles, and domain backgrounds, a single duplicate checking model often has difficulty adapting to the duplicate checking needs in different scenarios, resulting in insufficient generality and adaptability of the duplicate checking results.

[0004] In some advanced systems, although retrieval algorithms based on outlier recognition are introduced for preliminary screening of candidate documents, and metrics such as local outlier factors are used to measure the abnormality of candidate documents in the semantic space, these methods often only stay in the preliminary screening stage and do not further optimize the retrieval algorithm by combining historical query data. There is a lack of dynamic adjustment and adaptive optimization mechanisms for retrieval results. Traditional static retrieval methods are prone to problems such as unstable duplicate checking effects and high misjudgment rates when facing constantly changing user needs and document data distributions. In addition, although some studies attempt to improve system performance through dynamic model fine-tuning, they often do not combine the semantic representation ability of large language models with the dynamic update of vector indexes in vector databases, and cannot achieve the overall collaborative optimization of the duplicate checking system, resulting in the system still being difficult to meet the requirements of practical applications in terms of duplicate checking efficiency and accuracy.

[0005] Therefore, how to provide an intelligent document duplicate checking system and method based on a vector database and a large language model is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0006] An object of the present invention is to propose an intelligent document duplicate checking system and method based on a vector database and a large language model. The present invention makes full use of advanced technologies such as the deep semantic encoding of large language models, the efficient storage and retrieval of vector databases, self-supervised learning, and dynamic feedback optimization, and details key steps such as the generation, hierarchical compression, index construction, preliminary retrieval of candidate documents, fine-grained similarity calculation, and dynamic model fine-tuning based on user feedback of document semantic vectors. The present invention can accurately capture the deep semantic features of documents, efficiently process large-scale document data, and at the same time has the ability of self-learning and dynamic optimization, making the duplicate checking results more accurate and stable, adapting to various application scenarios, and significantly improving the speed and efficiency of document duplicate checking.

[0007] According to the intelligent document duplicate checking method based on a vector database and a large language model according to an embodiment of the present invention, the method includes the following steps:

[0008] S1. Collect document data in various formats and preprocess the document data;

[0009] S2. Semantically encode the preprocessed document data through a large language model to generate document semantic vectors, and pre-train the large language model;

[0010] S3. Store the document semantic vectors in a vector database, hierarchically compress the document semantic vectors, construct a vector index, and record historical query data;

[0011] S4. Receive the document to be checked for duplication, generate a query vector through a large language model, call the constructed vector index, perform approximate nearest neighbor search based on the outlier recognition-based retrieval algorithm, screen out a set of candidate documents with high preliminary similarity, and optimize the outlier recognition-based retrieval algorithm in combination with historical query data;

[0012] S5. For each document vector in the set of candidate documents with high preliminary similarity, calculate the semantic similarity with the document vector to be checked for duplication, combine the cosine similarity and the Mahalanobis distance to obtain a similarity score, and screen out the documents with high similarity;

[0013] S6. Generate a duplication check report based on the screened documents with high similarity, display the list of similar documents, similarity scores, highlighted duplicate content, and semantic analysis results through a graphical interface, and at the same time record the user's operation behavior in the duplication check report into the historical query data;

[0014] S7. Receive the user's feedback on the duplication check results, dynamically adjust the semantic representation ability of the large language model through the fine-tuning mechanism of the large language model in combination with self-supervised learning technology, and update the document semantic vectors and vector index in the vector database.

[0015] Optionally, the specific steps of S3 are as follows:

[0016] S31. Bind the document semantic vectors generated by the large language model with the corresponding document identifiers, store them in the vector database, and form a mapping relationship between the documents and the semantic vectors;

[0017] S32. Perform hierarchical compression processing on the stored document semantic vectors, divide the high-dimensional semantic vectors into multiple sub-vectors, and compress each sub-vector to obtain the low-dimensional representations of each subspace;

[0018] S33. Build an index for each sub-vector after hierarchical quantization, and define the index function as I(v):

[0019]

[0020] where C i is the representation obtained by quantization compression of the i-th subspace, w i is the dynamic weight factor, γ i is the auxiliary adjustment coefficient, ∈ and η are small constants, Z is the normalization coefficient, k represents the total number of sub-vectors, and v is the document semantic vector;

[0021] S34. During the index construction process, adaptively adjust the dynamic weight factors of each subspace according to the retrieval performance of each sub-vector in the historical query data;

[0022] S35. While constructing the vector index, record the historical data of each query operation in real time:

[0023]

[0024] where q is the query vector, sim(q, v i ) represents the similarity between the query vector and vector v i , θ is the preset similarity threshold, R is the retrieval hit rate, S is the set of similarity score distributions, O is the outlier factor score, N is the total number of candidate documents, and H is the historical query data;

[0025] S36. Associate and store the recorded historical query data H with the constructed vector index to build an optimized data set:

[0026]

[0027] where β is the adjustment coefficient, U(v) represents the vector index optimized by combining historical data, and sim(q, v) represents the similarity calculation function between the query vector q and the document semantic vector v.

[0028] Optionally, the specific steps of S4 include:

[0029] S41. Receive the document D to be checked for duplication Q , and convert the document D to be checked for duplication into a query vector q through a large language model;

[0030] S42. Call the constructed vector index I(v) to calculate the weighted Euclidean distance d(q, I(v)) between the query vector q and the index I(v) of each document:

[0031]

[0032] where ω j is the weight factor, q j is the j-th query vector, I(v) j is the j-th vector index, and n is the dimension of the vector space;

[0033] S43. Construct a retrieval algorithm based on outlier recognition to calculate the local outlier factor for each candidate document v:

[0034]

[0035] where N(v) is the local neighborhood, ζ is the smoothing constant, LOF(v) represents the local outlier factor of the candidate document v, u is the neighbor document, d(q, I(u)) represents the weighted Euclidean distance between the query vector q and the vector index I(u) of the neighbor document u, and exp is the exponential function;

[0036] S44. Combine the weighted Euclidean distance d(q, I(v)) with the local outlier factor LOF(v) to construct a comprehensive distance function for sorting and screening candidate documents:

[0037] d′(q, I(v)) = d(q, I(v))·[1 + λ·ln(1 + LOF(v))];

[0038] where λ is an adjustment coefficient, d′(q, I(v)) is the comprehensive distance, and ln is the logarithmic function;

[0039] S45. Optimize the comprehensive distance of candidate document v in combination with the recorded historical query data H:

[0040]

[0041] where β is the historical data adjustment coefficient, d″(q, I(v)) represents the optimized distance, H(v) represents the set of historical query data related to document v, q′ is the historical query vector, θ is the preset similarity threshold, and ||||2 represents the Euclidean norm;

[0042] S46. Sort all the documents in the vector database in ascending order according to the optimized distance d″(q, I(v)), and screen out the candidate document set S with high preliminary similarity C :

[0043] S C = {v ∈ D | d″(q, I(v)) ≤ τ};

[0044] where τ is the preset distance threshold, and D represents all the stored document semantic vectors in the vector database.

[0045] Optionally, the S5 specifically includes:

[0046] S51. Calculate the weighted cosine similarity sim n (q, v) between the semantic vector v ∈ R of each candidate document in the obtained candidate document set with high preliminary similarity and the query vector q ∈ R generated by the document to be checked for duplication n : wcos (q, v):

[0047]

[0048] where ω j is the weight factor, q j is the component of the semantic vector q representing the query document in the j-th dimension, v j represents the component value of the candidate document semantic vector v in the j-th dimension, and n represents the dimension of the semantic vector;

[0049] S52. Calculate the adaptive Mahalanobis distance between the semantic vector v and the query vector q

[0050]

[0051] where Σ is the covariance matrix, β o is the adjustment coefficient, and diag(ω) represents a diagonal matrix;

[0052] S53. Combine the weighted cosine similarity and the adaptive Mahalanobis distance to construct the comprehensive similarity score of the candidate documents:

[0053]

[0054] where α o is the weight coefficient, and S(v) is the comprehensive similarity score of the candidate document v;

[0055] S54. Use the local neighborhood N(v) of each document v in the candidate document set to calculate the refined local density factor RLD(v):

[0056]

[0057] where LOF(v) represents the local outlier factor of the candidate document v, and I(u) is the vector index of the candidate document u;

[0058] S55. Screen each document v in the candidate document set according to the comprehensive similarity score S(v) and the refined local density factor RLD(v), and select those that satisfy:

[0059] S(v)≥τ S and RLD(v)≤τ RLD ;

[0060] The high-similarity documents are used as the final duplicate check result, where τ S is the preset comprehensive similarity threshold, and τ RLD is the preset local density factor threshold.

[0061] Optionally, the S6 specifically includes:

[0062] S61. Generate a duplicate check report based on the selected high-similarity documents, including the basic information of the document to be checked, the list of similar documents, the similarity score, the detected duplicate content segments, and the semantic analysis results generated by the system;

[0063] S62. In the duplicate check report, display the list of similar documents through a graphical interface, sorted in descending order of similarity, and provide key attribute information for each similar document, including the title, source, similarity score, number of duplicate paragraphs, and similarity segment position index;

[0064] S63. Automatically mark the duplicate content in the document to be checked for plagiarism, highlight duplicate content of different degrees with different colors or formats, and provide corresponding annotation information, such as the similar document number, matching position, and similarity score;

[0065] S64. Combine semantic analysis in the process of checking for plagiarism, and display a detailed assessment of the text similarity by the plagiarism detection system, including semantic reconstruction analysis, syntactic structure change detection, and cross-paragraph semantic comparison;

[0066] S65. Provide user interaction functions on the plagiarism detection report interface, allowing users to manually adjust the similarity threshold, filter specific types of similar documents, mark misjudged or correctly matched content, and provide feedback options;

[0067] S66. Record the operation behaviors of users in the plagiarism detection report and store them in the historical query data.

[0068] Optionally, the specific content of S7 is as follows:

[0069] S71. Receive the feedback data of users on the plagiarism detection results, record the operation behaviors of users on the plagiarism detection report interface, and form a feedback data set;

[0070] S72. Integrate the collected user feedback data with the historical query data to construct a comprehensive feedback set;

[0071] S73. Calculate the error amount between the system prediction output and the user feedback based on the comprehensive feedback set to form a feedback error index;

[0072] S74. Use self-supervised learning technology and dynamic feedback error to fine-tune the parameters of the large language model:

[0073]

[0074] where θ old represents the parameters of the large language model before update, θ new represents the parameters of the large language model after update, η is the adaptive learning rate, m t is the first-order momentum term, v t is the second-order momentum term, ∈ is a small constant to prevent division by zero, ρ is the feedback error adjustment coefficient, and ΔF is the average feedback error;

[0075] S75. Use the updated parameters θ of the large language model new to regenerate the semantic vectors for the document set and obtain the updated document representation;

[0076] S76. Update the index in the vector database according to the updated semantic vectors using a dynamic index update strategy:

[0077]

[0078] Among them, I(v) is the index of the original document vector, and v new is the semantic vector of the document generated after update, and f update (v new ) is the new index value calculated based on the updated semantic vector, is the adjustment coefficient, and I(v new ) is the vector index after dynamic update.

[0079] The intelligent document duplicate checking system based on a vector database and a large language model according to an embodiment of the present invention includes the following modules:

[0080] The data collection and preprocessing module is used to extract, parse documents in various formats and preprocess the documents;

[0081] The large language model semantic encoding module uses the large language model to convert the preprocessed documents into high-dimensional semantic vectors and capture deep semantic features;

[0082] The vector database and index construction module stores the document semantic vectors and constructs a hierarchical index, and improves the retrieval efficiency through quantization and compression technologies;

[0083] The preliminary candidate document retrieval module uses a retrieval algorithm based on outlier recognition and historical query data optimization to quickly screen out candidate documents with high semantic similarity to the document to be checked for duplicates;

[0084] The fine-grained similarity calculation module combines weighted cosine similarity, adaptive Mahalanobis distance, and local density factor to perform fine similarity calculation and screening on the candidate documents;

[0085] The duplicate checking report generation and display module generates a duplicate checking report and displays similar documents, similarity scores, highlighted duplicate content, and semantic analysis results through a graphical interface;

[0086] The dynamic feedback and model optimization module receives user feedback and performs real-time fine-tuning on the parameters of the large language model and the vector index in combination with self-supervised learning technology.

[0087] The beneficial effects of the present invention are:

[0088] The present invention combines large language models with vector database technology, as well as methods such as self-supervised learning, outlier identification, and dynamic feedback optimization, to achieve precise capture and efficient retrieval of the deep semantic features of documents. In the problem that traditional document duplication detection methods cannot accurately handle semantic reconstruction, synonym replacement, or local text transformation, the present invention uses large language models to generate high-dimensional semantic vectors, and significantly improves the storage and retrieval efficiency of large-scale document data through hierarchical compression and dynamic indexing construction technology, enabling the system to maintain high duplication detection accuracy and effectively reduce the consumption of computing resources when facing complex text similarity detection tasks.

[0089] In addition, during the preliminary screening of candidate documents and the fine-grained similarity calculation process, the present invention introduces a retrieval algorithm based on outlier identification, and combines multiple metric indicators such as cosine similarity, adaptive Mahalanobis distance, and local density factor to achieve multi-dimensional evaluation of document similarity, thereby quickly screening out candidate documents with high semantic similarity in the preliminary retrieval stage and further improving the accuracy of the duplication detection results in the subsequent refined screening. Through the dynamic integration and feedback mechanism of historical query data, the present invention can continuously perform model fine-tuning to adapt to changes in different document types and user needs, ensuring that the duplication detection system has the ability of self-learning and continuous optimization.

[0090] Therefore, the present invention not only improves the accuracy and recall rate of document duplication detection, but also significantly enhances the response speed and processing efficiency of the system, reducing the risks of misjudgment and missed judgment. Its intelligent and adaptive duplication detection mechanism and dynamic optimization ability provide users with an efficient, accurate, and easily expandable document duplication detection solution, which has broad application prospects and obvious market competitive advantages in multiple fields such as academia, enterprises, law, and content review. Brief Description of the Drawings

[0091] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0092] Figure 1 is a flowchart of the intelligent document duplication detection method based on vector database and large language model proposed by the present invention;

[0093] Figure 2 is a schematic structural diagram of the intelligent document duplication detection system based on vector database and large language model proposed by the present invention. Detailed Description of the Invention

[0094] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0095] Reference Figure 1 , an intelligent document duplicate checking method based on a vector database and a large language model, includes the following steps:

[0096] S1. Collect document data in multiple formats and preprocess the document data;

[0097] S2. Semantically encode the preprocessed document data through a large language model to generate document semantic vectors, and pre-train the large language model;

[0098] S3. Store the document semantic vectors in a vector database, hierarchically compress the document semantic vectors, construct a vector index, and record historical query data;

[0099] S4. Receive the document to be checked for duplicates, generate a query vector through a large language model, call the constructed vector index, perform approximate nearest neighbor retrieval based on an outlier recognition-based retrieval algorithm, screen out a set of candidate documents with a high initial similarity, and optimize the outlier recognition-based retrieval algorithm in combination with historical query data;

[0100] S5. Calculate the semantic similarity between each document vector in the set of candidate documents with a high initial similarity and the document vector to be checked for duplicates, combine the cosine similarity and the Mahalanobis distance to obtain a similarity score, and screen out the documents with a high similarity;

[0101] S6. Generate a duplicate check report based on the screened documents with a high similarity, display the list of similar documents, similarity scores, highlighted duplicate content, and semantic analysis results through a graphical interface, and record the user's operation behavior in the duplicate check report in the historical query data at the same time;

[0102] S7. Receive the user's feedback on the duplicate check results, dynamically adjust the semantic representation ability of the large language model through the fine-tuning mechanism of the large language model in combination with self-supervised learning technology, and update the document semantic vectors and vector index in the vector database.

[0103] In this embodiment, the specific steps of S3 include:

[0104] S31. Bind the document semantic vectors generated by the large language model to the corresponding document identifiers, store them in the vector database, and form a mapping relationship between the documents and the semantic vectors;

[0105] S32. Perform hierarchical compression processing on the stored document semantic vectors, divide the high-dimensional semantic vectors into multiple sub-vectors, and compress each sub-vector to obtain a low-dimensional representation of each subspace;

[0106] S33. Construct an index for each sub-vector after hierarchical quantization, and define the index function as I(v):

[0107]

[0108] Among them, C i is the representation obtained by quantizing and compressing the i-th subspace, w i is the dynamic weight factor, γ i is the auxiliary adjustment coefficient, ∈ and η are small constants, Z is the normalization coefficient, k represents the total number of sub-vectors, and v is the document semantic vector;

[0109] S34. During the index construction process, adaptively adjust the dynamic weight factors of each subspace according to the retrieval performance of each sub-vector in the historical query data;

[0110] S35. While constructing the vector index, record the historical data of each query operation in real time:

[0111]

[0112] Among them, q is the query vector, sim(q, v i ) represents the similarity between the query vector and the vector v i , θ is the preset similarity threshold, R is the retrieval hit rate, S is the set of similarity score distributions, O is the outlier factor score, N is the total number of candidate documents, and H is the historical query data;

[0113] S36. Associatively store the recorded historical query data H with the constructed vector index to construct an optimized data set:

[0114]

[0115] Among them, β is the adjustment coefficient, U(v) represents the vector index optimized by combining historical data, and sim(q, v) represents the similarity calculation function between the query vector q and the document semantic vector v.

[0116] In this embodiment, the specific steps of S4 include:

[0117] S41. Receive the document D to be checked for duplication Q , and convert the document D to be checked for duplication into a query vector q through a large language model;

[0118] S42. Call the constructed vector index I(v), and calculate the weighted Euclidean distance d(q, I(v)) between the query vector q and the index I(v) of each document:

[0119]

[0120] Among them, ω j is the weight factor, q j is the j-th query vector, and I(v) jis the index of the j-th vector, and n is the dimension of the vector space;

[0121] S43. Construct a retrieval algorithm based on outlier recognition, and calculate the local outlier factor for each candidate document v:

[0122]

[0123] where N(v) is the local neighborhood, ζ is the smoothing constant, LOF(v) represents the local outlier factor of the candidate document v, u is the neighbor document, d(q, I(u)) represents the weighted Euclidean distance between the query vector q and the vector index I(u) of the neighbor document u, and exp is the exponential function;

[0124] S44. Combine the weighted Euclidean distance d(q, I(v)) with the local outlier factor LOF(v) to construct a comprehensive distance function for sorting and screening candidate documents:

[0125] d′(q, I(v)) = d(q, I(v)) · [1 + λ · ln(1 + LOF(v))];

[0126] where λ is the adjustment coefficient, d′(q, I(v)) is the comprehensive distance, and ln is the logarithmic function;

[0127] S45. Combine the recorded historical query data H to optimize the comprehensive distance of the candidate document v:

[0128]

[0129] where β is the historical data adjustment coefficient, d″(q, I(v)) represents the optimized distance, H(v) represents the set of historical query data related to the document v, q′ is the historical query vector, θ is the preset similarity threshold, and ||||2 represents the Euclidean norm;

[0130] S46. Sort all the documents in the vector database in ascending order according to the optimized distance d″(q, I(v)), and screen out the candidate document set S with high preliminary similarity C :

[0131] S C = {v ∈ D | d″(q, I(v)) ≤ τ};

[0132] where τ is the preset distance threshold, and D represents all the stored document semantic vectors in the vector database.

[0133] In this embodiment, the S5 specifically includes:

[0134] S51. For the semantic vector v of each candidate document in the obtained candidate document set with high preliminary similarity, v ∈ R nThe query vector q ∈ R generated from the document to be checked for duplication n Calculate the weighted cosine similarity sim wcos (q, v):

[0135]

[0136] where ω j is the weight factor, q j is the component of the semantic vector q representing the query document on the j-th dimension, and v j represents the component value of the candidate document semantic vector v on the j-th dimension. n represents the dimension of the semantic vector;

[0137] S52. Calculate the adaptive Mahalanobis distance between the semantic vector v and the query vector q

[0138]

[0139] where Σ is the covariance matrix, β o is the adjustment coefficient, and diag(ω) represents the diagonal matrix;

[0140] S53. Combine the weighted cosine similarity and the adaptive Mahalanobis distance to construct the comprehensive similarity score of the candidate document:

[0141]

[0142] where α o is the weight coefficient, and S(v) is the comprehensive similarity score of the candidate document v;

[0143] S54. Use the local neighborhood N(v) of each document v in the candidate document set to calculate the refined local density factor RLD(v):

[0144]

[0145] where LOF(v) represents the local outlier factor of the candidate document v, and I(u) is the vector index of the candidate document u;

[0146] S55. Screen each document v in the candidate document set according to the comprehensive similarity score S(v) and the refined local density factor RLD(v), and select those that satisfy:

[0147] S(v) ≥ τ S and RLD(v) ≤ τ RLD ;

[0148] The high-similarity documents that meet the criteria are used as the final duplicate-checking results. Among them, τ S is the preset comprehensive similarity threshold, and τ RLD is the preset local density factor threshold.

[0149] In this embodiment, S6 specifically includes:

[0150] S61. Generate a duplicate check report based on the screened highly similar documents, including the basic information of the document to be checked, the list of similar documents, the similarity score, the detected repeated content segments, and the semantic analysis results generated by the system;

[0151] S62. In the duplicate check report, display the list of similar documents through a graphical interface, sorted in descending order of similarity, and provide key attribute information for each similar document, including the title, source, similarity score, the number of repeated paragraphs, and the position index of the similar segments;

[0152] S63. Automatically mark the repeated content in the document to be checked, highlight the repeated content of different degrees with different colors or formats, and provide corresponding annotation information, such as the similar document number, the matching position, and the similarity score;

[0153] S64. Combine the semantic analysis in the duplicate check process to display the detailed evaluation of the text similarity by the duplicate check system, including semantic reconstruction analysis, syntactic structure change detection, and cross-paragraph semantic comparison;

[0154] S65. Provide user interaction functions on the duplicate check report interface, allowing users to manually adjust the similarity threshold, filter specific types of similar documents, mark misjudged or correctly matched content, and provide feedback options;

[0155] S66. Record the operation behaviors of users in the duplicate check report and store them in the historical query data.

[0156] In this embodiment, S7 specifically includes:

[0157] S71. Receive the feedback data of users on the duplicate check results, and record the operation behaviors of users in the duplicate check report interface to form a feedback data set;

[0158] S72. Integrate the collected user feedback data with the historical query data to construct a comprehensive feedback set;

[0159] S73. Calculate the error amount between the system prediction output and the user feedback based on the comprehensive feedback set to form a feedback error index;

[0160] S74. Use self-supervised learning technology and dynamic feedback error to fine-tune the parameters of the large language model:

[0161]

[0162] Among them, θ old represents the parameters of the large language model before update, θ newDenote the updated large language model parameters as θ, η as the adaptive learning rate, and m t as the first-order momentum term, and v t as the second-order momentum term, ∈ as a small constant to prevent division by zero, ρ as the feedback error adjustment coefficient, and ΔF as the average feedback error;

[0163] S75. Use the updated large language model parameters θ new to regenerate semantic vectors for the document collection and obtain the updated document representation;

[0164] S76. According to the updated semantic vectors, adopt a dynamic index update strategy to update the indexes in the vector database:

[0165]

[0166] where I(v) is the original document vector index, v new is the document semantic vector generated after update, and f update (v new ) is the new index value calculated based on the updated semantic vectors, is the adjustment coefficient, and I(v new ) is the vector index after dynamic update.

[0167] Reference Figure 2 , an intelligent document duplicate checking system based on a vector database and a large language model, includes the following modules:

[0168] A data collection and preprocessing module, used to extract, parse documents in multiple formats and preprocess the documents;

[0169] A large language model semantic encoding module, using the large language model to convert the preprocessed documents into high-dimensional semantic vectors to capture deep semantic features;

[0170] A vector database and index construction module, storing document semantic vectors and constructing a hierarchical index to improve the retrieval efficiency through quantization and compression technologies;

[0171] A preliminary candidate document retrieval module, adopting a retrieval algorithm based on outlier recognition and optimizing with historical query data to quickly screen out candidate documents with high semantic similarity to the document to be checked for duplicates;

[0172] A fine-grained similarity calculation module, combining weighted cosine similarity, adaptive Mahalanobis distance, and local density factor to perform fine similarity calculation and screening on candidate documents;

[0173] A duplicate checking report generation and display module, generating a duplicate checking report and displaying similar documents, similarity scores, highlighted duplicate content, and semantic analysis results through a graphical interface;

[0174] The dynamic feedback and model optimization module receives user feedback and combines self-supervised learning techniques to perform real-time fine-tuning on the parameters and vector indices of the large language model.

[0175] Example 1:

[0176] To verify the feasibility of the present invention in implementation, the present invention is applied to a scientific research institute of a certain university. This scientific research institute has hundreds of thousands of academic papers and various research reports and has long faced problems of academic misconduct and repeated citation of literature. When dealing with complex semantic matching tasks such as synonym replacement, paragraph reconstruction, and semantic reconstruction, traditional plagiarism detection systems often have problems of missed judgment and misjudgment, resulting in inaccurate and inefficient plagiarism detection results. To solve this problem, the system proposed by the present invention uses a large language model to generate high-dimensional semantic vectors, performs semantic encoding on documents, and realizes the efficient storage and index construction of massive vector data through a vector database. Then, combined with a retrieval algorithm based on outlier recognition, cosine similarity, adaptive Mahalanobis distance, and a dynamic feedback mechanism, it realizes the fine calculation of multi-dimensional semantic similarity and the efficient screening of candidate documents, thus greatly improving the accuracy and recall rate of plagiarism detection.

[0177] In practical applications, the system first uniformly preprocesses all the documents in the institute and converts various types of documents into standardized texts. Subsequently, it uses a large language model to generate high-dimensional semantic vectors for the preprocessed texts and stores these vectors in a vector database through hierarchical compression and index construction techniques. On this basis, the system generates a query vector by inputting the document to be detected for plagiarism, quickly preliminarily screens out a set of candidate documents with relatively high similarity from the massive literature using a retrieval algorithm based on outlier recognition, then performs refined calculations on the candidate documents using weighted cosine similarity and adaptive Mahalanobis distance, and further screens the candidate documents in combination with the local density factor. Finally, it outputs the document with the highest semantic similarity as the plagiarism detection result.

[0178] For example, in October 2024, when a scientific research institute of a certain university detected plagiarism for 5,000 newly submitted academic papers, the traditional plagiarism detection system took an average of about 15 minutes per paper, and the plagiarism detection accuracy was only 78%, with many problems of misjudgment and missed judgment. After using the intelligent document plagiarism detection system of the present invention, through the deep combination of the large language model and vector database technology, the entire plagiarism detection process achieved fully automated processing. The average time for the system to detect plagiarism per paper decreased to about 3 minutes, the plagiarism detection accuracy increased to over 92%, and the misjudgment rate and missed judgment rate were both significantly reduced. In addition, through the dynamic feedback and model optimization mechanism, during the continuous operation within a week, the system automatically fine-tuned the model parameters and index strategies according to user feedback, further increasing the plagiarism detection accuracy to 94% and ensuring efficient response under large data volumes.

[0179] During the duplicate checking process, the system not only detects the surface similarity of documents, but also delves into the semantic structure within the documents, and can accurately identify complex situations such as synonymous substitution, word order adjustment, and paragraph reconstruction. After using the system, researchers feedback that the system can accurately find the duplicate citations and semantic reconstructions in the literature, effectively assisting the originality review work of academic papers. Through comparative analysis, this system is superior to traditional duplicate checking solutions in terms of retrieval time, duplicate checking accuracy, and system stability, saving a large amount of duplicate checking labor and time costs for the institute, and at the same time providing a strong guarantee for the prevention of academic misconduct.

[0180] Table 1 Data comparison table of the performance of the document duplicate checking system

[0181] Index Traditional Plagiarism Detection System Plagiarism Detection System of the Present Invention Total Plagiarism Detection Time 1250 hours 250 hours Average Plagiarism Detection Time 15 minutes per article 3 minutes per article Plagiarism Detection Accuracy 78% 94% False Judgment Rate 12% 4% Omission Judgment Rate 10% 2% User Feedback Satisfaction 70% 92%

[0182] In this experiment, the data table shows that the intelligent document duplicate checking system of the present invention has significant advantages in processing large-scale documents. When the traditional duplicate checking system processes about 5000 papers, the total duplicate checking time is as high as 1250 hours, with an average of 15 minutes per paper, while the system of the present invention only needs 250 hours, with an average of 3 minutes per paper, and the duplicate checking efficiency is improved by more than 5 times. This improvement greatly shortens the document review time and makes the duplicate checking process more efficient.

[0183] At the same time, in terms of duplicate checking accuracy, the traditional system is only 78%, while the system of the present invention can reach 94%, with an accuracy improvement of 16 percentage points. The system generates high-dimensional semantic vectors through a large language model, and combines hierarchical indexing, outlier recognition, and multiple similarity calculations to effectively capture deep semantic features, thereby greatly reducing misjudgment and missed judgment. The misjudgment rate and missed judgment rate are reduced from 12% and 10% to 4% and 2% respectively. This refined semantic matching ability is particularly prominent when dealing with complex situations such as synonymous substitution and word order adjustment.

[0184] In addition, the user feedback satisfaction has increased from 70% of the traditional system to 92% of the system of the present invention, indicating that users have obtained obvious improvements in the display of duplicate checking reports, highlighting of duplicate content, and interactive experience. The system not only provides intuitive duplicate checking results, but also supports users to dynamically adjust the duplicate checking results, further optimizing the duplicate checking effect.

[0185] In summary, the intelligent document duplicate checking system of the present invention shows excellent advantages in terms of duplicate checking efficiency, accuracy, and user experience through the organic combination of a large language model, a vector database, and a dynamic feedback mechanism, providing an efficient, accurate, and adaptive solution for document duplicate checking.

[0186] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, making equivalent substitutions or changes should be covered within the protection scope of the present invention.

Claims

1. An intelligent document duplication detection method based on a vector database and a large language model, characterized in that, It includes the following steps: S1. Collect document data in multiple formats and preprocess the document data; S2. Semantically encode the preprocessed document data through a large language model to generate document semantic vectors, and pre-train the large language model; S3. Store the document semantic vectors in a vector database, hierarchically compress the document semantic vectors, build a vector index and record historical query data; S4. Receive the document to be checked for duplication, generate a query vector through a large language model, call the built vector index, perform approximate nearest neighbor search based on an outlier recognition-based retrieval algorithm, filter out a candidate document set with a high initial similarity, and optimize the outlier recognition-based retrieval algorithm in combination with historical query data; S5. For each document vector in the candidate document set with a high initial similarity, calculate the semantic similarity with the document vector to be checked for duplication, combine the cosine similarity and the Mahalanobis distance to obtain a similarity score, and filter out the documents with a high similarity; S6. Generate a duplication check report based on the filtered documents with a high similarity, display the list of similar documents, similarity scores, highlighted duplicate content, and semantic analysis results through a graphical interface, and record the user's operation behavior in the duplication check report in the historical query data at the same time; S7. Receive the user's feedback on the duplication check result, dynamically adjust the semantic representation ability of the large language model through the fine-tuning mechanism of the large language model in combination with self-supervised learning technology, and update the document semantic vectors and vector index in the vector database.

2. The intelligent document duplicate checking method based on a vector database and a large language model according to claim 1, wherein The specific steps of S3 include: S31. Bind the document semantic vectors generated by the large language model with the corresponding document identifiers and store them in the vector database to form a mapping relationship between the documents and the semantic vectors; S32. Perform hierarchical compression processing on the stored document semantic vectors, divide the high-dimensional semantic vectors into multiple sub-vectors, and compress each sub-vector to obtain low-dimensional representations of each subspace; S33. Build an index for each sub-vector after hierarchical quantization, and define the index function as I(v): Among them, C i is the representation obtained by quantizing and compressing the i-th subspace, w i is the dynamic weight factor, γ i is the auxiliary adjustment coefficient, ∈ and η are small constants, Z is the normalization coefficient, k represents the total number of sub-vectors, and v is the document semantic vector; S34. During the index construction process, adaptively adjust the dynamic weight factors of each subspace according to the retrieval performance of each sub-vector in the historical query data; S35. While building the vector index, record the historical data of each query operation in real time: Among them, q is the query vector, sim(q, v i ) represents the similarity between the query vector and the vector v i , θ is the preset similarity threshold, R is the retrieval hit rate, S is the set of similarity score distributions, O is the outlier factor score, N is the total number of candidate documents, and H is the historical query data; S36. Associatively store the recorded historical query data H with the built vector index to build an optimized data set: Among them, β is an adjustment coefficient, U(v) represents the vector index optimized by combining historical data, and sim(q, v) represents the similarity calculation function between the query vector q and the document semantic vector v.

3. The intelligent document duplicate checking method based on a vector database and a large language model according to claim 1, characterized in that, The specific steps of S4 include: S41. Receive the document D to be checked for duplication Q , and convert the document D to be checked for duplication into a query vector q through a large language model; S42. Call the built vector index I(v) and calculate the weighted Euclidean distance d(q, I(v)) between the query vector q and the index I(v) of each document; Among them, ω j is the weight factor, q j is the j-th query vector, I(v) j is the j-th vector index, and n is the dimension of the vector space; S43. Build an outlier recognition-based retrieval algorithm and calculate the local outlier factor for each candidate document v; Among them, N(v) is the local neighborhood, ζ is the smoothing constant, LOF(v) is the local outlier factor representing the candidate document v, u is the neighbor document, d(q, I(u)) is the weighted Euclidean distance between the query vector q and the vector index I(u) of the neighbor document u, and exp is the exponential function; S44. Combine the weighted Euclidean distance d(q, I(v)) with the local outlier factor LOF(v) to construct a comprehensive distance function for sorting and screening candidate documents: d′(q, I(v)) = d(q, I(v))·[1 + λ·ln(1 + LOF(v))]; Among them, λ is the adjustment coefficient, d′(q, I(v)) is the comprehensive distance, and ln is the logarithmic function; S45. Optimize the comprehensive distance of the candidate document v in combination with the recorded historical query data H: where β is the historical data adjustment coefficient, d ″ (q, I(v)) represents the optimized distance, H(v) represents the set of historical query data related to document v, q′ is the historical query vector, θ is the preset similarity threshold, and ||||2 represents the Euclidean norm; S46. Sort all the documents in the vector database in ascending order according to the optimized distance d ″ (q, I(v)), and filter out the candidate document set S with high preliminary similarity C : S C = {v ∈ D | d ″ (q, I(v)) ≤ τ}; Among them, τ is the preset distance threshold, and D represents all stored document semantic vectors in the vector database.

4. The intelligent document duplicate checking method based on a vector database and a large language model according to claim 1, wherein, The specific steps of S5 are as follows: S51. For each candidate document in the obtained candidate document set with a high preliminary similarity, the semantic vector v ∈ R n and the query vector q ∈ R generated by the document to be checked for duplication n Calculate the weighted cosine similarity sim wcos (q, v): Among them, ω j is the weight factor, q j is the component of the semantic vector q representing the query document in the j-th dimension, and v j represents the component value of the candidate document semantic vector v in the j-th dimension, and n represents the dimension of the semantic vector; S52. Calculate the adaptive Mahalanobis distance between the semantic vector v and the query vector q where Σ is the covariance matrix, and β o is the adjustment coefficient, and diag(ω) represents a diagonal matrix; S53. Combine the weighted cosine similarity and the adaptive Mahalanobis distance to construct a comprehensive similarity score for candidate documents: Among them, α o is the weight coefficient, and S(v) is the comprehensive similarity score of the candidate document v; S54. Use the local neighborhood N(v) of each document in the candidate document set to calculate the refined local density factor RLD(v): Among them, LOF(v) represents the local outlier factor of the candidate document v, and I(u) is the vector index of the candidate document u; S55. Screen each document v in the candidate document set according to the comprehensive similarity score S(v) and the refined local density factor RLD(v), and select those that meet: S(v) ≥ τ S and RLD(v) ≤ τ RLD ; The highly similar documents are used as the final duplicate check result, where τ S is a preset comprehensive similarity threshold, and τ RLD is a preset local density factor threshold.

5. The intelligent document duplicate checking method based on a vector database and a large language model according to claim 1, characterized in that The specific steps of S6 are as follows: S61. Generate a duplicate check report based on the selected highly similar documents, including the basic information of the document to be checked, the list of similar documents, the similarity score, the detected duplicate content segments, and the semantic analysis results generated by the system; S62. In the duplicate check report, display the list of similar documents through a graphical interface, sorted in descending order of similarity, and provide key attribute information for each similar document, including title, source, similarity score, number of duplicate paragraphs, and similarity segment position index; S63. Automatically mark the duplicate content in the document to be checked, highlight different degrees of duplicate content with different colors or formats, and provide corresponding annotation information, such as similar document number, matching position, and similarity score; S64. Combine the semantic analysis during the duplicate check process to display the detailed evaluation of the text similarity by the duplicate check system, including semantic reconstruction analysis, syntactic structure change detection, and cross-paragraph semantic comparison; S65. Provide user interaction functions on the duplicate check report interface, allowing users to manually adjust the similarity threshold, filter specific types of similar documents, mark misjudged or correctly matched content, and provide feedback options; S66. Record the operation behaviors of users in the duplicate check report and store them in the historical query data.

6. The intelligent document duplicate checking method based on a vector database and a large language model according to claim 1, wherein The specific steps of S7 are as follows: S71. Receive the feedback data of users on the duplicate check results, and record the operation behaviors of users in the duplicate check report interface to form a feedback data set; S72. Integrate the collected user feedback data with the historical query data to construct a comprehensive feedback set; S73. Calculate the error amount between the system prediction output and the user feedback based on the comprehensive feedback set to form a feedback error index; S74. Use self-supervised learning technology and dynamic feedback error to fine-tune the parameters of the large language model: Among them, θ old represents the parameters of the large language model before update, and θ new represents the parameters of the large language model after update. η is the adaptive learning rate, and m t is the first-order momentum term, v t is the second-order momentum term, ∈ is a small constant to prevent division by zero, ρ is the feedback error adjustment coefficient, and ΔF is the average feedback error; S75. Using the updated large language model parameters θ new Regenerate semantic vectors for the document set to obtain updated document representations; S76. According to the updated semantic vectors, update the indexes in the vector database using a dynamic index update strategy: Among them, I(v) is the original document vector index, v new is the document semantic vector generated after update, f update (v new ) is the new index value calculated based on the updated semantic vector, is the adjustment coefficient, I(v new ) is the vector index after dynamic update.

7. An intelligent document duplicate checking system based on a vector database and a large language model, the intelligent document duplicate checking method based on a vector database and a large language model according to any one of claims 1 to 6, characterized in that, It includes the following modules: Data collection and preprocessing module, which is used to extract, parse documents in multiple formats and preprocess the documents; Large language model semantic encoding module, which uses the large language model to convert the preprocessed documents into high-dimensional semantic vectors to capture deep semantic features; Vector database and index construction module, which stores the document semantic vectors and constructs a hierarchical index to improve the retrieval efficiency through quantization and compression technologies; Initial candidate document retrieval module, which uses a retrieval algorithm based on outlier recognition and historical query data optimization to quickly screen out candidate documents with high semantic similarity to the document to be checked for duplication; Fine-grained similarity calculation module, which combines weighted cosine similarity, adaptive Mahalanobis distance and local density factor to perform fine similarity calculation and screening on the candidate documents; Duplication check report generation and display module, which generates a duplication check report and displays similar documents, similarity scores, highlighted duplicate content and semantic analysis results through a graphical interface; Dynamic feedback and model optimization module, which receives user feedback and combines self-supervised learning technology to perform real-time fine-tuning on the parameters of the large language model and vector indexes.

Citation Information

Cited By

  • Semantic focusing test question duplicate checking method based on large language model

    CN121031565A

  • Data mining system and method for mail data

    CN121070985A

  • Pre-patrol report generation method and system based on vector model

    CN121235086A

  • A method and system for generating pre-trip reports based on a vector model

    CN121235086B

  • Automatic duplicate checking and rewriting method and device for document and program product

    CN121328511A