File open review auxiliary judgment method fusing synonym and large language model

Through the jieba word segmentation and the improved Word2Vec model, and the archive content is understood in combination with the large language model, the problem of low accuracy of open archive review in the existing technology is solved, and more efficient review results are achieved.

CN120387444APending Publication Date: 2025-07-29NINGBO INST OF TECH ZHEJIANG UNIV ZHEJIANG +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510512182.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing open archive review method relies on manual and simple sensitive words to match, resulting in low accuracy, especially when dealing with rich synonyms and language expressions, it is difficult to accurately judge information involving state secrets, trade secrets and personal privacy.

Method used

The jieba word segmentation tool and the improved Word2Vec model are used to expand the sensitive thesaurus, and the large language model is used to understand the archive content. Through multi-level judgments of initial sensitive words, extended sensitive words and large language model, open archive review is assisted.

Benefits of technology

It significantly improves the accuracy of open archive review, reduces the workload of manual review, and improves the review efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The invention relates to the technical field of archive opening and review, in particular to an archive opening and review auxiliary judgment method fusing synonyms and a large language model, which comprises the following steps: S1, an archive opening and review worker conceives an initial sensitive word; s2, collecting all archive files in an archive and digitalizing the archive files into text files by using OCR (Optical Character Recognition); s3, performing word segmentation processing on the text file obtained in the step S2 by using a jieba word segmentation tool, training by adopting a Word2Vec model, inputting the initial sensitive words in the step S1 into the trained Word2Vec model, and finding out vocabularies of which the similarity is greater than 0.7 as extended sensitive words; s4, deploying a large language model locally; s5, judging the archive which needs to be subjected to open auditing; and S6, manually auditing the final auditing archive. The method can greatly improve the accuracy of archive opening audit judgment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of open review of archives, and in particular to an auxiliary judgment method for open review of archives that integrates near-synonyms and large language models. Background Art

[0002] Currently, the open review of archives mainly relies on manual work. Since the review of archive content requires professional knowledge and an accurate understanding of policies and regulations, staff need to check the archive content word by word, line by line, and page by page to determine whether it involves information that is not suitable for public disclosure, such as state secrets, business secrets, and personal privacy. Even with computer-aided review, only simple sensitive word matching is used. By establishing a sensitive word library, the archive text is retrieved, and the archives containing sensitive words are screened out for further manual review. However, the accuracy of relying solely on sensitive word matching for auxiliary judgment is relatively low, mainly reflected in:

[0003] Firstly, the sensitive words are mainly determined by the review staff. However, due to differences in the knowledge reserves, business understanding, and awareness of various sensitive information among the staff, it is difficult for the constructed sensitive word library to cover all aspects comprehensively. For example, in terms of personal privacy, in addition to the basic word "privacy", closely related near-synonyms such as "private information", "personal confidential materials", and "details of personal information" are easily overlooked and not included in the sensitive word library due to the negligence of the staff.

[0004] Secondly, the language expression is extremely rich. Simply comparing text words mechanically and ignoring the in-depth understanding and judgment of the archive content will greatly reduce the judgment accuracy. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: to provide an auxiliary judgment method for open review of archives that integrates near-synonyms and large language models, which can greatly improve the accuracy of open review judgment of archives.

[0006] The technical solution adopted by the present invention is: an auxiliary judgment method for open review of archives that integrates near-synonyms and large language models, which includes the following steps:

[0007] S1. The staff for open review of archives conceive initial sensitive words according to the content of the "Interim Provisions on the Decryption and Classification of the Scope of Controlled Use of Archives in National Archives at All Levels".

[0008] S2. Collect all archive files in the archive and digitize the archive files into text files using OCR.

[0009] S3. Use the jieba word segmentation tool to segment the text file obtained in step S2, then train it using the Word2Vec model. After that, input the initial sensitive words in step S1 into the trained Word2Vec model, and find the words with a similarity greater than 0.7 as extended sensitive words;

[0010] S4. Deploy a large language model locally;

[0011] S5. Judge the files that need to be open for review, which specifically includes the following steps:

[0012] S51. Use the initial sensitive words in step S1 to perform text matching on all files that need to be open for review. If the initial sensitive words in step S1 exist in the file, the file is judged as "restricted access". If the initial sensitive words in step S1 do not exist in the file, the file is classified into the preliminary review files, and then go to step S52;

[0013] S52. Use the extended sensitive words in step S3 to perform text matching on the preliminary review files. If the extended sensitive words in step S3 exist in the file, the file is judged as "restricted access". If the extended sensitive words do not exist in the file, the file is classified into the secondary review files, and then go to step S53.

[0014] S53. Save the content in the "Interim Provisions on the Decryption and Classification of the Scope of Controlled Use of Archives in National Archives at All Levels" as a document as the knowledge base file of the large language model. Input the text content of the secondary review files as the framework chat content of the large language model, and add "Judge whether this file can be opened" at the end of the chat content, and then obtain the final review files that can be opened;

[0015] S6. Manually review the final review files.

[0016] Preferably, the improved Word2Vec model is used in step S3. Step S3 specifically includes the following steps:

[0017] S31. Use the jieba word segmentation tool to segment the text file obtained in step S2;

[0018] S32. Use the word segmentation in step S31 for one-hot vector encoding;

[0019] S33. Let x t be the one-hot vector encoding corresponding to a word segmentation in the text. Select the first n words and the last n words of this word segmentation, and calculate the one-hot vector encoding considering the position. The calculation formula is as follows:

[0020]

[0021] where \(i = t - n,\cdots,t - 1,t + 1,\cdots,t + n\), \(x\) i is the one - hot vector encoding of the word segmentation at position \(i\), and \(x'\) i is the one - hot vector encoding of the word segmentation at position \(i\) after considering the position factor;

[0022] Step S34: Calculate the vector \(h\) through the weight matrix \(W\) t :

[0023]

[0024] where \(W\) is a matrix with dimensions \(V\times D\), \(V\) is the number of different words obtained by word segmentation in Step S31, and \(D\) is the dimension of the vector \(h\) t ;

[0025] Step S35: Calculate the vector \(y\) through the weight matrix \(A\) t :

[0026] \(y\) t \(= Ah\) t

[0027] where \(A\) is a matrix with dimensions \(D\times V\);

[0028] Step S36: Calculate the central word vector \(e\) with the attention mechanism introduced, and the calculation steps are as follows: t

[0029] Step S361: Calculate the attention score \(e\) for each \(x'\) i : i

[0030] \(e\) i \(=(x\) i ') T \(y\) t

[0031] where \(i = t - n,\cdots,t - 1,t + 1,\cdots,t + n\);

[0032] Step S362: Calculate the attention weight \(\alpha\) i :

[0033]

[0034] where \(j = t - n,\cdots,t - 1,t + 1,\cdots,t + n\);

[0035] Step S363: Calculate the central word vector \(z\) with the attention mechanism introduced t :

[0036]

[0037] Step S37: Train the weight matrix W and weight matrix A in the model using the following loss function L:

[0038]

[0039] Obtain an improved Word2Vec model;

[0040] Step S38: Input the initial sensitive words in Step S1 into the improved Word2Vec model, and find the words with a similarity greater than 0.7 as the extended sensitive words.

[0041] Preferably, the large language model in Step S4 is Tongyi Qianwen large language model.

[0042] Compared with the prior art by adopting the above method, the present invention has the following advantages: for the sensitive word library constructed by the staff, the improved Word2Vec model is used to expand the sensitive word library, greatly improving the coverage of the sensitive word library; at the same time, the large language model is used, and the judgment rule is used as a prompt word for the large language model to learn, and the large language model is used to judge the understanding of the file content, which can greatly improve the accuracy of the file opening review judgment. Detailed implementation manners

[0043] The embodiments of the present invention are described in detail below. The examples of the embodiments show that the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by reference are exemplary and are intended to explain the present invention and should not be construed as a limitation of the present invention.

[0044] Embodiment 1:

[0045] A method for assisting in judging file opening review by integrating synonyms and large language models includes the following steps:

[0046] S1: The staff for file opening review conceive initial sensitive words according to the content of the "Interim Provisions on the Decryption and Division of the Scope of Controlled Use of the Archives in the National Archives at All Levels"; the "Interim Provisions on the Decryption and Division of the Scope of Controlled Use of the Archives in the National Archives at All Levels" is a standard regulation and can be searched in the departmental regulation library. For example, regarding citizens' personal privacy, such as educational information, the staff can generally think of education, diploma; materials involving requests and reports are not made public, and the staff can generally think of request documents, report materials, request reports;

[0047] S2: Collect all the file documents in the archive and digitize the file documents into text files using OCR;

[0048] S3. Use the Jieba word segmentation tool to segment the text file obtained in step S2, and then use the Word2Vec model for training. After that, input the initial sensitive words in step S1 into the trained Word2Vec model to find words with a similarity greater than 0.7 as extended sensitive words; using the ordinary Word2Vec model, extended sensitive words can be obtained. For example, if you input academic qualifications and diplomas, you can get academic qualifications, diplomas, academic ability, student status, homework, and cultural level; if you input request documents, reporting materials, and request reports, you can get request documents, request documents, reporting materials, situation reports, request reports, and request applications;

[0049] S4. Deploy the Tongyi Qianwen large language model locally. Since this is a large language model from the prior art, there are detailed guides on how to deploy it locally online, so we will not elaborate on it here. This application only uses this large language model for judgment, and the specific judgment method is also a function of the Tongyi Qianwen large language model itself, so we will not elaborate on it here.

[0050] S5. Determine the files that need to be open for review, including the following steps:

[0051] S51. Use the initial sensitive words in step S1 to perform text matching on all files that need to be open for review. If the initial sensitive words in step S1 are present in the file, the file is judged as "restricted open". If the initial sensitive words in step S1 are not present in the file, the file is classified as a preliminary review file, and then the process proceeds to step S52; this step is to delete files containing the initial sensitive words.

[0052] S52, use the extended sensitive words in step S3 to perform text matching on the preliminary review file. If the extended sensitive words in step S3 exist in the file, the file is judged as "restricted opening". If the extended sensitive words do not exist in the file, the file is classified as a secondary review file, and then enter step S53, which is to delete the file with the extended sensitive words.

[0053] S53. Save the contents of the "Interim Provisions on Declassification and Controlled Use of Archives in National Archives at All Levels" into a document as the knowledge base file for the large language model. Input the text content of the secondary review archive as the framework chat content of the large language model, and add "For this archive, determine whether it can be opened" to the end of the chat content to obtain the final review archive that can be opened. In this step, the AI further deletes some archives that are restricted from opening based on semantics.

[0054] S6. The final audit files will be manually reviewed. This will result in fewer files being screened than before, and manual review will be much more efficient.

[0055] Example 2:

[0056] The difference from Example 1 is that the Word2Vec model used in Example 2 is an improved one. Specifically, step S3 includes the following steps:

[0057] S31. Use the jieba word segmentation tool to segment the text file obtained in step S2;

[0058] S32. Use the segmentation in step S31 for one-hot vector encoding;

[0059] S33. Let x t be the one-hot vector encoding corresponding to a word segment in the text. Select the first n and the last n words of this word segment, and calculate the one-hot vector encoding considering the position. The calculation formula is as follows:

[0060]

[0061] where i = t - n, …, t - 1, t + 1, ……, t + n, x i is the one-hot vector encoding of the word segment at position i, and x’ i is the one-hot vector encoding of the word segment at position i considering the position factor;

[0062] Step S34. Calculate the vector h t through the weight matrix W

[0063]

[0064] where W is a matrix with dimensions V×D, V is the number of different words obtained by word segmentation in step S31, and D is the dimension of the vector h t ;

[0065] Step S35. Calculate the vector y t through the weight matrix A

[0066] y t = Ah t

[0067] where A is a matrix with dimensions D×V;

[0068] Step S36. Calculate the central word vector e t introducing the attention mechanism. The calculation steps are as follows:

[0069] Step S361. Calculate the attention score e i for each x’ i :

[0070] e i =(x i ') T yt

[0071] where i = tn,…,t-1,t+1,…,t+n;

[0072] Step S362: Calculate attention weight α i :

[0073]

[0074] where j = tn,…,t-1,t+1,…,t+n;

[0075] Step S363: Calculate the center word vector z introduced by the attention mechanism t :

[0076]

[0077] Step S37: Use the following loss function L to train the weight matrix W and weight matrix A in the model:

[0078]

[0079] Get the improved Word2Vec model;

[0080] Step S38: Input the initial sensitive words in step S1 into the improved Word2Vec model to find words with a similarity greater than 0.7 as extended sensitive words.

[0081] The above-mentioned improved Word2Vec model mainly adds position encoding and attention mechanism on the basis of the original Word2Vec model, so that the overall result is more accurate. For example, if academic qualifications and diplomas are input, the improved Word2Vec model will output academic qualifications, diplomas, academic ability, student status, academic titles, courses, academic level, educational resume, and cultural level; if request documents, reporting materials, and request reports are input, the improved Word2Vec model will output request documents, request documents, request documents, request letters, reporting materials, situation reports, work reports, performance reports, request reports, request applications, report requests, and petition reports.

[0082] That is, the improved Word2Vec model can output more and more accurate extended sensitive words, so that the overall accuracy of archive open review is higher.

[0083] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

[0084] For those skilled in the art, after reading the above description, various changes and modifications will undoubtedly be obvious. Therefore, the appended claims should be regarded as covering all changes and modifications that fall within the true intent and scope of the present invention. Any and all equivalent ranges and contents within the scope of the claims should be considered to still fall within the intent and scope of the present invention.

Claims

1. An auxiliary judgment method for the open review of archives that integrates synonyms and large language models, characterized in that It includes the following steps: S1. The staff responsible for the review of the opening of archives conceive the initial sensitive words according to the content of the "Interim Provisions on the Decryption and Classification of the Scope of Controlled Use of Archives in National Archives at All Levels". S2. Collect all the archival documents in the archive and digitize the archival documents into text files using OCR. S3. Use the jieba word segmentation tool to segment the text files obtained in step S2, then train using the Word2Vec model, and then input the initial sensitive words in step S1 into the trained Word2Vec model to find the words with a similarity greater than 0.7 as the extended sensitive words. S4. Deploy a large language model locally. S5. Judge the archives that need to be reviewed for opening, which specifically includes the following steps: S51. Use the initial sensitive words in step S1 to perform text matching on all the archives that need to be reviewed for opening. If the initial sensitive words in step S1 exist in the archives, the archives are judged as "restricted opening". If the initial sensitive words in step S1 do not exist in the archives, the archives are classified into the preliminary review archives, and then go to step S52; S52. Use the extended sensitive words in step S3 to perform text matching on the preliminary review archives. If the extended sensitive words in step S3 exist in the archives, the archives are judged as "restricted opening". If the extended sensitive words do not exist in the archives, the archives are classified into the secondary review archives, and then go to step S53; S53. Save the content in the "Interim Provisions on the Decryption and Classification of the Scope of Controlled Use of Archives in National Archives at All Levels" as a knowledge base file for the large language model, input the text content of the secondary review archives as the framework chat content of the large language model, and add "Judge whether this archive can be opened" at the end of the chat content to obtain the final reviewed archives that can be opened; S6. Conduct a manual review of the final reviewed archives.

2. The method for assisting in judging the opening review of archives by integrating synonyms and large language models according to claim 1, wherein: The improved Word2Vec model is adopted in step S3, and step S3 specifically includes the following steps: S31. Use the jieba word segmentation tool to segment the text files obtained in step S2; S32. Perform one-hot vector encoding using the word segmentation in step S31; S33. Let x t be the one-hot vector encoding corresponding to a word segment in the text. Select the first n and the last n words of this word segment, and calculate the position-aware one-hot vector encoding. The calculation formula is as follows: where \(i = t - n,\ldots,t - 1,t + 1,\ldots,t + n\), \(x\) i is the one-hot vector encoding of the word segmentation at position \(i\), and \(x'\) i is the one-hot vector encoding of the word segmentation at position \(i\) after considering the position factor; Step S34: Calculate vector h using weight matrix W t : Where W is a matrix with dimensions V×D, V is the number of different words obtained by word segmentation in step S31, and D is the dimension of the vector h t ; Step S35: Calculate the vector y using the weight matrix A t : y t = Ah t where A is a matrix with dimensions D×V; Step S36: Calculate the central word vector e incorporating the attention mechanism t , and the calculation steps are as follows: Step S361. For each x' i Calculate the attention score e i : e i =(x i ) T y t where i = t - n, …, t - 1, t + 1, ……, t + n; Step S362, calculate the attention weight α i : where j = t - n, …, t - 1, t + 1, ……, t + n; Step S363: Calculate the central word vector z incorporating the attention mechanism t : Step S37. Use the following loss function L to train the weight matrix W and the weight matrix A in the model: Obtain the improved Word2Vec model; Step S38. Input the initial sensitive words in step S1 into the improved Word2Vec model to find the words with a similarity greater than 0.7 as the extended sensitive words.

3. The method for assisting in judging the opening review of archives by integrating synonyms and large language models according to claim 1, wherein: The large language model in step S4 is Tongyi Qianwen large language model.