Document checking method and device, electronic equipment and storage medium
By calculating the hash value of the document and filtering similar reference documents, the inspectors are assisted in document inspection, solving the problems of low efficiency and low accuracy of manual review, and achieving more efficient and accurate document inspection.
Patent Information
- Application Number
- CN202510215993.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-17
AI Technical Summary
In the prior art, document inspection relies on manual review, resulting in low efficiency and low accuracy, especially when technical knowledge is not fully covered.
By calculating the hash value of the document to be inspected based on the reference vocabulary, filtering out the reference documents similar to the document to be inspected, and extracting the feature vectors of the sentences, and determining the reference text based on the similarity to assist the inspector in the inspection.
It improves the efficiency and accuracy of document inspections, reduces the requirements for inspectors to fully master technical knowledge, and enhances the automation and accuracy of the review process.
Smart Images

Figure CN120162819A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a document inspection method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of information technology, electronic documents have become increasingly common, specifically including various documents generated during software development and system maintenance. Usually, after these documents are written, they need to be inspected to ensure their accuracy, consistency, integrity, standardization, and / or security. Currently, inspections are often carried out by manual review. For example, for database operation and maintenance documents, usually one person writes and another person reviews. However, the technical knowledge of the corresponding reviewers is not fully covered, and it is inevitable to encounter unfamiliar fields, and there are inevitably omissions in the process of manual review. Therefore, the efficiency and accuracy of reviewing and inspecting the corresponding documents cannot be guaranteed. Summary of the Invention
[0003] Embodiments of the present invention provide a document inspection method, apparatus, electronic device, and storage medium, which can improve the efficiency and accuracy of reviewing and inspecting documents to be inspected.
[0004] In a first aspect, an embodiment of the present invention provides a document inspection method, including:
[0005] Calculating a similar hash value of a document to be inspected based on a reference vocabulary to obtain a hash value of the document to be inspected, where the reference vocabulary includes each valid word in each candidate reference document;
[0006] Determining a reference document similar to the document to be inspected in each candidate reference document based on the hash value of the document to be inspected and the candidate document hash values corresponding to each candidate reference document, where the candidate document hash values corresponding to each candidate reference document are similar hash values of the corresponding candidate reference documents calculated based on the reference vocabulary;
[0007] Extracting feature vectors of each sentence in the reference document and the document to be inspected respectively to obtain a plurality of reference sentence vectors and a plurality of sentence vectors to be inspected; and
[0008] Based on the similarity between each sentence vector to be inspected and each reference sentence vector, determining a reference text corresponding to each sentence in the document to be inspected in the reference document, so that an inspector can inspect each sentence in the document to be inspected based on the corresponding reference text.
[0009] In a second aspect, an embodiment of the present invention provides a document inspection apparatus, including:
[0010] A module for obtaining the hash value of the document to be checked, which is used to calculate the similar hash value of the document to be checked based on the reference vocabulary to obtain the hash value of the document to be checked, and the reference vocabulary includes each valid word in each candidate reference document;
[0011] A reference document obtaining module, which is used to determine the reference documents similar to the document to be checked in each candidate reference document based on the hash value of the document to be checked and the candidate document hash values corresponding to each candidate reference document, and the candidate document hash values corresponding to each candidate reference document are the similar hash values of the corresponding candidate reference documents calculated based on the reference vocabulary;
[0012] A sentence vector obtaining module, which is used to extract the feature vectors of each sentence in the reference document and the document to be checked respectively, and correspondingly obtain a plurality of reference sentence vectors and a plurality of sentence vectors to be checked; and
[0013] A reference text obtaining module, which is used to determine the reference text corresponding to each sentence in the document to be checked in the reference document based on the similarity between each sentence vector to be checked and each reference sentence vector, so that the inspector can check each sentence in the document to be checked based on the corresponding reference text.
[0014] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the document inspection method according to any one of the embodiments of the present invention.
[0015] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the document inspection method according to any one of the embodiments of the present invention.
[0016] A document inspection method, device, electronic device, and storage medium provided by an embodiment of the present invention can efficiently screen out reference documents similar to the document to be checked by calculating the hash value of the document to be checked based on the reference vocabulary and determining the reference documents similar to the document to be checked in each candidate reference document based on the hash value of the document to be checked and the candidate document hash values corresponding to each candidate reference document; the embodiment of the present invention further extracts the sentence vectors of the reference document and the document to be checked, and determines the reference text similar to the sentences in each document to be checked more accurately based on the similarity between the sentence vectors of the reference document and the document to be checked, so that the inspector can check each sentence in the document to be checked based on the corresponding reference text, and can improve the efficiency and accuracy of the inspector when checking and reviewing the document without requiring the inspector to comprehensively master the corresponding technical knowledge. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.
[0018] Figure 1 is a schematic flowchart of a document inspection method provided by an embodiment of the present invention;
[0019] Figure 2 is another schematic flowchart of a document inspection method provided by an embodiment of the present invention;
[0020] Figure 3 is another schematic flowchart of a document inspection method provided by an embodiment of the present invention;
[0021] Figure 4 is another schematic flowchart of a document inspection method provided by an embodiment of the present invention;
[0022] Figure 5 is a schematic structural diagram of a document inspection device provided by an embodiment of the present invention;
[0023] Figure 6 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0024] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0026] Figure 1 This is a flowchart of the document inspection method provided by an embodiment of the present invention. This embodiment is applicable to scenarios where electronic documents generated during software development and system maintenance, such as database operation and maintenance documents, are inspected. This method can be executed by the document inspection device provided by an embodiment of the present invention, and the device can be implemented in software and / or hardware. In a specific embodiment, the device can be integrated into an electronic device, such as a computer, a server, etc. The following embodiments will be described by taking the integration of the device into an electronic device as an example. Refer to Figure 1 , the method specifically may include the following steps:
[0027] Step 101, calculate the similar hash value of the document to be inspected based on the reference vocabulary to obtain the hash value of the document to be inspected. The reference vocabulary includes each valid word in each candidate reference document. This step can facilitate determining the reference document from each candidate reference document based on the hash value of the document to be inspected and the hash value of the candidate document.
[0028] Specifically, the above hash value of the document to be inspected may be a 64-bit or 128-bit hash value.
[0029] Specifically, the above document to be inspected may be an electronic document generated during software development and system maintenance, such as a database operation and maintenance document.
[0030] Specifically, the above document to be inspected may also be other types of electronic documents, such as various office documents.
[0031] Specifically, the above reference vocabulary may be a vocabulary corresponding to each candidate reference document established based on the Bag-of-Words Model (BoW) or Term Frequency-Inverse Document Frequency (TF-IDF).
[0032] Specifically, each valid word in each candidate reference document can be understood as other words after removing the stop words and / or low-frequency words in the candidate reference document.
[0033] Optionally, the process of calculating the similar hash value of the document to be inspected based on the reference vocabulary to obtain the hash value of the document to be inspected includes: calculating the Simhash value of the document to be inspected based on the reference vocabulary through the Simhash algorithm to obtain the above hash value of the document to be inspected.
[0034] Optionally, the process of calculating the similarity hash value of the document to be inspected based on the reference vocabulary to obtain the hash value of the document to be inspected includes: determining the first weight of each word in the reference vocabulary in the document to be inspected based on the term frequency-inverse document frequency method; determining the second weight of each word in the reference vocabulary in the document to be inspected based on the position of each word in the reference vocabulary in the document to be inspected; and determining the comprehensive weight of each word in the vocabulary in the document to be inspected based on the first weight and the second weight, and calculating the hash value of the document to be inspected based on the comprehensive weight.
[0035] Optionally, the process of calculating the hash value of the document to be inspected based on the comprehensive weight includes: determining the word vectors of the corresponding words of the document to be inspected based on the comprehensive weights of the words, and obtaining the hash value of the document to be inspected through hash operation and bit operation based on the word vectors of the corresponding words of the document to be inspected.
[0036] Optionally, when using weight0 to represent the first weight of each word in the document to be inspected and using K to represent the first weight of each word in the document to be inspected, the comprehensive weight weight = weight0 * K.
[0037] Optionally, the process of determining the second weight of each word in the reference vocabulary in the document to be inspected based on the position of each word in the document to be inspected includes: determining the corresponding second weight based on whether each word is in the title of the corresponding document.
[0038] Specifically, the weight coefficient K of the corresponding word can be set according to the headings at all levels. The first-level heading is 2, and the keywords of each level of heading are +1 based on the weight coefficient of each level of heading. If the keyword is not included in the heading, the corresponding weight coefficient is determined according to the preset rule.
[0039] Step 102, determining the reference documents similar to the document to be inspected in each candidate reference document based on the hash value of the document to be inspected and the candidate document hash values corresponding to each candidate reference document. The candidate document hash values corresponding to each candidate reference document are the similarity hash values of the corresponding candidate reference documents calculated based on the reference vocabulary. This step can efficiently screen out the reference documents similar to the document to be inspected, and is conducive to extracting the sentence feature vectors of the reference documents and the document to be inspected, and then determining the reference text of each sentence in the document to be inspected based on the obtained sentence feature vectors.
[0040] Specifically, the number of bits of the above candidate document hash value is the same as that of the hash value of the document to be inspected, and can be a 64-bit or 128-bit hash value.
[0041] Specifically, the number of reference documents similar to the document to be inspected can be one or more.
[0042] Optionally, the process of determining the reference documents similar to the document to be inspected from each candidate reference document based on the hash value of the document to be inspected and the candidate document hash values corresponding to each candidate reference document includes: querying the reference document text library to attempt to obtain a target candidate document hash value whose Hamming distance from the hash value of the document to be inspected is not greater than a preset Hamming distance threshold; and when the target candidate document hash value is obtained, determining the candidate reference document corresponding to the target candidate document hash value as the reference document.
[0043] Specifically, the above Hamming distance threshold can be set based on experience, for example, it can be set to 3 or 2.
[0044] It can be understood that the Hamming distance can be used to calculate the similarity between samples. The smaller the Hamming distance, the higher the similarity between the corresponding samples. Therefore, a target candidate document hash value whose Hamming distance from the hash value of the document to be inspected is not greater than the Hamming distance threshold can be obtained, and the candidate reference document corresponding to the target candidate document hash value can be used as the reference document.
[0045] Optionally, the process of calculating the similar hash value of the corresponding candidate reference document based on the reference vocabulary includes: for each candidate reference document, determining the first weight of each word in the reference vocabulary in the current candidate reference document based on the term frequency-inverse document frequency method; determining the second weight of each word in the reference vocabulary in the current candidate reference document based on the position of each word in the reference vocabulary in the current candidate reference document; and determining the comprehensive weight of each word in the reference vocabulary in the current candidate reference document based on the corresponding first weight and the corresponding second weight; determining the word vectors of each word in the current candidate reference document based on the corresponding comprehensive weight, and obtaining the similar hash value of the current candidate reference document through hash operation and bit operation based on the corresponding word vectors.
[0046] Step 103: Extract the feature vectors of each sentence in the reference document and the document to be inspected respectively, and correspondingly obtain a plurality of reference sentence vectors and a plurality of sentence vectors to be inspected. This step can facilitate determining the reference text similar to the sentences in each document to be inspected more accurately based on the similarity of the sentence vectors of the reference document and the document to be inspected.
[0047] It can be understood that the inspector needs to check each sentence when checking the review document. Since the reference document determined based on the hash value of the document to be inspected includes multiple sentences, it is necessary to extract the sentence vectors of each sentence in the document to be inspected and the reference document, and then determine the reference text similar to the sentences in each document to be inspected more accurately based on the similarity of the sentence vectors of the reference document and the document to be inspected.
[0048] Optionally, the process of separately extracting the feature vectors of each sentence in the reference document and the document to be inspected includes: extracting the feature vectors of each sentence in the reference document and the document to be inspected based on the reference vocabulary.
[0049] Optionally, the process of extracting the feature vectors of each sentence in the reference document and the document to be inspected based on the reference vocabulary includes: for each sentence in the reference document and the document to be inspected, obtaining the word vectors of each word in the current sentence and combining the word vectors of each word in the current sentence to obtain the sentence vector of the current sentence.
[0050] Specifically, the process of extracting the feature vectors of each sentence in the reference document and the document to be inspected based on the reference vocabulary can also be performed based on the term frequency-inverse document frequency algorithm, and specifically can include: for each sentence in the reference document and the document to be inspected, calculating the term frequency-inverse document frequency values of each word in the current sentence, and using the term frequency-inverse document frequency values of each word in the current sentence as the vector elements of the sentence vector of the current sentence to obtain the sentence vector of the current sentence.
[0051] Specifically, when the number of reference documents is multiple, the process of separately extracting the feature vectors of each sentence in the reference documents and the document to be inspected, and correspondingly obtaining multiple reference sentence vectors and multiple sentence vectors to be inspected includes: extracting the feature vectors of each sentence in each reference document to obtain the above-mentioned multiple reference sentence vectors.
[0052] Step 104, based on the similarity between each sentence vector to be inspected and each reference sentence vector, determine the reference text in the reference document corresponding to each sentence in the document to be inspected, so that the inspector can inspect each sentence in the document to be inspected based on the corresponding reference text. Based on Steps 101 to 103, this step can efficiently screen out the reference documents similar to the document to be inspected by calculating the hash value of the document to be inspected based on the reference vocabulary and determining the reference documents similar to the document to be inspected among each candidate reference document based on the hash value of the document to be inspected and the candidate document hash values corresponding to each candidate reference document; the embodiment of the present invention further extracts the sentence vectors of the reference document and the document to be inspected, and based on the similarity between the sentence vectors of the reference document and the document to be inspected, more accurately determines the reference text similar to the sentences in each document to be inspected, so that the inspector can inspect each sentence in the document to be inspected based on the corresponding reference text, which can improve the efficiency and accuracy of the inspector when checking and reviewing the document without requiring the inspector to comprehensively master the corresponding technical knowledge.
[0053] Optionally, the process of determining the reference text corresponding to each sentence in the document to be inspected in the reference document based on the similarity between each sentence vector to be inspected and each reference sentence vector includes: calculating the cosine similarity between each sentence vector to be inspected and each reference sentence vector; and when the cosine similarity is greater than the similarity threshold, determining the sentence in the document to be inspected corresponding to the corresponding reference sentence vector as the reference text of the sentence in the document to be inspected.
[0054] Specifically, for text content, using the method of calculating cosine similarity of text vectors can more accurately match the content of the text.
[0055] Specifically, it is also possible to calculate the Euclidean distance between each sentence vector to be inspected and each reference sentence vector, and when the calculated Euclidean distance is less than the preset Euclidean distance threshold, determine the sentence in the document to be inspected corresponding to the corresponding reference sentence vector as the reference text of the sentence in the document to be inspected.
[0056] Optionally, the document inspection method provided by the embodiments of the present invention further includes: querying for sensitive words in the document to be inspected based on a sensitive word data source and returning the position information of the sensitive words to desensitize the sensitive words.
[0057] Optionally, the above-mentioned sensitive words can be, for example, words representing sensitive information such as login passwords, user contact information, and household ID numbers.
[0058] Specifically, the above-mentioned sensitive word data source can be the words adjacent to the position of the sensitive word. For example, for a login password, it is usually adjacent to a login account that is not a sensitive word. Therefore, the login account can be used as the sensitive word data source to try to match the possible login password.
[0059] Specifically, the process of desensitizing the sensitive words can be performed manually or automatically.
[0060] Specifically, the sensitive words can be desensitized by deleting or replacing the corresponding sensitive words.
[0061] The following further introduces the document inspection method provided by the embodiments of the present invention.
[0062] Optionally, each candidate reference document and the corresponding respective reference substring storage tables are stored in a reference document text library.
[0063] Optionally, as Figure 2 shown, the document inspection method provided by the embodiments of the present invention may include the following steps:
[0064] Step 201, calculating the similar hash value of the document to be inspected based on a reference vocabulary table to obtain the hash value of the document to be inspected.
[0065] Step 202: Query the reference document text library and attempt to obtain a target candidate document hash value whose Hamming distance from the hash value of the document to be inspected is not greater than a preset Hamming distance threshold.
[0066] Step 203: When the target candidate document hash value is obtained, determine the candidate reference document corresponding to the target candidate document hash value as the reference document.
[0067] Step 204: When the target candidate document hash value cannot be obtained, generate corresponding prompt information to enable the inspector to conduct a self-verification. Obtain and store the document to be inspected after the self-verification by the inspector as a new candidate reference document in the reference document text library, and store the hash value of the document to be inspected as the corresponding candidate document hash value in the hash value of the document to be inspected.
[0068] By establishing and updating the reference document text library in the embodiments of the present invention, it is possible to facilitate the provision of more comprehensive reference documents, making the review and inspection of the document to be inspected more convenient.
[0069] The following further introduces the document inspection method provided by the embodiments of the present invention.
[0070] Optionally, as Figure 3 shown, the document inspection method provided by the embodiments of the present invention may include the following steps:
[0071] Step 301: Calculate the similar hash value of the document to be inspected based on the reference vocabulary to obtain the hash value of the document to be inspected.
[0072] Step 302: Calculate the candidate document hash values corresponding to each candidate reference document based on the reference vocabulary.
[0073] Step 303: For each candidate document hash value, divide the string of the current candidate document hash value into multiple parts based on the first division rule to obtain multiple reference substrings.
[0074] Specifically, the above first division rule includes the number of parts for dividing the current candidate document, the share of each part, and the positions of each character in the divided substring in the string before division.
[0075] Specifically, the string of the current candidate document hash value can be evenly divided into four parts, or the string of the current candidate document hash value can be divided into four parts based on the first preset ratio.
[0076] Specifically, the string of the current candidate document hash value can also be divided into other numbers of parts, such as five parts.
[0077] Optionally, the process of dividing the string of the current candidate document hash value into multiple parts according to the first division rule to obtain multiple reference substrings includes: dividing the string of the current candidate document hash value into four parts to obtain a first reference substring, a second reference substring, a third reference substring, and a fourth reference substring.
[0078] Step 304: Based on the second division rule, divide the complementary substring of each reference substring into multiple segments to obtain multiple complementary substring segments corresponding to each reference substring.
[0079] Specifically, the second division rule includes the number of segments for dividing each reference substring, the share of each segment, and the positions of each character in the segment obtained in the complementary substring before the corresponding division.
[0080] Specifically, the complementary substring of each reference substring can be evenly divided into four segments, or the complementary substring of each reference substring can be divided into four parts based on other second preset ratios.
[0081] Specifically, the complementary substring of each reference substring can also be divided into other numbers of segments, such as five segments.
[0082] Optionally, the process of dividing the complementary substring of each reference substring into multiple segments based on the second division rule to obtain multiple complementary substring segments corresponding to each reference substring includes: dividing the complementary substring of the first reference substring, the complementary substring of the second reference substring, the complementary substring of the third reference substring, and the complementary substring of the fourth reference substring into four segments respectively, corresponding to obtaining four complementary substring segments of the first reference substring, four complementary substring segments of the second reference substring, four complementary substring segments of the third reference substring, and four complementary substring segments of the fourth reference substring.
[0083] Step 305: Combine each reference substring with the corresponding complementary substring segments respectively to obtain multiple storage substrings corresponding to each reference substring.
[0084] Optionally, the process of combining each reference substring with the corresponding complementary substring segments respectively to obtain multiple storage substrings corresponding to each reference substring includes: combining the first reference substring with the corresponding four complementary substring segments respectively to obtain four storage substrings corresponding to the first reference substring; combining the second reference substring with the corresponding four complementary substring segments respectively to obtain four storage substrings corresponding to the second reference substring; combining the third reference substring with the corresponding four complementary substring segments respectively to obtain four storage substrings corresponding to the third reference substring; combining the fourth reference substring with the corresponding four complementary substring segments respectively to obtain four storage substrings corresponding to the fourth reference substring.
[0085] Specifically, in the embodiments of the present invention, 16 stored substrings corresponding to each candidate document can be obtained.
[0086] Step 306: Store each of the stored substrings corresponding to the current candidate document hash value into the corresponding reference substring storage table, where the reference substring storage table corresponds one-to-one with each of the stored substrings corresponding to the current candidate document hash value.
[0087] Optionally, the number of the above reference substring storage tables is 16, each corresponding to one stored substring of each current candidate document hash value.
[0088] Specifically, steps 302 to 306 can also be executed before step 301.
[0089] Step 307: Query each reference substring storage table simultaneously based on the hash value of the document to be inspected, and determine the reference document based on the query results.
[0090] Specifically, the hash value of the document to be inspected can be directly used to query each reference substring storage table simultaneously, or after the string of the hash value of the document to be inspected is divided based on the above first division rule and second division rule, the corresponding string fragments can be used to query each reference substring storage table simultaneously.
[0091] Step 308: Extract the feature vectors of each sentence in the reference document and the document to be inspected respectively, and correspondingly obtain a plurality of reference sentence vectors and a plurality of sentence vectors to be inspected.
[0092] Step 309: Based on the similarity between each sentence vector to be inspected and each reference sentence vector, determine the reference text in the reference document corresponding to each sentence in the document to be inspected, so that the inspector can inspect each sentence in the document to be inspected based on the corresponding reference text.
[0093] In a specific example of the present invention, for the current candidate document hash value:
[0094] data = {
[0095] 0, 1, 0, 0, 1, 1, 1, 1, 0, 0, 1, 1, 1, 0, 1, 0,
[0096] 0, 0, 1, 0, 0, 1, 0, 0, 1, 1, 1, 0, 1, 0, 1, 1,
[0097] 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 1,
[0098] 0, 1, 1, 0, 1, 0, 0, 1, 1, 1, 1, 0, 0, 0, 1, 0,
[0099] }
[0100] It can be evenly divided into 4 parts to obtain 4 reference substrings:
[0101] A1: {0, 1, 0, 0, 1, 1, 1, 1, 0, 0, 1, 1, 1, 0, 1, 0,}
[0102] A2: {0, 0, 1, 0, 0, 1, 0, 0, 1, 1, 1, 0, 1, 0, 1, 1,}
[0103] A3: {1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 1,}
[0104] A4: {0, 1, 1, 0, 1, 0, 0, 1, 1, 1, 1, 0, 0, 0, 1, 0}
[0105] Then the complementary substring of A1 is:
[0106] data = {
[0107] 0, 0, 1, 0, 0, 1, 0, 0, 1, 1, 1, 0, 1, 0, 1, 1,
[0108] 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 1,
[0109] 0, 1, 1, 0, 1, 0, 0, 1, 1, 1, 1, 0, 0, 0, 1, 0,
[0110] }
[0111] Dividing it into 4 parts gives 4 complementary substring segments of A1:
[0112] B1[1]: {0, 0, 1, 0, 0, 1, 0, 0, 1, 1, 1, 0,}
[0113] B1[2] {1, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,}
[0114] B1[3]: {1, 1, 1, 1, 1, 1, 0, 1, 0, 1, 1, 0,}
[0115] B1[4]: {1, 0, 0, 1, 1, 1, 1, 0, 0, 0, 1, 0,}
[0116] Combining A1 with the corresponding 4 complementary substrings respectively, 4 storage substrings corresponding to A1 can be obtained:
[0117] A1 B1[1]: {0, 1, 0, 0, 1, 1, 1, 1, 0, 0, 1, 1, 1, 0, 1, 0, 0, 0, 1, 0, 0, 1, 0, 0, 1, 1, 1, 0,}
[0118] A1 B1[2]: {0,1,0,0,1,1,1,1,0,0,1,1,1,0,1,0,1,0,1,1,1,1,1,1,1,1,1,1}
[0119] A1 B1[3]: {0,1,0,0,1,1,1,1,0,0,1,1,1,0,1,0,1,1,1,1,1,1,0,1,0,1,1,0}
[0120] A1 B1[4]: {0,1,0,0,1,1,1,1,0,0,1,1,1,0,1,0,1,0,0,1,1,1,1,0,0,0,1,0}
[0121] And so on, the 4 storage substrings corresponding to A2, A3, and A4 respectively can be obtained: that is, the 16 storage substrings corresponding to the current candidate document hash value can be obtained.
[0122] A1B1[1], A1B1[2], A1B1[3], A1B1[4]
[0123] A2B2[1], A2B2[2], A2B2[3], A2B2[4]
[0124] A3B3[1], A3B3[2], A3B3[3], A3B3[4]
[0125] A4B4[1], A4B4[2], A4B4[3], A4B4[4]
[0126] The embodiments of the present invention can help reduce the time required to determine the reference document, thereby improving the inspection efficiency.
[0127] It can be understood that when the number of bits of the candidate document hash value and the document hash value to be inspected is 64 bits, the data range of each candidate document hash value is:
[0128] 0x0000000000000000 - 0xFFFFFFFFFFFFFFFF, which can represent 2 to the 64th power of data. If the reference document is directly determined by comparing the document hash value to be inspected and each candidate document hash value, it will take a long time. The embodiments of the present invention store each candidate document hash value in a sub-table, similar to setting up a first-level index, such as A1, A2, A3, and A4, and a second-level index, such as B1, B2, B3, and B4. The process of searching for a data table of 64-bit data can be transformed into the process of simultaneously searching for 16 data tables of 28-bit data, which can greatly save the time required for searching.
[0129] The following further introduces the document inspection method provided by the embodiments of the present invention. As Figure 4As shown, namely Figure 3 Step 307 in
[0130] Step 3071: Divide the string of the document hash value to be inspected into multiple parts based on the first division rule to obtain multiple substrings to be inspected.
[0131] Step 3072: Divide the complementary substrings of each substring to be inspected into multiple parts based on the second division rule to obtain multiple complementary substring segments of each substring to be inspected.
[0132] Step 3073: Combine each substring to be inspected with its corresponding complementary substring respectively to obtain multiple query substrings corresponding to each substring to be inspected.
[0133] Step 3074: Determine the reference substring storage table corresponding to each query substring based on the first division rule and the second division rule.
[0134] Optionally, the process of determining the reference substring storage table corresponding to each query substring based on the first division rule and the second division rule includes:
[0135] For each query substring, determine the reference substring storage table corresponding to the current query substring based on the positions of each character in the substring obtained by the corresponding division in the string before the division, and the positions of each character in the substring segment obtained by the corresponding division of the current query substring in the complementary substring before the division.
[0136] Step 3075: Try to query the target storage substring with the same characters as the corresponding query substring from the corresponding reference substring storage table based on each query substring.
[0137] Step 3076: When the target storage substring is queried, determine the reference document based on the number of target storage substrings corresponding to the same candidate reference document.
[0138] Optionally, the process of determining the reference document based on the number of target storage substrings corresponding to the same candidate reference document includes: When the number of target storage substrings corresponding to the same candidate reference document is not less than the preset storage substring number threshold, determine the candidate reference document as the reference document.
[0139] Optionally, when the number of target storage substrings corresponding to the same candidate reference document is less than the preset storage substring number threshold, it is determined that the reference document cannot be obtained, generate the corresponding prompt information to enable the inspector to check by himself, obtain and store the document to be inspected after the inspector's self-check as a new candidate reference document in the reference document text library, and store each query substring as a new storage substring in the corresponding reference substring storage table.
[0140] Specifically, the above-mentioned storage substring quantity threshold can be determined based on the aforementioned Hamming distance threshold, the number of reference substrings corresponding to each candidate document, and the number of complementary substring segments.
[0141] Optionally, when the Hamming distance threshold is 3, the number of reference substrings corresponding to each candidate document is 4, and the number of complementary substring segments corresponding to each candidate document is 4, the above-mentioned storage substring quantity threshold is determined to be 2.
[0142] It can be understood that based on the drawer principle: if there are 3 drawers and 4 books, then there must be two books in the same drawer. The Hamming distance threshold of 3 requires that the Hamming distance between the hash value of the document to be inspected and the similar hash value of the reference document is not greater than 3. If the number of reference substrings and complementary substring segments corresponding to each candidate document is 4, each candidate document can be divided into 5 segments. Since the Hamming distance between the hash value of the document to be inspected and the similar hash value of the reference document is not greater than 3, at least 2 query substrings in the query substrings corresponding to the hash value of the document to be inspected are equal to the storage substrings corresponding to the similar hash value of the reference document.
[0143] Therefore, determining the above-mentioned storage substring quantity threshold to be 2 can ensure that the Hamming distance between the hash value of the document to be inspected and the similar hash value of the reference document is not greater than 3, and a reference document highly similar to the document to be inspected can be obtained.
[0144] The embodiment of the present invention can further efficiently query and obtain a reference document highly similar to the document to be inspected based on the candidate document hash values stored in sub-tables.
[0145] Figure 5 FIG. is a schematic structural diagram of a document inspection device provided by an embodiment of the present invention. This device is suitable for executing the document inspection method provided by the embodiment of the present invention. As Figure 5 shown, this device may specifically include:
[0146] A module 501 for obtaining the hash value of the document to be inspected, which is used to calculate the similar hash value of the document to be inspected based on a reference vocabulary table to obtain the hash value of the document to be inspected. The reference vocabulary table includes each valid word in each candidate reference document. This module can facilitate determining a reference document from each candidate reference document based on the hash value of the document to be inspected and the candidate document hash values.
[0147] Optionally, the above-mentioned to-be-checked document hash value obtaining module 501 can specifically be used to determine the first weight of each word in the reference vocabulary in the to-be-checked document based on the term frequency-inverse document frequency method; determine the second weight of each word in the reference vocabulary in the to-be-checked document based on the position of each word in the reference vocabulary in the to-be-checked document; and determine the comprehensive weight of each word in the vocabulary in the to-be-checked document based on the first weight and the second weight, and calculate the to-be-checked document hash value based on the comprehensive weight.
[0148] The reference document obtaining module 502 is used to determine the reference documents similar to the to-be-checked document in each candidate reference document based on the to-be-checked document hash value and the candidate document hash values corresponding to each candidate reference document. The candidate document hash values corresponding to each candidate reference document are the similar hash values of the corresponding candidate reference documents calculated based on the reference vocabulary. This module can efficiently screen out the reference documents similar to the to-be-checked document, and is conducive to extracting the sentence feature vectors of the reference document and the to-be-checked document, and then determining the reference text of each sentence in the to-be-checked document based on the obtained sentence feature vectors.
[0149] Optionally, each of the above-mentioned candidate reference documents and the corresponding candidate document hash values are stored in the reference document text library.
[0150] Optionally, the above-mentioned reference document obtaining module 502 can specifically be used to query the reference document text library to try to obtain a target candidate document hash value whose Hamming distance from the to-be-checked document hash value is not greater than a preset Hamming distance threshold; query the reference document text library to try to obtain a target candidate document hash value whose Hamming distance from the to-be-checked document hash value is not greater than a preset Hamming distance threshold.
[0151] The sentence vector obtaining module 503 is used to extract the feature vectors of each sentence in the reference document and the to-be-checked document respectively, and correspondingly obtain a plurality of reference sentence vectors and a plurality of to-be-checked sentence vectors.
[0152] Optionally, the above-mentioned sentence vector obtaining module 503 can specifically be used to extract the feature vectors of each sentence in the reference document and the to-be-checked document based on the reference vocabulary. This module can facilitate and accurately determine the reference text similar to the sentence in each to-be-checked document based on the similarity of the sentence vectors of the reference document and the to-be-checked document.
[0153] A reference text acquisition module 504 is used to determine, based on the similarity between each sentence vector to be checked and each reference sentence vector, the reference text in the reference document corresponding to each sentence in the document to be checked, so that the inspector can check each sentence in the document to be checked based on the corresponding reference text. By combining modules 501 to 503, this module can calculate the hash value of the document to be checked based on the reference vocabulary table, and determine the reference document similar to the document to be checked in each candidate reference document based on the hash value of the document to be checked and the candidate document hash values corresponding to each candidate reference document, and can efficiently screen out the reference document similar to the document to be checked; in the embodiment of the present invention, further by extracting the sentence vectors of the reference document and the document to be checked, and based on the similarity between the sentence vectors of the reference document and the document to be checked, the reference text similar to the sentence in each document to be checked is determined more accurately, so that the inspector can check each sentence in the document to be checked based on the corresponding reference text, and can improve the efficiency and accuracy of the inspector when checking and reviewing the document without requiring the inspector to comprehensively master the corresponding technical knowledge.
[0154] Optionally, the above-mentioned reference text acquisition module 504 can specifically be used to calculate the cosine similarity between each sentence vector to be checked and each reference sentence vector; and when the cosine similarity is greater than the similarity threshold, determine the sentence in the document to be checked corresponding to the corresponding reference sentence vector as the reference text of the sentence in the corresponding document to be checked.
[0155] Optionally, the document inspection device provided in the embodiment of the present invention further includes a candidate document hash value storage module, which is used for: calculating the candidate document hash values corresponding to each candidate reference document based on the reference vocabulary table; dividing the string of the current candidate document hash value into multiple parts according to the first division rule to obtain multiple reference substrings; respectively dividing the complementary substrings of each reference substring into multiple segments according to the second division rule to obtain multiple complementary substring segments corresponding to each reference substring; combining each reference substring with the corresponding multiple complementary substring segments respectively to obtain multiple storage substrings corresponding to each reference substring; storing each storage substring corresponding to the current candidate document hash value into the corresponding reference substring storage table, and the reference substring storage table corresponds one-to-one with each storage substring corresponding to the current candidate document hash value.
[0156] Optionally, the aforementioned reference document acquisition module 502 can specifically be used to query each reference substring storage table based on the hash value of the document to be checked, and determine the reference document based on the query result.
[0157] Optionally, the foregoing reference document acquisition module 502 can specifically be configured to divide the string of the document hash value to be inspected into multiple parts based on a first division rule to obtain multiple substrings to be inspected; divide the complementary substrings of each substring to be inspected into multiple parts based on a second division rule to obtain multiple complementary substring segments of each substring to be inspected; combine each substring to be inspected with the corresponding complementary substrings respectively to obtain multiple query substrings corresponding to each substring to be inspected; determine a reference substring storage table corresponding to each query substring based on the first division rule and the second division rule; attempt to query a target stored substring with the same characters as the corresponding query substring from the corresponding reference substring storage table based on each query substring; when a target stored substring is queried, determine the reference document based on the number of target stored substrings corresponding to the same candidate reference document.
[0158] Optionally, the document inspection device provided by an embodiment of the present invention further includes a sensitive word processing module, configured to query sensitive words in the document to be inspected based on a sensitive word data source and return the position information of the sensitive words, so as to desensitize the sensitive words.
[0159] Optionally, the document inspection device provided by an embodiment of the present invention further includes a prompt and library update module, configured to generate corresponding prompt information to enable an inspector to check by himself when the hash value of the target candidate document cannot be obtained, obtain and store the document to be inspected after the inspector's self-check as a new candidate reference document in the reference document text library, and store the hash value of the document to be inspected as the corresponding candidate document hash value in the hash value of the document to be inspected.
[0160] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working process of the above-described functional modules can refer to the corresponding process in the foregoing method embodiment, and will not be elaborated herein.
[0161] An embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the document inspection method provided by any one of the foregoing embodiments is implemented.
[0162] An embodiment of the present invention further provides a computer-readable medium, on which a computer program is stored. When the program is executed by a processor, the document inspection method provided by any one of the foregoing embodiments is implemented.
[0163] An embodiment of the present invention further provides a computer program product, including a computer program which, when executed by a processor, implements the document inspection method as described in any one of the embodiments of the present invention.
[0164] Reference is now made to Figure 6 , which shows a schematic structural diagram of a computer system 600 of an electronic device suitable for implementing the embodiments of the present invention. Figure 6 The illustrated electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0165] As Figure 6 shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the system 600 are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0166] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as required. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as required so that a computer program read from it can be installed into the storage section 608 as required.
[0167] Specifically, according to the disclosed embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present invention discloses a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 609 and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above-mentioned functions defined in the system of the present invention are executed.
[0168] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0169] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0170] The modules and / or units involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules and / or units can also be provided in a processor. For example, it can be described as: a processor includes a module for obtaining the hash value of the document to be inspected, a module for obtaining the reference document, a module for obtaining the sentence vector, and a module for obtaining the reference text. Among them, the names of these modules do not constitute a limitation on the module itself in some cases.
[0171] As another aspect, the present invention also provides a computer-readable medium, which can be included in the device described in the above embodiments; or can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device is enabled to: calculate the similar hash value of the document to be inspected based on the reference vocabulary to obtain the hash value of the document to be inspected, where the reference vocabulary includes each valid word in each candidate reference document; determine the reference document similar to the document to be inspected in each candidate reference document based on the hash value of the document to be inspected and the candidate document hash values corresponding to each candidate reference document, and the candidate document hash values corresponding to each candidate reference document are the similar hash values of the corresponding candidate reference documents calculated based on the reference vocabulary; extract the feature vectors of each sentence in the reference document and the document to be inspected respectively to obtain a plurality of reference sentence vectors and a plurality of sentence vectors to be inspected correspondingly; and determine the reference text corresponding to each sentence in the document to be inspected in the reference document based on the similarity between each sentence vector to be inspected and each reference sentence vector, so that the inspector can inspect each sentence in the document to be inspected based on the corresponding reference text.
[0172] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A document inspection method, characterized in that: include: Calculate similar hash values of the document to be checked based on a reference vocabulary to obtain a hash value of the document to be checked, wherein the reference vocabulary includes each valid word in each candidate reference document; Determine a reference document similar to the document to be checked among the candidate reference documents based on the hash value of the document to be checked and the candidate document hash values corresponding to the candidate reference documents, wherein the candidate document hash values corresponding to the candidate reference documents are similar hash values of the corresponding candidate reference documents calculated based on the reference vocabulary; Extracting feature vectors of each sentence in the reference document and the document to be checked respectively, and obtaining a plurality of reference sentence vectors and a plurality of sentence vectors to be checked correspondingly; and Based on the similarity between each sentence vector to be checked and each reference sentence vector, a reference text in the reference document corresponding to each sentence in the document to be checked is determined, so that the checker can check each sentence in the document to be checked based on the corresponding reference text.
2. The document checking method according to claim 1, characterized in that: Before determining reference documents similar to the document to be checked among the candidate reference documents based on the hash value of the document to be checked and the candidate document hash values corresponding to the candidate reference documents, the document checking method further includes: Calculate a candidate document hash value corresponding to each candidate reference document based on the reference vocabulary; For each candidate document hash value, based on the first division rule, the character string of the current candidate document hash value is divided into multiple parts to obtain multiple reference substrings; Based on the second division rule, the complementary substring of each reference substring is divided into multiple segments to obtain multiple complementary substring segments corresponding to each reference substring; Combining each reference substring with each corresponding complementary substring fragment to obtain a plurality of storage substrings corresponding to each reference substring; and Store each storage substring corresponding to the hash value of the current candidate document into a corresponding reference substring storage table, wherein the reference substring storage table corresponds one-to-one to each storage substring corresponding to the hash value of the current candidate document; The determining of reference documents similar to the document to be checked among the candidate reference documents based on the hash value of the document to be checked and the candidate document hash values corresponding to the candidate reference documents includes: Each reference substring storage table is queried simultaneously based on the hash value of the document to be checked, and the reference document is determined based on the query result.
3. The document checking method according to claim 2, characterized in that: The step of simultaneously querying each reference substring storage table based on the hash value of the document to be checked, and determining the reference document based on the query result, includes: Divide the character string of the hash value of the document to be checked into multiple parts based on the first division rule to obtain multiple substrings to be checked; Based on the second division rule, the complementary substring of each substring to be checked is divided into multiple parts to obtain multiple complementary substring fragments of each substring to be checked; Combining each substring to be checked with each corresponding complementary substring to obtain multiple query substrings corresponding to each substring to be checked; Determine a reference substring storage table corresponding to each query substring based on the first division rule and the second division rule; Based on each query substring, try to search the corresponding reference substring storage table for a target storage substring with the same characters as the corresponding query substring; and When the target storage substring is found, the reference document is determined based on the number of the target storage substrings corresponding to the same candidate reference document.
4. The document checking method according to claim 1, characterized in that: Each candidate reference document and the corresponding candidate document hash value are stored in a reference document text library; The determining of reference documents similar to the document to be checked among the candidate reference documents based on the hash value of the document to be checked and the candidate document hash values corresponding to the candidate reference documents includes: Querying the reference document text library to try to obtain a target candidate document hash value whose Hamming distance with the hash value of the document to be checked is not greater than a preset Hamming distance threshold; and When the target candidate document hash value is obtained, determining the candidate reference document corresponding to the target candidate document hash value as the reference document; The document checking method further comprises: When the target candidate document hash value cannot be obtained, a corresponding prompt message is generated to enable the inspector to verify it by himself; the document to be checked after the inspector's self-verification is obtained and stored in the reference document text library as a new candidate reference document, and the hash value of the document to be checked is stored as the corresponding candidate document hash value in the hash value of the document to be checked.
5. The document checking method according to claim 1, characterized in that: The step of calculating the similar hash value of the document to be checked based on the reference vocabulary to obtain the hash value of the document to be checked includes: Determining a first weight of each word in the reference vocabulary in the document to be checked based on a word frequency-inverse document frequency method; Determining a second weight of each word in the reference vocabulary in the document to be checked based on the position of each word in the reference vocabulary in the document to be checked; and The comprehensive weight of each word in the vocabulary in the document to be checked is determined based on the first weight and the second weight, and the hash value of the document to be checked is calculated based on the comprehensive weight.
6. The document checking method according to claim 1, characterized in that: The step of respectively extracting feature vectors of each sentence in the reference document and the document to be checked comprises: Extracting feature vectors of each sentence in the reference document and the document to be checked based on the reference vocabulary; The determining, based on the similarity between each to-be-checked sentence vector and each reference sentence vector, a reference text in the reference document corresponding to each sentence in the to-be-checked document comprises: Calculating the cosine similarity between each of the to-be-checked sentence vectors and each of the reference sentence vectors; and When the cosine similarity is greater than the similarity threshold, the sentence in the document to be checked corresponding to the corresponding reference sentence vector is determined as the reference text of the sentence in the document to be checked.
7. The document checking method according to claim 1, characterized in that: Also includes: Based on the sensitive word data source, sensitive words in the document to be checked are queried and location information of the sensitive words is returned to perform desensitization processing on the sensitive words.
8. A document inspection device, characterized in that: include: A module for obtaining a hash value of a document to be checked, configured to calculate a similar hash value of the document to be checked based on a reference vocabulary to obtain a hash value of the document to be checked, wherein the reference vocabulary includes each valid word in each candidate reference document; A reference document acquisition module, configured to determine a reference document similar to the document to be checked among the candidate reference documents based on a hash value of the document to be checked and candidate document hash values corresponding to each candidate reference document, wherein the candidate document hash value corresponding to each candidate reference document is a similar hash value of the corresponding candidate reference document calculated based on the reference vocabulary; A sentence vector acquisition module, used to extract feature vectors of each sentence in the reference document and the document to be checked, and obtain a plurality of reference sentence vectors and a plurality of sentence vectors to be checked respectively; as well as The reference text acquisition module is used to determine the reference text in the reference document corresponding to each sentence in the document to be checked based on the similarity between each sentence vector to be checked and each reference sentence vector, so that the checker can check each sentence in the document to be checked based on the corresponding reference text.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the document checking method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the document checking method according to any one of claims 1 to 7 is implemented.