Document screening method and device, equipment and medium

By employing a dual detection mechanism of explicit rule matching and semantic similarity detection, the problem of the inability to identify implicit sensitive fields in existing technologies is solved, achieving high-accuracy screening of bank documents and ensuring comprehensive identification of both explicit and implicit sensitive fields.

CN120874818APending Publication Date: 2025-10-31AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511031893.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies cannot effectively identify implicitly sensitive fields with semantic variations such as paraphrasing and typos when filtering bank documents, resulting in low accuracy in filtering target documents.

Method used

A dual detection mechanism of explicit rule matching and semantic similarity detection is adopted. By extracting document content and segmenting it into short sentences, sensitive fields are identified using preset keywords and regular expressions, and short sentences with high semantic similarity are identified through semantic similarity matching. A preset sentence library is constructed for semantic similarity calculation to determine the target document.

Benefits of technology

It improves the accuracy of target document filtering, ensures that explicit sensitive fields are not missed, and can capture implicit sensitive fields, thereby enhancing the comprehensiveness and accuracy of filtering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120874818A_ABST
    Figure CN120874818A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a document screening method and device, equipment and a medium, and relates to the technical field of data processing. The method comprises the following steps: acquiring a plurality of documents to be detected; the method comprises the following steps: extracting text contents of a plurality of to-be-detected documents, and segmenting the text contents into short sentences to generate a short sentence set; the short sentences in the short sentence set are matched with preset keywords and / or preset regular expressions, and the short sentences meeting the matching condition are determined as a first short sentence set; semantic similarity matching is carried out on the short sentences in the short sentence set and preset key sentences, and the short sentences with the semantic similarity higher than a preset threshold value are determined as a second short sentence set; determining a target short sentence by taking a union set of the first short sentence set and the second short sentence set; and determining the to-be-detected document comprising the target short sentence as a target document. Therefore, according to the document screening method provided by the embodiment of the invention, through a dual detection mechanism of explicit rule matching and semantic similarity detection, the screening accuracy of the target document can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a document screening method, apparatus, device, and medium. Background Technology

[0002] In the daily operations of banks, a large number of documents containing sensitive information (hereinafter referred to as "target documents") need to be processed, such as customer contracts, audit reports, and investment prospectuses. These target documents need to undergo rigorous anonymization before being circulated internally or disclosed externally to ensure compliance with legal and regulatory requirements. Therefore, accurately identifying target documents has become a pressing technical problem that needs to be solved.

[0003] Currently, explicit sensitive fields (such as "bank card number" or "ID card number") in a document are usually matched using methods such as keyword matching or regular expression matching. If a document contains explicit sensitive fields, it is determined to be the target document.

[0004] However, the above methods are insufficient in recognizing semantic variations such as paraphrasing and misspellings. For example, these methods cannot detect implicitly sensitive fields such as "bank card number" or "ID card number" that have been intentionally or unintentionally altered, resulting in low accuracy in filtering target documents. Summary of the Invention

[0005] To address the aforementioned issues, this application provides a document screening method, apparatus, device, and medium that can improve the accuracy of screening target documents.

[0006] The embodiments of this application disclose the following technical solutions:

[0007] Firstly, this application discloses a document filtering method, the method comprising:

[0008] Retrieve multiple documents to be detected;

[0009] By extracting the text content of the multiple documents to be detected and segmenting the text content into short sentences, a set of short sentences is generated;

[0010] By matching the short sentences in the short sentence set with preset keywords and / or preset regular expressions, the short sentences that meet the matching conditions are determined as the first short sentence set;

[0011] By performing semantic similarity matching between the short sentences in the short sentence set and preset key sentences, short sentences with semantic similarity higher than a preset threshold are identified as the second short sentence set;

[0012] The target short sentence is determined by taking the union of the first short sentence set and the second short sentence set;

[0013] The document to be detected, which includes the target phrase, is identified as the target document.

[0014] Optionally, the step of performing semantic similarity matching between the short sentences in the short sentence set and preset key sentences includes:

[0015] The short sentences in the short sentence set are semantically similar to the preset key sentences in the preset sentence library. The preset sentence library is constructed as follows:

[0016] Construct a first type of preset key sentence based on the preset keywords and / or the preset regular expression;

[0017] Based on entity naming recognition technology and regular expressions, the entities in the first type of preset key sentences are identified, and the entities in the first type of preset key sentences are replaced using a language model to obtain the second type of preset key sentences;

[0018] The second type of preset key sentences are subjected to synonym conversion to obtain the third type of preset key sentences;

[0019] A preset sentence library is constructed by integrating the first type of preset key sentences, the second type of preset key sentences, and the third type of preset key sentences.

[0020] Optionally, the step of segmenting the text content into short sentences includes:

[0021] Based on the punctuation and segmentation information in the text content, the text content is divided into short sentences.

[0022] Optionally, the short sentences with semantic similarity higher than a preset threshold are identified as a second set of short sentences, including:

[0023] Determine the business scenario information of the short sentence;

[0024] Short sentences with semantic similarity higher than a preset threshold corresponding to the business scenario information are identified as the second set of short sentences.

[0025] Optionally, the method further includes:

[0026] Determine the number of times the target phrase appears in each of the target documents;

[0027] Target documents whose frequency exceeds a certain threshold are identified as risky documents.

[0028] Secondly, this application discloses a document filtering device, the device comprising: a document acquisition module, a sentence segmentation module, a first matching module, a second matching module, a sentence merging module, and a document determination module;

[0029] The document acquisition module is used to acquire multiple documents to be detected;

[0030] The short sentence segmentation module is used to extract the text content of the multiple documents to be detected, segment the text content into short sentences, and generate a short sentence set;

[0031] The first matching module is used to determine the short sentences that meet the matching conditions as the first short sentence set by matching the short sentences in the short sentence set with preset keywords and / or preset regular expressions;

[0032] The second matching module is used to determine the short sentences with a semantic similarity higher than a preset threshold as the second short sentence set by performing semantic similarity matching between the short sentences in the short sentence set and preset key sentences;

[0033] The short sentence merging module is used to determine the target short sentence by taking the union of the first short sentence set and the second short sentence set;

[0034] The document determination module is used to determine that the document to be detected, which includes the target phrase, is the target document.

[0035] Optionally, the second matching module is specifically used to: perform semantic similarity matching between short sentences in the short sentence set and preset key sentences in the preset sentence library, wherein the construction unit of the preset sentence library is as follows:

[0036] The first construction unit is used to construct a first type of preset key sentence based on the preset keywords and / or the preset regular expression;

[0037] The second construction unit is used to identify entities in the first type of preset key sentences based on entity naming recognition technology and regular expressions, and to replace the entities in the first type of preset key sentences with language models to obtain the second type of preset key sentences.

[0038] The third construction unit is used to perform synonym conversion on the second type of preset key sentences to obtain the third type of preset key sentences;

[0039] The fourth construction unit is used to construct a preset sentence library by integrating the first type of preset key sentences, the second type of preset key sentences, and the third type of preset key sentences.

[0040] Optionally, the sentence segmentation module is specifically used to: segment the text content into sentences based on punctuation and segmentation information in the text content.

[0041] Optionally, the second matching module is specifically used to: determine the business scenario information of the short sentence; and determine the short sentences with a semantic similarity higher than a preset threshold corresponding to the business scenario information as the second set of short sentences.

[0042] Optionally, the device further includes: a number of times determination module and a document determination module;

[0043] The frequency determination module is used to determine the frequency of occurrence of the target phrase in each target document;

[0044] The document determination module is used to determine that target documents whose occurrence frequency exceeds a threshold are risk documents.

[0045] Thirdly, this application discloses a document screening device, the device comprising: a memory and a processor;

[0046] The memory is used to store programs;

[0047] The processor is configured to execute the program to implement the various steps of the document filtering method as described in the first aspect.

[0048] Fourthly, this application discloses a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the various steps of the document filtering method as described in the first aspect.

[0049] Compared with the prior art, this application has the following beneficial effects:

[0050] This application provides a document screening method, apparatus, device, and medium. The method includes: acquiring multiple documents to be detected; extracting the text content of the multiple documents to be detected and segmenting the text content into short sentences to generate a set of short sentences; matching the short sentences in the set of short sentences with preset keywords and / or preset regular expressions to determine short sentences that meet the matching conditions as a first set of short sentences; performing semantic similarity matching between the short sentences in the set of short sentences and preset key sentences to determine short sentences with semantic similarity higher than a preset threshold as a second set of short sentences; determining a target short sentence by taking the union of the first set of short sentences and the second set of short sentences; and determining the document to be detected that includes the target short sentence as the target document. Therefore, the document screening method provided by this application, through a dual detection mechanism of explicit rule matching and semantic similarity detection, ensures that explicit sensitive fields are not missed (first set of short sentences) and that implicit sensitive fields are captured (second set of short sentences), thereby improving the accuracy of target document screening. Attached Figure Description

[0051] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0052] Figure 1 A flowchart illustrating a document filtering method provided in this application embodiment;

[0053] Figure 2 A flowchart illustrating a method for constructing a preset sentence library as provided in an embodiment of this application;

[0054] Figure 3 A schematic diagram of a document filtering device provided in an embodiment of this application;

[0055] Figure 4 This is a schematic diagram of a computer-readable medium provided in an embodiment of this application. Detailed Implementation

[0056] As described earlier, explicit sensitive fields (such as "bank card number" or "ID card number") in a document are usually matched using keywords or regular expressions. If a document contains explicit sensitive fields, then the document is identified as the target document.

[0057] However, the aforementioned techniques have two drawbacks: First, the methods are insufficient in recognizing semantic variations such as paraphrasing and misspellings. For example, they cannot detect implicitly sensitive fields such as "bank card number" or "ID card number" that have been intentionally or unintentionally altered, resulting in low accuracy in filtering target documents. Second, when processing scanned or PDF documents using these methods, OCR text extraction is required. However, OCR text extraction is prone to character obfuscation (e.g., "0 / O", "1 / I") and symbol omissions (e.g., missing "¥"), making sensitive fields easily misclassified as non-sensitive fields, again leading to low accuracy in filtering target documents.

[0058] The inventors, through research, have proposed a document screening method, apparatus, device, and medium. The method includes: acquiring multiple documents to be detected; extracting the text content of the multiple documents to be detected and segmenting the text content into short sentences to generate a set of short sentences; matching the short sentences in the set of short sentences with preset keywords and / or preset regular expressions to determine short sentences that meet the matching conditions as a first set of short sentences; performing semantic similarity matching between the short sentences in the set of short sentences and preset key sentences to determine short sentences with semantic similarity higher than a preset threshold as a second set of short sentences; determining the target short sentence by taking the union of the first and second sets of short sentences; and determining the document to be detected containing the target short sentence as the target document. Therefore, the document screening method provided in this application, through a dual detection mechanism of explicit rule matching and semantic similarity detection, ensures that explicit sensitive fields are not missed (first set of short sentences) and that implicit sensitive fields are captured (second set of short sentences), thereby improving the accuracy of target document screening.

[0059] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0060] See Figure 1 The figure is a flowchart of a document filtering method provided in an embodiment of this application. The method includes:

[0061] S101: Obtain multiple documents to be detected.

[0062] In one specific implementation, the document to be detected may include the following formats: .doc / .docx, .pptx, .pdf, .xlsx, and .txt. This application does not limit the specific format.

[0063] S102: Extract the text content of multiple documents to be detected, and divide the text content into short sentences to generate a set of short sentences.

[0064] In one specific implementation, text content such as body text, tables, and annotations in the document to be detected can be extracted, and the text content can be divided into short sentences according to punctuation information (such as periods, semicolons, etc.) and segmentation information (such as line breaks, etc.) to generate a set of short sentences.

[0065] S103: By matching the short sentences in the short sentence set with preset keywords and / or preset regular expressions, the short sentences that meet the matching conditions are determined as the first short sentence set.

[0066] In one specific implementation, multiple sensitive fields can be set as preset keywords, such as "bank card number," "ID card number," and "password." Then, short phrases containing these preset keywords are added to the first set of short phrases. For example, the phrase "Please provide your bank card number" includes the preset keyword "bank card number," and is therefore added to the first set of short phrases.

[0067] In one specific implementation, a regular expression can be set as a preset regular expression, such as "phone number: ^1[3456789]\d{9}$". Subsequently, short phrases that can be matched by the preset regular expression are added to the first short phrase set. For example, the short phrase "Zhang San's phone number is 139XXXXXXXX" matches the preset regular expression and is therefore added to the first short phrase set.

[0068] S104: By matching the semantic similarity of short sentences in the short sentence set with preset key sentences, short sentences with semantic similarity higher than a preset threshold are identified as the second short sentence set.

[0069] In one specific implementation, the short sentences in the short sentence set can be semantically similar to the preset key sentences in the preset sentence library, and the short sentences with a semantic similarity higher than a preset threshold can be identified as the second short sentence set.

[0070] See Figure 2 This figure is a flowchart illustrating a method for constructing a preset sentence library according to an embodiment of this application. Figure 2 As shown, the preset key sentences in the preset sentence library can come from the following four sources:

[0071] Source 1: Users manually design preset key sentences according to actual needs and add them to the preset sentence library. For example, short sentences that have been identified by preset keywords and / or preset regular expressions in the past can be added to the preset sentence library.

[0072] Source 2: Based on the recognition results of preset keywords and preset regular expressions, determine the newly added preset key sentences and add them to the preset sentence library.

[0073] It should be noted that when adding new preset key sentences to the preset sentence library, the cosine similarity between the new preset key sentence and existing sentences in the preset sentence library needs to be calculated. Only when the similarity between the new preset key sentence and existing sentences in the preset sentence library is less than the cosine similarity threshold (e.g., 0.9) will the new preset key sentence be added to the preset sentence library. This avoids having too many duplicate or highly similar sentences, thus reducing the computational load of subsequent semantic similarity calculations.

[0074] Source 3: First, using named entity recognition technology and regular expressions, identify entities (such as name, mobile phone number, and amount) in the preset key sentences from Source 1 and Source 2, and replace the entities with placeholders (such as [Name], [Mobile Phone Number]) to generate templated preset key sentences. For example, the preset key sentence from Source 1 is: Customer Zhang San, mobile phone number 13800XXXXXX, has consumed 200 yuan in phone bills this month. Then, the templated preset key sentence is: Customer [Name], mobile phone number [Mobile Phone Number] has consumed [Amount] in phone bills this month.

[0075] Next, set the generic placeholder to the masking symbol MASK to generate the initial variant. For example, the initial variant is: Customer [MASK], Mobile number [MASK] has consumed [MASK] this month.

[0076] Subsequently, a pre-trained language model (such as a generative model) is used to generate k semantically correct candidate values ​​(e.g., 3, 5, etc.) for each mask. For example, the 3 candidate values ​​for name [MASK] could be: Zhang San, Li Si, Wang Wu; the 3 candidate values ​​for amount [MASK] could be: 200 yuan, 120.5 yuan, 300 yuan; and the 3 candidate values ​​for date [MASK] could be: today, within 24 hours, within 3 days.

[0077] Finally, the k candidate values ​​are filled into the initial variant to obtain multiple preset key sentences, and the preset key sentences are added to the preset sentence library. For example, the preset key sentence could be: Customer Wang Wu, mobile phone number 13800XXXXXX, has consumed 120.5 yuan in phone bills this month.

[0078] Source 4: The preset key sentences generated from Source 3 are processed using a synonym list to obtain multiple new preset key sentences, which are then added to the preset sentence library. For example, the synonym list can be shown in Table 1 below:

[0079] Table 1

[0080]

[0081] For example, the new preset key sentence could be: User Wang Wu, mobile phone number 13800XXXXXX, has spent 120.5 yuan on phone bills this month.

[0082] Understandably, a pre-defined sentence library is constructed by integrating the pre-defined key sentences determined from Source 1, Source 2, Source 3, and Source 4.

[0083] After determining the preset key sentences in the preset sentence library using the above method, the semantic similarity between each short sentence in the short sentence set and the preset key sentences in the preset sentence library can be calculated. Specifically, Sentence-BERT is first used to encode each short sentence in the short sentence set into a corresponding first vector, and each preset key sentence in the preset sentence library is encoded into a corresponding second vector. Sentence-BERT is a model for generating sentence vectors, which can encode the semantic information of sentences into numerical vectors, facilitating subsequent similarity calculation. Subsequently, when it is necessary to calculate the semantic similarity of a certain first vector, the semantic similarity is calculated using the following formula (1):

[0084] (1)

[0085] Where u is the first vector, v is the second vector, and sim is the semantic similarity.

[0086] The sim value calculated by formula (1) is between -1 and 1. The closer the value is to 1, the more similar the two sentences are semantically. If the sim value is higher than the preset threshold (e.g., 0.8), the short sentence is added to the second set of short sentences. This means that when the short sentence is semantically similar enough to the preset key sentence, the short sentence will be selected and enter the subsequent processing flow.

[0087] In one specific implementation, the aforementioned preset threshold can be dynamically adjusted according to the business scenario. This means that the standard for judging the semantic similarity between short sentences and preset key sentences can be changed under different business needs to better adapt to diverse business situations.

[0088] Specifically, the first step is to determine the business scenario information of the short sentences, such as scenarios involving the external transmission of data files (from the internal network to the Internet) or the internal circulation of data files. Next, based on the business scenario information, a corresponding preset threshold is determined, and the semantic similarity between the short sentence and a preset key sentence is compared with this specific preset threshold. If the semantic similarity is higher than the corresponding preset threshold, then the short sentence is added to the second set of short sentences.

[0089] In one example, if information in a data transmission scenario is not properly filtered, it can easily lead to the leakage of sensitive information. Any sentence with a certain semantic similarity to a preset key sentence should be added to the second set of short sentences. Therefore, the preset threshold for the data file transmission scenario can be 0.8. This means that as long as the semantic similarity between a short sentence and a preset key sentence reaches 0.8 or higher, they are considered to be semantically matched, thereby identifying risky documents as much as possible and preventing sensitive information from flowing to the outside world.

[0090] In another example, information circulating within a data file, such as internal training materials, primarily circulates only within the internal network and among employees, minimizing the risk of sensitive information leakage. Therefore, the preset threshold for this data file circulation scenario could be 0.9. That is, only sentences with a semantic similarity exceeding 0.9 are considered semantically related and added to the second set of short sentences.

[0091] S105: Determine the target short sentence by taking the union of the first set of short sentences and the second set of short sentences.

[0092] In set operations, the union of two sets combines all elements from two sets, removing duplicates to form a new set. The first set of short sentences is filtered by matching short sentences against preset keywords and / or preset regular expressions. It primarily captures short sentences containing explicit sensitive keywords or conforming to a specific format. The second set of short sentences is filtered through semantic similarity matching, identifying short sentences that are highly semantically similar to preset keywords. This set can discover short sentences that, while not containing preset keywords or conforming to preset regular expressions, are semantically related to sensitive content. By taking the union, the results of these two different filtering methods can be combined to collect potentially problematic short sentences as comprehensively as possible, thus identifying target short sentences. This avoids the limitations of a single filtering method and improves the comprehensiveness of capturing sensitive or target content.

[0093] S106: Identify the document to be detected that contains the target phrase as the target document.

[0094] Iterate through all documents to be tested, checking each document for the target phrase identified in the previous steps. If a document contains at least one target phrase, it is identified as the target document.

[0095] In one specific implementation, after identifying the target documents, the frequency of occurrence of the target phrase in each target document can be determined, and target documents with a frequency exceeding a threshold are identified as risk documents. For example, the frequency threshold can be 1. Therefore, identifying risk documents helps to further filter out documents with a higher risk level, enabling relevant personnel to prioritize the de-identification of these risk documents.

[0096] In summary, the embodiments of this application provide a document filtering method. The document filtering method provided by the embodiments of this application, through a dual detection mechanism of explicit rule matching and semantic similarity detection, ensures that explicit sensitive fields are not missed (first short sentence set) and implicit sensitive fields are captured (second short sentence set), thereby improving the filtering accuracy of target documents.

[0097] See Figure 3 The figure is a schematic diagram of a document filtering device provided in an embodiment of this application. The document filtering device 300 includes: a document acquisition module 301, a sentence segmentation module 302, a first matching module 303, a second matching module 304, a sentence merging module 305, and a document determination module 306;

[0098] Document acquisition module 301 is used to acquire multiple documents to be detected;

[0099] The short sentence segmentation module 302 is used to extract the text content of multiple documents to be detected, segment the text content into short sentences, and generate a short sentence set;

[0100] The first matching module 303 is used to determine the short sentences that meet the matching conditions as the first short sentence set by matching the short sentences in the short sentence set with preset keywords and / or preset regular expressions;

[0101] The second matching module 304 is used to determine the short sentences with a semantic similarity higher than a preset threshold as the second short sentence set by performing semantic similarity matching between the short sentences in the short sentence set and the preset key sentences.

[0102] The short sentence merging module 305 is used to determine the target short sentence by taking the union of the first short sentence set and the second short sentence set;

[0103] Document determination module 306 is used to determine the document to be detected that includes the target phrase as the target document.

[0104] In one specific implementation, the second matching module 304 is specifically used to: perform semantic similarity matching between short sentences in the short sentence set and preset key sentences in the preset sentence library, wherein the construction unit of the preset sentence library is as follows:

[0105] The first construction unit is used to construct the first type of preset key sentences based on preset keywords and / or preset regular expressions;

[0106] The second building unit is used to identify entities in the first type of preset key sentences based on entity naming recognition technology and regular expressions, and to replace the entities in the first type of preset key sentences with language models to obtain the second type of preset key sentences.

[0107] The third building unit is used to perform synonym conversion on the second type of preset key sentences to obtain the third type of preset key sentences;

[0108] The fourth building unit is used to construct a preset sentence library by integrating the first type of preset key sentences, the second type of preset key sentences, and the third type of preset key sentences.

[0109] In one specific implementation, the short sentence segmentation module 302 is specifically used to: segment the text content into short sentences based on the punctuation and segmentation information in the text content.

[0110] In one specific implementation, the second matching module 304 is specifically used to: determine the business scenario information of the short sentences; and determine the short sentences with semantic similarity higher than the preset threshold corresponding to the business scenario information as the second set of short sentences.

[0111] In one specific implementation, the document filtering device 300 further includes: a count determination module and a document determination module;

[0112] The frequency determination module is used to determine the frequency of the target phrase in each target document;

[0113] The document identification module is used to identify target documents that appear more than a certain number of times as risk documents.

[0114] In summary, the embodiments of this application provide a document filtering device. The document filtering device provided by the embodiments of this application, through a dual detection mechanism of explicit rule matching and semantic similarity detection, ensures that explicit sensitive fields are not missed (first short sentence set) and implicit sensitive fields are captured (second short sentence set), thereby improving the filtering accuracy of target documents.

[0115] This application also provides corresponding document filtering devices and computer-readable media for implementing the document filtering method provided in this application.

[0116] The document screening device includes a memory and a processor. The memory is used to store instructions or code, and the processor is used to execute the instructions or code to cause the device to perform a document screening method according to any embodiment of this application.

[0117] See Figure 4 This figure is a schematic diagram of a computer-readable medium provided in an embodiment of this application. The computer-readable medium 400 stores a computer program 411, which, when executed by a processor, implements the above-described... Figure 1 The steps of the document filtering method.

[0118] It should be noted that, in the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0119] It should be noted that the machine-readable medium described above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0120] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0121] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

[0122] While several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0123] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.

Claims

1. A document filtering method, characterized in that, The method includes: Retrieve multiple documents to be detected; By extracting the text content of the multiple documents to be detected and segmenting the text content into short sentences, a set of short sentences is generated; By matching the short sentences in the short sentence set with preset keywords and / or preset regular expressions, the short sentences that meet the matching conditions are determined as the first short sentence set; By performing semantic similarity matching between the short sentences in the short sentence set and preset key sentences, short sentences with semantic similarity higher than a preset threshold are identified as the second short sentence set; The target short sentence is determined by taking the union of the first short sentence set and the second short sentence set; The document to be detected, which includes the target phrase, is identified as the target document.

2. The method according to claim 1, characterized in that, The step of performing semantic similarity matching between the short sentences in the short sentence set and preset key sentences includes: The short sentences in the short sentence set are semantically similar to the preset key sentences in the preset sentence library. The preset sentence library is constructed as follows: Construct a first type of preset key sentence based on the preset keywords and / or the preset regular expression; Based on entity naming recognition technology and regular expressions, the entities in the first type of preset key sentences are identified, and the entities in the first type of preset key sentences are replaced using a language model to obtain the second type of preset key sentences; The second type of preset key sentences are subjected to synonym conversion to obtain the third type of preset key sentences; A preset sentence library is constructed by integrating the first type of preset key sentences, the second type of preset key sentences, and the third type of preset key sentences.

3. The method according to claim 1, characterized in that, The step of segmenting the text content into short sentences includes: Based on the punctuation and segmentation information in the text content, the text content is divided into short sentences.

4. The method according to claim 1, characterized in that, The short sentences with semantic similarity higher than a preset threshold are identified as the second set of short sentences, including: Determine the business scenario information of the short sentence; Short sentences with semantic similarity higher than a preset threshold corresponding to the business scenario information are identified as the second set of short sentences.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: Determine the number of times the target phrase appears in each of the target documents; Target documents whose frequency exceeds a certain threshold are identified as risky documents.

6. A document filtering device, characterized in that, The device includes: a document acquisition module, a sentence segmentation module, a first matching module, a second matching module, a sentence merging module, and a document determination module; The document acquisition module is used to acquire multiple documents to be detected; The short sentence segmentation module is used to extract the text content of the multiple documents to be detected, segment the text content into short sentences, and generate a short sentence set; The first matching module is used to determine the short sentences that meet the matching conditions as the first short sentence set by matching the short sentences in the short sentence set with preset keywords and / or preset regular expressions; The second matching module is used to determine the short sentences with a semantic similarity higher than a preset threshold as the second short sentence set by performing semantic similarity matching between the short sentences in the short sentence set and preset key sentences; The short sentence merging module is used to determine the target short sentence by taking the union of the first short sentence set and the second short sentence set; The document determination module is used to determine that the document to be detected, which includes the target phrase, is the target document.

7. The apparatus according to claim 6, characterized in that, The second matching module is specifically used to: perform semantic similarity matching between short sentences in the short sentence set and preset key sentences in the preset sentence library, wherein the construction unit of the preset sentence library is shown below: The first construction unit is used to construct a first type of preset key sentence based on the preset keywords and / or the preset regular expression; The second construction unit is used to identify entities in the first type of preset key sentences based on entity naming recognition technology and regular expressions, and to replace the entities in the first type of preset key sentences with language models to obtain the second type of preset key sentences. The third construction unit is used to perform synonym conversion on the second type of preset key sentences to obtain the third type of preset key sentences; The fourth construction unit is used to construct a preset sentence library by integrating the first type of preset key sentences, the second type of preset key sentences, and the third type of preset key sentences.

8. The apparatus according to claim 6, characterized in that, The sentence segmentation module is specifically used to segment the text content into sentences based on punctuation and segmentation information in the text content.

9. A document filtering device, characterized in that, The device includes: a memory and a processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the document filtering method as described in any one of claims 1 to 5.

10. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements each step of the document filtering method as described in any one of claims 1 to 5.