Open format document sensitive word detection method and system, program product and equipment
By analyzing open-type documents and extracting text information, combined with sensitive vocabulary screening, the problems of slow detection speed and low accuracy in the existing technology are solved, and sensitive words in open-type documents are fully and effectively detected in one analysis process.
Patent Information
- Application Number
- CN202510555872.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-08
AI Technical Summary
When detecting sensitive words in open-type documents, the prior art has problems such as slow detection speed, low accuracy, and missing, out of order or damage to the content of the streaming document, making it difficult to achieve one-stop comprehensive and effective detection.
By analyzing the open-type documents to be uploaded, text information of the main text, notes, attached pictures, watermarks, seals and signatures are extracted, and filtered based on the pre-established sensitive thesaurus to achieve one-stop detection.
It realizes the comprehensive extraction of text information involved in various contents in open-type documents during a parsing process, improves detection efficiency and accuracy, avoids content omissions and damage, and ensures the effectiveness of detection.
Smart Images

Figure CN120448549A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of document processing technology, and more specifically, to a method, system, program product, and device for detecting sensitive words in open-format documents. Background Art
[0002] Open Fixed-layout Document (OFD) is a fixed-layout document format defined in my country. OFD documents can contain multiple types of content.
[0003] Currently, users typically manually edit and input portions of text from OFD documents multiple times on a sensitive word detection website for sensitive word detection, or they convert the OFD document into a streaming document and then input the streaming document into the sensitive word detection website for sensitive word detection. The first method relies on multiple manual edits and inputs, resulting in lower detection speed and accuracy. The second method is prone to problems such as missing content, out-of-order content, and even document corruption in the streaming document, making it difficult to ensure detection effectiveness. Therefore, how to comprehensively and effectively detect sensitive words in OFD documents in a one-stop manner has become a major challenge that needs to be solved urgently. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a method, system, program product and device for detecting sensitive words in open-layout documents, so as to achieve the technical effect of one-stop, comprehensive and effective detection of sensitive words in open-layout documents.
[0005] In a first aspect, an embodiment of the present application provides a method for detecting sensitive words in an open-format document, comprising:
[0006] Parse the uploaded open layout document and obtain the parsed folder;
[0007] Extracting page text information of the open layout document from the parsed folder; wherein the page text information includes one or more of body text information, annotation text information, attached figure text information, watermark text information, seal text information, and signature text information;
[0008] Sensitive words in the page text information are filtered based on a pre-established sensitive word library.
[0009] In the above implementation process, the uploaded open layout document is parsed to obtain a parsed folder, and one or more of the main text information, annotation text information, attached figure text information, watermark text information, seal text information and signature text information of the open layout document are extracted from the parsed folder as page text information. Sensitive words in the page text information are filtered based on a pre-established sensitive word library. By parsing the open layout document once, the text information involved in various types of content in the open layout document can be extracted for sensitive word detection, thereby realizing one-stop, comprehensive and effective detection of sensitive words in OFD documents.
[0010] Furthermore, the open layout document to be uploaded is parsed to obtain a parsed folder, including:
[0011] Obtaining the open layout document uploaded by the user through the browser;
[0012] Decompressing the open format document to obtain multiple parsed files;
[0013] The multiple parsed files are stored in a newly created folder locally in the browser to obtain the parsed folder.
[0014] In the above implementation process, by obtaining the open layout document uploaded by the user through the browser, the multiple parsed files obtained by decompressing the open layout document are stored in a newly created folder locally in the browser to obtain a parsed folder. The open layout document can be parsed directly on the browser, which is conducive to improving the efficiency of sensitive word detection in open layout documents and user convenience.
[0015] Furthermore, the parsed folder includes a page folder, an annotation folder, a resource folder, and a signature folder; the first parsed file in the page folder records the text content of the open layout document, the second parsed file in the annotation folder records the annotation content of the open layout document, the resource folder stores the text drawings, image watermarks, and seal images of the open layout document, and the signature folder stores the signature value file of the open layout document;
[0016] The extracting the page text information of the open layout document from the parsed folder includes:
[0017] Extracting the body text information from the first parsed file;
[0018] Extracting the annotation text information from the second parsed file;
[0019] Identify text information in the accompanying drawings of the main text to obtain text information of the accompanying drawings;
[0020] Identify the text information in the image watermark to obtain the image watermark text information; wherein the watermark text information includes the image watermark text information;
[0021] Recognizing text information in the seal image to obtain the seal text information;
[0022] According to the parsing result of the signature value file, searching for the seal image corresponding to the signature value file, and identifying the text information in the seal image to obtain the signature text information;
[0023] It is determined that the page text information includes the main text information, the annotation text information, the drawing text information, the watermark text information, the seal text information and the signature text information.
[0024] In the above implementation process, when an open layout document contains multiple types of content such as main text, annotations, drawings, picture watermarks, seals and signatures, the open layout document is parsed once to obtain page folders, annotation folders, resource folders and signature folders, the main text information is extracted from the first parsed file in the page folder, the annotation text information is extracted from the second parsed file in the annotation folder, the main text drawings, picture watermarks and seal images are extracted from the resource folder to identify the text information in the main text drawings, picture watermarks and seal images, and obtain the drawing text information, picture watermark text information and seal text information, and according to the parsing result of the signature value file in the signature folder, the signature image corresponding to the signature value file is searched to identify the text information in the signature image, and obtain the signature text information, and finally determine that the page text information includes the main text information, annotation text information, drawing text information, picture watermark text information, seal text information and signature text information, so that when the open layout document contains multiple types of content such as main text, annotations, drawings, picture watermarks, seals and signatures, the text information involved in each type of content in the open layout document can be completely extracted.
[0025] Furthermore, the first parsed file and / or the second parsed file further records a text watermark of the open layout document;
[0026] The extracting the page text information of the open layout document from the parsed folder includes:
[0027] Extracting text watermark information from the first parsed file and / or the second parsed file; wherein the watermark text information includes the text watermark information.
[0028] In the above implementation process, by extracting the text watermark text information from the first parsed file and / or the second parsed file when the open layout document contains a text watermark, and using the text watermark text information as the watermark text information, it is possible to effectively avoid missing the text information related to the text watermark in the open layout document, thereby better realizing one-stop, comprehensive and effective detection of sensitive words in the open layout document.
[0029] Furthermore, the method further comprises:
[0030] storing information belonging to the same page of the open layout document in the body text information in the same text object;
[0031] storing information in the annotation text information belonging to the same page of the open layout document in the same annotation object;
[0032] The images of the text attachment, the image watermark and the seal image that belong to the same page of the open layout document are stored in the same image object; wherein the text object, the annotation object and the image object are structural objects.
[0033] In the above implementation process, by selecting structural objects as text objects, annotation objects and picture objects, the information belonging to the same page of the open layout document in the main text information is stored in the same text object, the information belonging to the same page of the open layout document in the annotation text information is stored in the same annotation object, and the pictures belonging to the same page of the open layout document in the main text attachments, picture watermarks and seal pictures are stored in the same picture object. This enables structured storage of various types of content in the open layout document by page, ensuring orderly storage of data while saving storage memory.
[0034] Furthermore, the attached drawing text information, the picture watermark text information, the seal text information and the signature text information are identified by using an optical character recognition (OCR) recognition tool; wherein the OCR recognition tool is configured according to the language type of the open layout document.
[0035] In the above implementation process, by configuring the OCR recognition tool according to the language type of the open layout document, this OCR recognition tool is used to recognize the text information in the text attachments, picture watermarks, seal pictures and signature pictures, and the attachment text information, picture watermark text information, seal text information and signature text information are obtained. It can flexibly adapt to various language types of open layout documents, and completely and accurately recognize the text information in the picture, which is conducive to improving the recognition efficiency of picture text information.
[0036] Furthermore, before screening the sensitive words in the page text information based on the pre-established sensitive word library, the method further includes:
[0037] Determine whether the page text information has passed user review.
[0038] In the above implementation process, by screening sensitive words in the page text information based on a pre-established sensitive word library when determining that the page text information has passed user review, it is possible to ensure effective and accurate sensitive word detection for open layout documents.
[0039] Furthermore, screening the sensitive words in the page text information based on a pre-established sensitive word library includes:
[0040] A pre-configured sensitive word filtering tool is used to determine a sensitive word in the page text information that matches any sensitive word in the sensitive word library.
[0041] In the above implementation process, by using a pre-configured sensitive word filtering tool, the sensitive words in the page text information that match any sensitive word in the sensitive word library are determined, and the step-by-step matching function of the sensitive word filtering tool can be utilized to completely and accurately filter the sensitive words in the page text information.
[0042] Furthermore, the method further comprises:
[0043] Counting the total number of sensitive words in the text information of the page;
[0044] Assessing the risk level of the open-format document based on the total number of sensitive words;
[0045] If the total number of sensitive words and / or the risk level of the open layout document does not meet the preset file security conditions, the open layout document is refused to be uploaded.
[0046] In the above implementation process, by counting the total number of sensitive words based on the sensitive words in the page text information, evaluating the risk level of the open layout document based on the total number of sensitive words, and refusing to upload the open layout document if the total number of sensitive words and / or the risk level of the open layout document does not meet the pre-set file security conditions, the security of the open layout document can be effectively detected, and uploading of risky open layout documents can be avoided.
[0047] In a second aspect, an embodiment of the present application provides a system for detecting sensitive words in open-format documents, including:
[0048] The parsing module is used to parse the uploaded open format document and obtain the parsed folder;
[0049] an extraction module, configured to extract page text information of the open layout document from the parsed folder; wherein the page text information includes one or more of body text information, annotation text information, attached figure text information, watermark text information, seal text information, and signature text information;
[0050] The detection module is used to filter sensitive words in the page text information based on a pre-established sensitive word library.
[0051] In a third aspect, an embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed by a computer, the computer implements the method described above.
[0052] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; when the processor executes the computer program, the method described above is implemented.
[0053] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0055] Figure 1 A flowchart of a method for detecting sensitive words in an open-format document provided in the first embodiment of the present application;
[0056] Figure 2 A schematic diagram of the structure of a sensitive word detection system for open-format documents provided in the second embodiment of the present application;
[0057] Figure 3 A schematic structural diagram of an electronic device provided in the fourth embodiment of the present application. DETAILED DESCRIPTION
[0058] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0059] It should be noted that in the description of this application, the terms "first" and "second" are used only to distinguish descriptions and should not be understood to indicate or imply relative importance. Furthermore, the step numbers herein are used only to facilitate the explanation of the embodiments of this application and do not limit the order in which the steps are to be executed.
[0060] Open Fixed-layout Document (OFD) is a fixed-layout document format defined in my country. OFD documents can contain various types of content, including text, annotations, drawings, watermarks, seals, and signatures.
[0061] In related technologies, users typically manually edit and input portions of text from OFD documents multiple times on a sensitive word detection website for sensitive word detection, or they first convert the OFD document into a streaming document and then input the streaming document on a sensitive word detection website for sensitive word detection. The first method relies on multiple manual edits and inputs, resulting in low detection speed and accuracy. The second method is prone to problems such as missing content, disordered order, and even document damage in streaming documents, making it difficult to ensure detection effectiveness. Therefore, how to comprehensively and effectively detect sensitive words in OFD documents in a one-stop manner has become a major problem that urgently needs to be solved.
[0062] To this end, the present application proposes a method for detecting sensitive words in open layout documents. The method parses the uploaded open layout document to obtain a parsed folder, extracts one or more of the main text information, annotation text information, drawing text information, watermark text information, seal text information and signature text information of the open layout document from the parsed folder as page text information, and filters the sensitive words in the page text information based on a pre-established sensitive word library. The method can extract text information related to various types of content in the open layout document for sensitive word detection by parsing the open layout document once, thereby realizing one-stop, comprehensive and effective detection of sensitive words in OFD documents.
[0063] The method provided in the embodiment of the present application can be executed by a relevant terminal device, such as a user terminal or a server, and the following description will be given using the server as an example of the execution entity.
[0064] Please see Figure 1 , Figure 1 The first embodiment of the present application provides a method for detecting sensitive words in an open-format document, including steps S101 to S103:
[0065] S101, parsing the open format document to be uploaded to obtain a parsed folder;
[0066] S102: extracting page text information of the open layout document from the parsed folder; wherein the page text information includes one or more of body text information, annotation text information, attached figure text information, watermark text information, seal text information, and signature text information;
[0067] S103: Filter sensitive words in the page text information based on a pre-established sensitive word library.
[0068] As an example, after obtaining the open format document (hereinafter referred to as OFD document) to be uploaded by the user, the server parses the OFD document to obtain a parsed folder, wherein the parsed folder is used to store parsed files recording various contents of the OFD document.
[0069] Extract page text information of the OFD document from the parsed folder, wherein the page text information includes one or more of body text information, annotation text information, drawing text information, watermark text information, seal text information, and signature text information.
[0070] It can be understood that the content of the parsed file record in the parsed folder depends on the content contained in the OFD document, and the type of information included in the extracted page text information depends on the content type of the parsed file record in the parsed folder. For example, if the OFD document contains a main text, the parsing folder contains a parsing file recording the main text of the OFD document, and the extracted page text information includes the text information in the main text, that is, the main text text information; if the OFD document contains annotations, the parsing folder contains a parsing file recording the annotations of the OFD document, and the extracted page text information includes the text information in the annotations, that is, the annotation text information; if the OFD document contains drawings, the parsing folder contains a parsing file recording the drawings of the OFD document, and the extracted page text information includes the text information in the drawings, that is, the drawing text information; if the OFD document contains a watermark, the parsing folder contains a parsing file recording the watermark of the OFD document, and the extracted page text information includes the text information in the watermark, that is, the watermark text information; if the OFD document contains a seal, the parsing folder contains a parsing file recording the seal image of the OFD document, and the extracted page text information includes the text information in the seal image, that is, the seal text information; if the OFD document contains a signature, the parsing folder contains a parsing file recording the signature-related file of the OFD document, and the extracted page text information includes the text information in the signature image corresponding to the signature-related file, that is, the signature text information.
[0071] Based on the actual business scenario, a sensitive word library is pre-established. After obtaining the page text information, sensitive words in the page text information are filtered based on the pre-established sensitive word library to complete the sensitive word detection of the OFD document.
[0072] In actual applications, the pre-established sensitive word library may include the sensitive word library provided by the Ansj Chinese word segmentation library and the sensitive word library built by the user.
[0073] The embodiment of the present application parses the uploaded open layout document to obtain a parsed folder, extracts one or more of the main text information, annotation text information, drawing text information, watermark text information, seal text information and signature text information of the open layout document from the parsed folder as page text information, and filters sensitive words in the page text information based on a pre-established sensitive word library. By parsing the open layout document once, the text information related to various types of content in the open layout document can be extracted for sensitive word detection, thereby realizing one-stop, comprehensive and effective detection of sensitive words in OFD documents.
[0074] In an optional embodiment, parsing the open layout document to be uploaded to obtain a parsed folder includes: obtaining the open layout document uploaded by the user through the browser; decompressing the open layout document to obtain multiple parsed files; storing the multiple parsed files in a newly created folder locally in the browser to obtain a parsed folder.
[0075] As an example, in actual application, a user may upload an OFD document through a browser.
[0076] The server obtains the OFD document uploaded by the user through the browser, decompresses the OFD document according to the organizational structure of the OFD document, obtains multiple parsed files, and requests permission to operate the browser's local folder, creates a new general folder locally in the browser, stores multiple parsed files in this new folder, and obtains a parsed folder.
[0077] The embodiment of the present application obtains an open layout document uploaded by a user through a browser, stores multiple parsed files obtained by decompressing the open layout document in a newly created folder locally in the browser, and obtains a parsed folder. This can directly complete the parsing of the open layout document once on the browser, which is beneficial to improving the efficiency of sensitive word detection in open layout documents and user convenience.
[0078] In an optional embodiment, the parsing folder includes a page folder, an annotation folder, a resource folder and a signature folder; the first parsing file in the page folder records the main text content of the open layout document, the second parsing file in the annotation folder records the annotation content of the open layout document, the resource folder stores the main text drawings, picture watermarks and seal pictures of the open layout document, and the signature folder stores the signature value file of the open layout document; the extracting of page text information of the open layout document from the parsing folder includes: extracting main text information from the first parsing file; extracting annotation text information from the second parsing file; identifying text information in the main text drawings to obtain drawing text information; identifying text information in picture watermarks to obtain picture watermark text information; wherein the watermark text information includes picture watermark text information; identifying text information in the seal picture to obtain seal text information; according to the parsing result of the signature value file, searching for the signature picture corresponding to the signature value file, and identifying the text information in the signature picture to obtain signature text information; determining that the page text information includes main text information, annotation text information, drawing text information, watermark text information, seal text information and signature text information.
[0079] As an example, when an OFD document contains multiple types of content such as text, annotations, drawings, picture watermarks, seals and signatures, the server parses the OFD document after obtaining the OFD document, and the resulting parsed folders include a page folder, an annotation folder, a resource folder and a signature folder, wherein the first parsed file in the page folder records the text content of the OFD document, the second parsed file in the annotation folder records the annotation content of the OFD document, the resource folder stores the text drawings, picture watermarks and seal pictures of the OFD document, and the signature folder stores the signature value file of the OFD document.
[0080] It should be noted that the signature is usually an electronic signature produced according to a standard. Due to the special nature of the signature, the parsed signature-related file is not a simple signature image, but a signature text with a seal structure that combines electronic signature technology, that is, a signature value file.
[0081] For example, the parsed folders obtained by the server include the Pages folder, the Annots folder, the Res folder and the Signs folder, among which the Pages folder is a page folder, the Content.xml file in the Pages folder is the first parsed file, and the Content.xml file records the text content of each page of the OFD document, etc.; the Annots folder is an annotation folder, the Annotation.xml file in the Annots folder is the second parsed file, and the Annotation.xml file records the annotation content of each page of the OFD document, etc.; the Res folder is a resource folder, and the Res folder stores files of resources related to pictures and fonts of each page of the OFD document, such as picture files and font files, etc.; the Signs folder is a signature folder, which stores signature value files corresponding to each signature of the OFD document, such as the SignedValue.dat signature value file.
[0082] After obtaining the parsed folder, the server executes: extracting the main text information from the first parsed file; extracting the annotation text information from the second parsed file; extracting the main text drawings from the relevant files of the resource folder, identifying the text information in the main text drawings, and obtaining the drawing text information; extracting the image watermark from the relevant files of the resource folder, identifying the text information in the image watermark, and obtaining the image watermark text information, so as to determine the image watermark text information as the watermark text information; extracting the seal image from the relevant files of the resource folder, identifying the text information in the seal image, and obtaining the seal text information; for each signature value file in the signature folder, parsing the signature value file, searching for the signature image corresponding to the signature value file according to the parsing result of the signature value file, identifying the text information in the signature image, and obtaining the corresponding signature text information; determining that the page text information includes the acquired main text information, annotation text information, drawing text information, watermark text information, seal text information and signature text information.
[0083] In actual applications, each of all images in an OFD document, including text attachments, image watermarks, and seal images, can be stored as a separate image file, such as picture1.jpg file and picture2.png file. All images in an OFD document, including text attachments, image watermarks, and seal images, can also be recorded in the same file.
[0084] In actual applications, the main text attachments, picture watermarks and seal pictures can be extracted from the relevant files of the resource folder in turn, and image extraction and text recognition operations can be performed on the main text attachments, picture watermarks and seal pictures in turn. The main text attachments, picture watermarks and seal pictures can also be extracted from the relevant files of the resource folder at one time, and text recognition operations can be performed on the main text attachments, picture watermarks and seal pictures in a unified manner.
[0085] The embodiment of the present application parses the open layout document once to obtain a page folder, an annotation folder, a resource folder, and a signature folder, extracts the main text information from the first parsed file in the page folder, extracts the annotation text information from the second parsed file in the annotation folder, extracts the main text drawings, image watermarks, and seal images from the resource folder to identify the text information in the main text drawings, image watermarks, and seal images, and obtains the drawing text information, image watermark text information, and seal text information, and searches for the signature image corresponding to the signature value file according to the parsing result of the signature value file in the signature folder to identify the text information in the signature image, and obtain the signature text information, and finally determines that the page text information includes the main text information, annotation text information, drawing text information, image watermark text information, seal text information, and signature text information. In this case, when the open layout document contains multiple types of content such as the main text, annotations, drawings, image watermarks, seals, and signatures, the text information related to each type of content in the open layout document can be completely extracted.
[0086] In an optional embodiment, the first parsed file and / or the second parsed file also records the text watermark of the open layout document; extracting the page text information of the open layout document from the parsed folder includes: extracting the text watermark text information from the first parsed file and / or the second parsed file; wherein the watermark text information includes text watermark text information.
[0087] As an example, watermarks are mainly divided into two types: image watermarks and text watermarks. In actual applications, OFD documents may also include text watermarks.
[0088] In the case that the OFD document contains a text watermark, the text watermark extracted by parsing the OFD document once may be recorded in the first parsed file and / or the second parsed file.
[0089] After obtaining the parsed folder, the server extracts the text information in the text watermark from the first parsed file, i.e., the text watermark text information, and / or extracts the text watermark text information from the second parsed file, and uses the text watermark text information as the watermark text information.
[0090] It should be noted that if the text watermark extracted by parsing the OFD document once is recorded in the first parsed file, the server extracts the text information of the text watermark from the first parsed file; if the text watermark extracted by parsing the OFD document once is recorded in the second parsed file, the server extracts the text information of the text watermark from the second parsed file; if the text watermark extracted by parsing the OFD document once is recorded in the first parsed file and the second parsed file, the server extracts the text information of the text watermark from the first parsed file and the second parsed file respectively.
[0091] The embodiment of the present application extracts text watermark text information from the first parsed file and / or the second parsed file when the open layout document contains a text watermark, and uses the text watermark text information as the watermark text information, thereby effectively avoiding missing the text information related to the text watermark in the open layout document, thereby better realizing one-stop, comprehensive and effective detection of sensitive words in the open layout document.
[0092] In an optional embodiment, the method further includes steps S104 to S106:
[0093] S104, storing the information belonging to the same page of the open layout document in the body text information in the same text object;
[0094] S105, storing the information belonging to the same page of the open layout document in the annotation text information in the same annotation object;
[0095] S106. Store the images of the text attachment, image watermark, and seal image that belong to the same page of the open layout document in the same image object; wherein the text object, annotation object, and image object are structural objects.
[0096] As an example, after the server extracts the main text information, the main text information includes the text information in the main content of each page of the OFD document, and stores the information in the main text information belonging to the same page of the OFD document and the page number of the current page in the same text object; after extracting the annotation text information, the annotation text information includes the text information in the annotation content of each page of the OFD document, and stores the information in the annotation text information belonging to the same page of the OFD document and the page number of the current page in the same annotation object; after extracting the pictures of each page of the OFD document, namely the main text drawings, picture watermarks and seal pictures, the pictures in the main text drawings, picture watermarks and seal pictures belonging to the same page of the OFD document are stored in the same picture object, wherein the text object, annotation object and picture object are structured objects, such as JSON objects.
[0097] The embodiment of the present application selects structural objects as text objects, annotation objects and picture objects, stores information belonging to the same page of an open layout document in the main text information in the same text object, stores information belonging to the same page of an open layout document in the annotation text information in the same annotation object, and stores pictures belonging to the same page of an open layout document in the main text illustrations, picture watermarks and seal pictures in the same picture object. It can perform structured storage of various types of content in an open layout document by page, ensure orderly storage of data, and save storage memory.
[0098] In an optional embodiment, the text information of the attached drawings, the text information of the picture watermark, the text information of the seal and the text information of the signature are recognized by an optical character recognition (OCR) recognition tool; wherein the OCR recognition tool is configured according to the language type of the open layout document.
[0099] As an example, the OCR recognition tool is configured according to the language type of the OFD document.
[0100] In actual application, if the language type of the OFD document is a single language, such as Chinese, English or other languages, the OCR recognition tool is directly configured as a tool dedicated to recognizing text in this language, so that this OCR recognition tool can be used to recognize the text in the picture; if the language type of the OFD document is multiple languages, such as Chinese and English, two OCR recognition tools can be configured, one OCR recognition tool can be configured as a tool dedicated to recognizing Chinese text, and the other OCR recognition tool can be configured as a tool dedicated to recognizing English text, so that these two OCR recognition tools can be used to recognize the text in the picture respectively. Alternatively, the OCR recognition tool can be first configured as a tool dedicated to recognizing Chinese text, and the current OCR recognition tool can be used to recognize the text in the picture, and then the OCR recognition tool can be configured as a tool dedicated to recognizing English text, and the current OCR recognition tool can be used to recognize the text in the picture.
[0101] After obtaining the main text and attached drawings, the server converts the format of the main text and attached drawings into base64 encoding format, and uses an OCR recognition tool to recognize the text information in the main text and attached drawings to obtain the text information of the attached drawings.
[0102] After obtaining the image watermark, the server converts the image watermark format into base64 encoding format, and uses OCR recognition tools to recognize the text information in the image watermark to obtain the image watermark text information.
[0103] After obtaining the seal image, the server converts the seal image format into base64 encoding format, and uses OCR recognition tools to recognize the text information in the seal image to obtain the seal text information.
[0104] After obtaining the signature image, the server converts the signature image format into base64 encoding format and uses OCR recognition tools to recognize the text information in the signature image to obtain the signature text information.
[0105] In actual applications, image extraction and text recognition operations can be performed separately for the text attachments, image watermarks, seal images and signature images, or after obtaining the text attachments, image watermarks, seal images and signature images, text recognition operations can be performed uniformly on the text attachments, image watermarks, seal images and signature images.
[0106] Base64 is one of the most common encoding methods for transmitting 8-bit bytecode on the Internet. By first converting the image format into base64 encoding format and then using OCR recognition tools for text recognition, the OCR recognition tool can directly read the image, which is conducive to improving the recognition speed of image text information.
[0107] The embodiment of the present application configures an OCR recognition tool according to the language type of the open layout document, and uses this OCR recognition tool to recognize the text information in the text drawings, picture watermarks, seal pictures and signature pictures, thereby obtaining the text information of the drawings, picture watermarks, seals and signatures. It can flexibly adapt to various language types of open layout documents, completely and accurately recognize the text information in the pictures, and is conducive to improving the recognition efficiency of picture text information.
[0108] In an optional embodiment, before filtering the sensitive words in the page text information based on the pre-established sensitive word library, the method further includes: determining whether the page text information has passed user review.
[0109] As an example, after obtaining the page text information, the server returns the page text information to the user, so that the user can review whether the page text information has any errors or omissions and send the review result to the server.
[0110] If the server determines that the page text information has passed the user review, it considers that the page text information is correct and continues to filter sensitive words in the page text information based on the pre-established sensitive word library; if the server determines that the page text information has not passed the user review, it considers that the page text information has errors and omissions. At this time, the above-mentioned parsing and extraction operations can be performed again on the OFD document, or the page text information corrected by the user can be obtained, and the sensitive words in the corrected page text information can be filtered based on the pre-established sensitive word library.
[0111] The embodiment of the present application can ensure effective and accurate sensitive word detection for open-format documents by screening sensitive words in page text information based on a pre-established sensitive word library when determining that the page text information has passed user review.
[0112] In an optional embodiment, screening sensitive words in the page text information based on a pre-established sensitive word library includes: using a pre-configured sensitive word filtering tool to determine a sensitive word in the page text information that matches any sensitive word in the sensitive word library.
[0113] As an example, after obtaining the page text information, the server uses a pre-configured sensitive word filtering tool to match the page text information with each sensitive word in the sensitive word library step by step, and determines the sensitive word in the page text information that matches any sensitive word in the sensitive word library.
[0114] For example, assuming that there is "xyz" in the text information of the page, the sensitive word filtering tool will first match "x", "y" and "z" with the sensitive words in the sensitive word library respectively, then match "xy" and "yz" with the sensitive words in the sensitive word library respectively, and finally match "xyz" with the sensitive words in the sensitive word library respectively, thereby filtering out the sensitive words in "xyz".
[0115] In practical applications, sensitive word filtering tools can be configured to perform text matching based on the DFA (Deterministic Finite Automaton) algorithm. The DFA algorithm is an efficient algorithm for string matching. It considers optimal transition paths, reduces unnecessary backtracking, and is suitable for large texts. It is also fast and relatively simple to implement. Using the DFA algorithm for text matching can improve the efficiency of sensitive word screening.
[0116] The embodiment of the present application uses a pre-configured sensitive word filtering tool to determine the sensitive words in the page text information that match any sensitive word in the sensitive word library, and can use the step-by-step matching function of the sensitive word filtering tool to completely and accurately filter the sensitive words in the page text information.
[0117] In an optional embodiment, the method further includes steps S107 to S109:
[0118] S107: Count the total number of sensitive words in the page text information;
[0119] S108. Evaluate the risk level of the open-format document based on the total number of sensitive words;
[0120] S109: If the total number of sensitive words and / or the risk level of the open-layout document does not meet the pre-set file security conditions, the open-layout document is refused to be uploaded.
[0121] As an example, after obtaining the sensitive words in the page text information, the server counts the total number of sensitive words according to the sensitive words in the page text information, and evaluates the risk level of the OFD document according to the total number of sensitive words.
[0122] In practical applications, a mapping relationship table between the total number of sensitive words and multiple risk levels can be pre-configured according to actual business needs, such as Table 1.
[0123] Table 1
[0124]
[0125]
[0126] The pre-configured mapping relationship table is searched to determine the risk level corresponding to the total number of sensitive words counted, and this risk level is determined as the risk level of the OFD document.
[0127] After determining the risk level of the OFD document, the server determines whether the total number of sensitive words counted and / or the risk level of the OFD document meet the pre-set file security conditions. If so, the OFD document is allowed to be uploaded; if not, the OFD document is rejected from being uploaded.
[0128] It can be understood that the file security conditions are set for the total number of sensitive words and / or the risk level of the OFD document. That is, the file security conditions can be set for the total number of sensitive words, the risk level of the OFD document, or the total number of sensitive words and the risk level of the OFD document.
[0129] When the file security condition is set for the total number of sensitive words, the server determines whether the total number of sensitive words counted meets the file security condition, such as whether the total number of sensitive words counted is greater than the preset threshold for the total number of sensitive words; when the file security condition is set for the risk level of the OFD document, the server determines whether the risk level of the OFD document meets the file security condition, such as whether the risk level of the OFD document is higher than the preset risk level threshold; when the file security condition is set for the total number of sensitive words and the risk level of the OFD document, the server determines whether the total number of sensitive words counted and the risk level of the OFD document meet the file security condition, such as whether the total number of sensitive words counted is greater than the preset threshold for the total number of sensitive words, and whether the risk level of the OFD document is higher than the preset risk level threshold.
[0130] The embodiment of the present application counts the total number of sensitive words based on the sensitive words in the page text information, evaluates the risk level of the open layout document based on the total number of sensitive words, and refuses to upload the open layout document if the total number of sensitive words and / or the risk level of the open layout document does not meet the pre-set file security conditions. This can effectively detect the security of the open layout document and avoid uploading risky open layout documents.
[0131] Please see Figure 2 , Figure 2 A schematic diagram of the structure of a multi-database driven self-adaptive system provided in the second embodiment of the present application. The second embodiment of the present application provides a sensitive word detection system for open-layout documents, comprising: a parsing module 201 for parsing an uploaded open-layout document to obtain a parsed folder; an extraction module 202 for extracting page text information of the open-layout document from the parsed folder; wherein the page text information includes one or more of body text information, annotation text information, illustration text information, watermark text information, seal text information, and signature text information; and a detection module 203 for screening sensitive words in the page text information based on a pre-established sensitive word library.
[0132] In an optional embodiment, parsing the open layout document to be uploaded to obtain a parsed folder includes: obtaining the open layout document uploaded by the user through the browser; decompressing the open layout document to obtain multiple parsed files; storing the multiple parsed files in a newly created folder locally in the browser to obtain a parsed folder.
[0133] In an optional embodiment, the parsed folder includes a page folder, an annotation folder, a resource folder, and a signature folder; the first parsed file in the page folder records the text content of the open layout document, the second parsed file in the annotation folder records the annotation content of the open layout document, the resource folder stores the text drawings, image watermarks, and seal images of the open layout document, and the signature folder stores the signature value file of the open layout document;
[0134] The method of extracting page text information of an open layout document from a parsed folder includes: extracting main text information from a first parsed file; extracting annotation text information from a second parsed file; identifying text information in main text drawings to obtain drawing text information; identifying text information in image watermarks to obtain image watermark text information; wherein the watermark text information includes image watermark text information; identifying text information in a seal image to obtain seal text information; searching for a signature image corresponding to the signature value file according to a parsed result of the signature value file, and identifying text information in the signature image to obtain signature text information; and determining that the page text information includes main text information, annotation text information, drawing text information, watermark text information, seal text information, and signature text information.
[0135] In an optional embodiment, the first parsed file and / or the second parsed file also records the text watermark of the open layout document; extracting the page text information of the open layout document from the parsed folder includes: extracting the text watermark text information from the first parsed file and / or the second parsed file; wherein the watermark text information includes text watermark text information.
[0136] In an optional embodiment, the extraction module 202 is further used to: store information in the main text information that belongs to the same page of the open layout document in the same text object; store information in the annotation text information that belongs to the same page of the open layout document in the same annotation object; store images in the main text drawings, image watermarks and seal images that belong to the same page of the open layout document in the same image object; wherein the text object, annotation object and image object are structural objects.
[0137] In an optional embodiment, the text information of the attached drawings, the text information of the picture watermark, the text information of the seal and the text information of the signature are recognized by an optical character recognition (OCR) recognition tool; wherein the OCR recognition tool is configured according to the language type of the open layout document.
[0138] In an optional embodiment, the detection module 203 is further configured to determine whether the page text information has passed user review before filtering the sensitive words in the page text information based on the pre-established sensitive word library.
[0139] In an optional embodiment, screening sensitive words in the page text information based on a pre-established sensitive word library includes: using a pre-configured sensitive word filtering tool to determine a sensitive word in the page text information that matches any sensitive word in the sensitive word library.
[0140] In an optional embodiment, the detection module 203 is further used to: count the total number of sensitive words based on the sensitive words in the page text information; evaluate the risk level of the open layout document based on the total number of sensitive words; and refuse to upload the open layout document if the total number of sensitive words and / or the risk level of the open layout document does not meet the pre-set file security conditions.
[0141] The implementation process of the functions and effects of each module in the above system is specifically described in the implementation process of the corresponding steps in the above method, which will not be repeated here.
[0142] The third embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed by a computer, the computer implements the method described in the first embodiment of the present application and can achieve the same beneficial effects.
[0143] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in accordance with the fourth embodiment of the present application. The fourth embodiment of the present application provides an electronic device 30, comprising a processor 301, a memory 302, and a computer program stored in the memory 302 and configured to be executed by the processor 301. The memory 302 is coupled to the processor 301, and when the processor 301 executes the computer program, the method described in accordance with the first embodiment of the present application is implemented, achieving the same beneficial effects as described in accordance with the first embodiment.
[0144] In which, when the processor 301 reads the computer program from the memory 302 through the bus 303 and executes the computer program, it can implement the method of any embodiment included in the method described in the first embodiment of the present application.
[0145] Processor 301 can process digital signals and can include various computing architectures, such as a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements a combination of multiple instruction sets. In some examples, processor 301 can be a microprocessor.
[0146] The memory 302 can be used to store instructions executed by the processor 301 or data related to the execution of instructions. These instructions and / or data may include code for implementing some or all functions of one or more modules described in the embodiments of this application. The processor 301 of the embodiment of the present disclosure can be used to execute the instructions in the memory 302 to implement the method described in the first embodiment of this application. The memory 302 includes dynamic random access memory, static random access memory, flash memory, optical storage, or other memory known to those skilled in the art.
[0147] The fifth embodiment of the present application provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method described in the first embodiment of the present application, and can achieve the same beneficial effects as the method.
[0148] The method described in the first embodiment of the present application can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in each embodiment of the present application are executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, a core network device, an OAM (Open Application Model), or other programmable device.
[0149] The computer program or instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired or wireless method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; an optical medium, such as a digital video disk; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.
[0150] In summary, embodiments of the present application provide a method, system, program product, and device for detecting sensitive words in open-layout documents. The method comprises: parsing an uploaded open-layout document to obtain a parsed folder; extracting page text information of the open-layout document from the parsed folder; wherein the page text information includes one or more of main text information, annotation text information, illustration text information, watermark text information, seal text information, and signature text information; and screening sensitive words in the page text information based on a pre-established sensitive word library. Embodiments of the present application parse an uploaded open-layout document to obtain a parsed folder, extracting one or more of the main text information, annotation text information, illustration text information, watermark text information, seal text information, and signature text information of the open-layout document from the parsed folder as page text information, and screening sensitive words in the page text information based on a pre-established sensitive word library. This method can extract text information related to various types of content in the open-layout document for sensitive word detection by parsing the open-layout document once, thereby achieving one-stop, comprehensive, and effective detection of sensitive words in OFD documents.
[0151] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0152] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0153] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0154] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for detecting sensitive words in open-format documents, characterized in that: include: Parse the uploaded open layout document and obtain the parsed folder; Extracting page text information of the open layout document from the parsed folder; wherein the page text information includes one or more of body text information, annotation text information, attached figure text information, watermark text information, seal text information, and signature text information; Sensitive words in the page text information are filtered based on a pre-established sensitive word library.
2. The method according to claim 1, characterized in that The open layout document to be uploaded is parsed to obtain a parsed folder, including: Obtaining the open layout document uploaded by the user through the browser; Decompressing the open format document to obtain multiple parsed files; The multiple parsed files are stored in a newly created folder locally in the browser to obtain the parsed folder.
3. The method according to claim 1, characterized in that The parsed folder includes a page folder, an annotation folder, a resource folder, and a signature folder; the first parsed file in the page folder records the text content of the open layout document, the second parsed file in the annotation folder records the annotation content of the open layout document, the resource folder stores the text drawings, picture watermarks, and seal pictures of the open layout document, and the signature folder stores the signature value file of the open layout document; The extracting the page text information of the open layout document from the parsed folder includes: Extracting the body text information from the first parsed file; Extracting the annotation text information from the second parsed file; Identify text information in the accompanying drawings of the main text to obtain text information of the accompanying drawings; Identify the text information in the image watermark to obtain the image watermark text information; wherein the watermark text information includes the image watermark text information; Recognizing text information in the seal image to obtain the seal text information; According to the parsing result of the signature value file, searching for the seal image corresponding to the signature value file, and identifying the text information in the seal image to obtain the signature text information; It is determined that the page text information includes the main text information, the annotation text information, the drawing text information, the watermark text information, the seal text information and the signature text information.
4. The method according to claim 3, characterized in that The first parsed file and / or the second parsed file further records a text watermark of the open layout document; The extracting the page text information of the open layout document from the parsed folder includes: Extracting text watermark information from the first parsed file and / or the second parsed file; wherein the watermark text information includes the text watermark information.
5. The method according to claim 3, characterized in that The method further comprises: storing information belonging to the same page of the open layout document in the body text information in the same text object; storing information in the annotation text information belonging to the same page of the open layout document in the same annotation object; The images of the text attachment, the image watermark and the seal image that belong to the same page of the open layout document are stored in the same image object; wherein the text object, the annotation object and the image object are structural objects.
6. The method according to claim 3, characterized in that The attached drawing text information, the picture watermark text information, the seal text information and the signature text information are identified by using an optical character recognition (OCR) recognition tool; wherein the OCR recognition tool is configured according to the language type of the open layout document.
7. The method according to claim 1, characterized in that Before screening the sensitive words in the page text information based on the pre-established sensitive word library, the method further includes: Determine whether the page text information has passed user review.
8. The method according to claim 1, characterized in that The screening of sensitive words in the page text information based on a pre-established sensitive word library includes: A pre-configured sensitive word filtering tool is used to determine a sensitive word in the page text information that matches any sensitive word in the sensitive word library.
9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: Counting the total number of sensitive words in the text information of the page; Assessing the risk level of the open-format document based on the total number of sensitive words; If the total number of sensitive words and / or the risk level of the open layout document does not meet the preset file security conditions, the open layout document is refused to be uploaded.
10. A sensitive word detection system for open format documents, characterized in that: include: The parsing module is used to parse the uploaded open format document and obtain the parsed folder; an extraction module, configured to extract page text information of the open layout document from the parsed folder; wherein the page text information includes one or more of body text information, annotation text information, attached figure text information, watermark text information, seal text information, and signature text information; The detection module is used to filter sensitive words in the page text information based on a pre-established sensitive word library.
11. A computer program product, characterized in that The computer program product comprises instructions which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 8.
12. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; when the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.