Filtering Method, Filtering Device and Processor for Document Page Content

By converting the PDF document into a target picture, determining the area to be blocked and generating a replacement area, the problem of difficulty in intelligently blocking the content of PDF document in the prior art is solved, and an efficient content blocking effect is achieved.

CN114036900BActive Publication Date: 2025-06-13ZHUHAI KINGSOFT OFFICE SOFTWARE +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111223525.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-20
Publication Date
2025-06-13
Estimated Expiration
2041-10-20

AI Technical Summary

Technical Problem

The prior art is difficult to intelligently block PDF document content, resulting in the inability to effectively block unwanted content during automatic analysis or manual reading.

Method used

By converting the target document page into a target picture, the location information of the area to be blocked is determined, and a replacement area matching the size and shape is generated to replace the area to be blocked. This method uses neural network model and target index identification information to accurately locate and generate replacement areas.

Benefits of technology

Intelligent blocking of document content is realized, unnecessary content can be blocked selectively and efficiently, and the problem of difficulty in realizing intelligent blocking in the prior art is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114036900B_ABST
    Figure CN114036900B_ABST
Patent Text Reader

Abstract

The present application provides a method for filtering document page content, a filtering device and a processor. The document page includes multiple pages. The filtering method includes: obtaining a target image, where the target image is an image obtained by converting a target document page, and the target document page is a to-be-blocked area including content to be blocked; determining the position information of the to-be-blocked area according to the target image; generating a target replacement area at least according to the position information of the to-be-blocked area, and using the target replacement area to replace the to-be-blocked area, where the size and shape of the target replacement area match the size and shape of the to-be-blocked area. In this solution, generating a target replacement area according to the position information of the to-be-blocked area and using the generated target replacement area to replace the to-be-blocked area realizes blocking the content that needs to be blocked, thus solving the problem that it is difficult to intelligently block document content in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of filtering of document pages. Specifically, it relates to a method for filtering document page content, a filtering device, a computer-readable storage medium, and a processor. Background Art

[0002] The Portable Document Format (PDF for short). This file format is independent of the operating system platform. That is to say, PDF files are common whether in the Windows, Unix, or Apple's Mac OS operating system. This feature makes it an ideal document format for electronic document distribution and digital information dissemination on the Internet. More and more e-books, product descriptions, company announcements, online materials, and emails are starting to use PDF format documents.

[0003] The content components of PDF documents are complex, including annotations, audio, video, links, pictures, controls (text boxes, buttons, check boxes, drop-down lists, etc.). When automatically parsing or manually reading the content of PDF documents, in some scenarios, it is necessary to block some document content.

[0004] Typical scenarios of automatic parsing include PDF document translation. If a PDF document contains annotations, illustrations, and controls, it is not easy to determine the text range to be translated, and it is easy to cause layout problems in the translation result due to translating non-body content; typical scenarios of manual reading include document review. For example, when manually reviewing a batch of PDF documents, these documents all contain sensitive content such as ID card photos, and it is hoped to conduct a review after targeted blocking.

[0005] In the current scenario of manual reading of PDF documents, generally, the need for document content shielding is bypassed through confidentiality agreements and personnel permission control. The need for document content shielding has not been truly solved, which requires high personnel quality and has relatively high risks.

[0006] Therefore, there is an urgent need for a method that can intelligently shield document content.

[0007] The above information disclosed in the background art section is only used to enhance the understanding of the background art of the technology described in this article. Therefore, the background art may contain certain information that is not prior art known to those skilled in the art in this country. Summary of the Invention

[0008] The main purpose of this application is to provide a method for filtering document page content, a filtering device, a computer-readable storage medium, and a processor to solve the problem that it is difficult to intelligently shield document content in the prior art.

[0009] According to one aspect of an embodiment of the present invention, a method for filtering document page content is provided. The document page includes multiple pages, and the method includes: obtaining a target picture, where the target picture is a picture converted from a target document page, and the target document page includes a to-be-blocked area with to-be-blocked content; determining position information of the to-be-blocked area according to the target picture; generating a target replacement area at least according to the position information of the to-be-blocked area, and replacing the to-be-blocked area with the target replacement area, where the size and shape of the target replacement area match the size and shape of the to-be-blocked area.

[0010] Optionally, determining the position information of the to-be-blocked area according to the target picture includes: determining position information of a target image area corresponding to the to-be-blocked area in the target picture according to the target picture and a neural network model. The neural network model is obtained by using multiple sets of training data, and the multiple sets of training data include: historical pictures and corresponding historical position information, where the historical position information is the position information corresponding to the to-be-blocked area in the historical pictures; determining the position information of the to-be-blocked area according to the position information of the target image area, the size information of the target picture, and the size information of the target document page.

[0011] Optionally, the target image area is a rectangular area, the to-be-blocked area is a rectangular area, and the position information of the target image area includes coordinate information of a first position point and coordinate information of a second position point, and the first position point and the second position point are located on the same diagonal of the rectangular area.

[0012] Optionally, determining the position information of the to-be-blocked area according to the position information of the target image area, the size information of the target picture, and the size information of the target document page includes: determining a height ratio according to the height of the target picture and the height of the target document page; determining a width ratio according to the width of the target picture and the width of the target document page; adjusting the coordinate information of the first position point according to the height ratio and the width ratio to obtain coordinate information of a third position point; adjusting the coordinate information of the second position point according to the height ratio and the width ratio to obtain coordinate information of a fourth position point, and the coordinate information of the third position point and the coordinate information of the fourth position point constitute the position information of the to-be-blocked area.

[0013] Optionally, the area in the target picture corresponding to the area to be masked is the target image area. Generating a target replacement area based at least on the position information of the area to be masked includes: obtaining target index identification information, where the target index identification information includes first identification information and second identification information. The first identification information is the unique identification information of the target image area in the target picture, and the second identification information is the unique identification information of the target document; generating the target replacement area carrying the target index identification information according to the target index identification information and the position information of the area to be masked.

[0014] Optionally, obtaining target index identification information includes: obtaining the hash value of the target document page as the first identification information; obtaining the serial number of the target image area in the target picture as the second identification information.

[0015] Optionally, obtaining target index identification information further includes: obtaining the current date and index extension information, where the current date is the date when the target document page needs to be filtered currently, and the index extension information is other identification information characterizing the target document page.

[0016] Optionally, generating the target replacement area carrying the target index identification information according to the target index identification information and the position information of the area to be masked includes: generating an initial replacement area according to the target index identification information; generating the target replacement image area according to the position information of the area to be masked and the initial image area.

[0017] Optionally, after determining the position information of the area to be masked according to the target picture, the method further includes: obtaining the target replacement area corresponding to the area to be masked in the target picture and storing it in a memory.

[0018] Optionally, after generating the target replacement area based at least on the position information of the area to be masked and replacing the area to be masked with the target replacement area, the method further includes: in response to an operation of scanning the target replacement area, determining the corresponding area to be masked according to the corresponding target index identification information, and replacing the target replacement area in the target document with the corresponding area to be masked.

[0019] According to another aspect of the embodiments of the present invention, there is also provided a filtering device for document page content. The document page includes multiple pages, and the device includes: a first acquisition unit configured to acquire a target picture, where the target picture is a picture converted from a target document page, and the target document page includes a to-be-blocked area with to-be-blocked content; a determination unit configured to determine the position information of the to-be-blocked area according to the target picture; and a replacement unit configured to generate a target replacement area at least according to the position information of the to-be-blocked area, and replace the to-be-blocked area with the target replacement area, where the size and shape of the target replacement area match the size and shape of the to-be-blocked area.

[0020] According to still another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium. The computer-readable storage medium includes a stored program, where the program executes any one of the above filtering methods.

[0021] According to yet another aspect of the embodiments of the present invention, there is also provided a processor. The processor is configured to run a program, where when the program runs, it executes any one of the above filtering methods.

[0022] In the embodiments of the present invention, in the filtering method for document page content, first, the target document page is converted to obtain a target picture, then the position information of the to-be-blocked area is determined according to the target picture, and finally, a target replacement area is generated at least according to the position information of the to-be-blocked area, and the generated target replacement area is used to replace the to-be-blocked area. Here, the target document page is a page including a to-be-blocked area with to-be-blocked content, and the size and shape of the target replacement area match the size and shape of the to-be-blocked area. In this solution, the position information of the to-be-blocked area is determined according to the target picture, which ensures that the content to be blocked can be selectively and intelligently selected. Subsequently, a target replacement area is generated according to the position information of the to-be-blocked area, and the generated target replacement area is used to replace the to-be-blocked area, that is, the content to be blocked is blocked, thus solving the problem in the prior art that it is difficult to intelligently block document content. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The schematic diagrams in the specification drawings that form a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0024] Figure 1 Shows a schematic diagram of a filtering method for document page content according to an embodiment of this application;

[0025] Figure 2 Shows a schematic diagram of a to-be-blocked area according to an embodiment of this application;

[0026] Figure 3 Shows a schematic diagram of replacing an area to be masked according to an embodiment of the present application;

[0027] Figure 4 Shows a schematic diagram of parsing a document page according to an embodiment of the present application;

[0028] Figure 5 Shows a schematic diagram of replacing a target image area according to an embodiment of the present application;

[0029] Figure 6 Shows a schematic diagram of a filtering device for document page content according to an embodiment of the present application.

[0030] Among them, the above-mentioned drawings include the following reference numerals:

[0031] 100, area to be masked; 101, initial replacement area; 102, target replacement area. Detailed implementation manners

[0032] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0033] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to describe the embodiments of the present application here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.

[0035] As described in the background art, it is difficult to intelligently mask the content of a document in the prior art. To solve the above problems, in a typical implementation manner of the present application, a method for filtering the content of a document page, a filtering device, a computer-readable storage medium, and a processor are provided.

[0036] According to an embodiment of the present application, a method for filtering the content of a document page is provided.

[0037] Figure 1 It is a flowchart of the method for filtering the content of a document page according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0038] Step S101, obtain a target picture, where the target picture is a picture converted from a target document page, and the target document page includes a to-be-masked area of to-be-masked content;

[0039] Step S102, determine the position information of the to-be-masked area according to the target picture;

[0040] Step S103, generate a target replacement area at least according to the position information of the to-be-masked area, and replace the to-be-masked area with the target replacement area, where the size and shape of the target replacement area match the size and shape of the to-be-masked area.

[0041] In the above method for filtering the content of a document page, first, the target document page is converted to obtain a target picture, then the position information of the to-be-masked area is determined according to the target picture, and finally, a target replacement area is generated at least according to the position information of the to-be-masked area, and the generated target replacement area is used to replace the to-be-masked area. Among them, the target document page is a to-be-masked area including to-be-masked content, and the size and shape of the target replacement area match the size and shape of the to-be-masked area. In this solution, the position information of the to-be-masked area is determined according to the target picture, which ensures that the content to be masked can be selectively and intelligently selected. Subsequently, a target replacement area is generated according to the position information of the to-be-masked area, and the generated target replacement area is used to replace the to-be-masked area, that is, the content to be masked is masked, thus solving the problem that it is difficult to intelligently mask the content of a document in the prior art.

[0042] Specifically, the size and shape of the above-mentioned target replacement area and the size and shape of the above-mentioned area to be masked may be the same or different; in the case of being different, the size of the above-mentioned target replacement area may be slightly smaller than the size of the above-mentioned area to be masked, or the size of the above-mentioned target replacement area may be slightly larger than the size of the above-mentioned area to be masked, but the difference in their sizes will not be too large. When the above-mentioned target replacement area is circular, the above-mentioned area to be masked may also be oval, rectangular, etc.

[0043] In the actual application process, the target document page is converted into a target image, and then the position information of the above-mentioned area to be masked is determined according to the target image, which ensures that the determined position information of the area to be masked is more accurate. In this application, it is not limited to the method of converting the target document page into a target image, as long as it can achieve the conversion of the target document page into a target image. In addition, this application does not limit the image format of the above-mentioned converted target image. The above-mentioned target image may be in PNG format or JPG format, which will not be elaborated here one by one. In addition, the size of the above-mentioned target image may be a fixed size, but the size of the above-mentioned target image may also not be restricted. When the size of the above-mentioned target image is a fixed size, it further ensures that the position information of the area to be masked can be obtained more efficiently.

[0044] In a specific embodiment of this application, the above-mentioned target document includes but is not limited to PDF documents. When the above-mentioned target document is a PDF document, this application is used to mask visually homogeneous content in the PDF document. Specific scenarios may include: a certain type of illustration, a certain layout combination of multiple small illustrations, a layout combination of a certain type of illustration and text, a text block with a specific title, a specific clause in a standardized document such as a contract document, a signature in a standardized document such as a contract document, and a seal in an official document such as a public document. In addition, since there is only one large image object the same size as the page in each page of a scanned PDF, it is impossible to replace a part of a page. Therefore, this solution is not applicable to the content masking task of processing scanned PDF documents.

[0045] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0046] In an embodiment of the present application, determining the position information of the to-be-masked area according to the above-mentioned target picture includes: determining the position information of the target image area corresponding to the to-be-masked area in the above-mentioned target picture according to the above-mentioned target picture and the neural network model. The above-mentioned neural network model is obtained by using multiple groups of training data, and the above-mentioned multiple groups of training data include: historical pictures and corresponding historical position information, and the above-mentioned historical position information is the position information corresponding to the to-be-masked area in the above-mentioned historical picture; determining the position information of the to-be-masked area according to the position information of the above-mentioned target image area, the size information of the above-mentioned target picture, and the size information of the above-mentioned target document page. In this solution, the position information of the target image area corresponding to the to-be-masked area in the target picture is determined according to the target picture and the neural network model, which ensures that the position information of the target image area can be determined more efficiently and accurately, and further ensures that the determined to-be-masked area is more accurate.

[0047] In a specific embodiment of the present application, the above-mentioned neural network model includes, but is not limited to, trained object detection deep learning models such as YOLOv5 and FasterRCNN. This deep learning model needs to provide a certain number of labeled data during the training phase. Each data sample includes a historical picture converted to a fixed size (W0, H0) and the corresponding historical position information, that is, the position information of the rectangular frame of the to-be-masked area.

[0048] In another specific embodiment of the present application, as Figure 2 shown ( Figure 2 where x represents the substantial content of the document, and xx that appears in the following figures are all the substantial content of the document), the above-mentioned neural network model needs to provide the converted target picture (W0, H0) during the prediction phase. The model will predict the rectangular frame of the to-be-masked area 100 in the current target picture, that is, the coordinate points (x0, y0) and (x1, y1) of the two diagonal points.

[0049] In order to ensure that the position information of the target image area can be obtained efficiently and accurately, in another embodiment of the present application, the above-mentioned target image area is a rectangular area, the above-mentioned to-be-masked area is a rectangular area, the position information of the above-mentioned target image area includes the coordinate information of the first position point and the coordinate information of the second position point, and the above-mentioned first position point and the above-mentioned second position point are located on the same diagonal line of the rectangular area.

[0050] In the actual application process, when the above-mentioned target image area is a rectangular area, the above-mentioned first position point and the above-mentioned second position point can be represented as (x0, y0) and (x1, y1) respectively.

[0051] In another embodiment of the present application, according to the position information of the above-mentioned target image area, the size information of the above-mentioned target picture, and the size information of the above-mentioned target document, determining the position information of the above-mentioned area to be masked includes: determining a height ratio according to the height of the above-mentioned target picture and the height of the above-mentioned target document page; determining a width ratio according to the width of the above-mentioned target picture and the width of the above-mentioned target document page; adjusting the coordinate information of the above-mentioned first position point according to the above-mentioned height ratio and the above-mentioned width ratio to obtain the coordinate information of the third position point; adjusting the coordinate information of the above-mentioned second position point according to the above-mentioned height ratio and the above-mentioned width ratio to obtain the coordinate information of the fourth position point, and the coordinate information of the above-mentioned third position point and the coordinate information of the above-mentioned fourth position point constitute the position information of the above-mentioned area to be masked. In this solution, the first position point and the second position point are adjusted according to the height ratio and the width ratio, so as to ensure that the obtained third position point and fourth position point are more accurate, that is, the position information of the above-mentioned area to be masked can be further accurately determined.

[0052] In a specific embodiment of the present application, the size of the above-mentioned target picture is (W0, H0), the size of the above-mentioned target document page is (Wi, Hi), the first position point is (x0, y0), and the second position point is (x1, y1). That is, the height ratio can be Hi / H0, and the width ratio can be Wi / W0. According to the above information, the width of the calculated third position point is x0' = x0 * Wi / W0, the height of the third position point is y0' = y0 * Hi / H0, the width of the calculated fourth position point is x1' = x1 * Wi / W0, and the height of the fourth position point is y1' = y1 * Hi / H0.

[0053] In order to ensure that the target image area can be generated more efficiently and accurately, and to facilitate subsequent querying of the target image area according to the target index information, in another implementation of the present application, the area corresponding to the above-mentioned area to be masked in the above-mentioned target picture is the target image area, and the target replacement area is generated at least according to the position information of the above-mentioned area to be masked, including: obtaining target index identification information, the above-mentioned target index identification information includes first identification information and second identification information, the above-mentioned first identification information is the unique identification information of the above-mentioned target image area in the above-mentioned target picture, and the above-mentioned second identification information is the unique identification information of the above-mentioned target document; generating the above-mentioned target replacement area carrying the above-mentioned target index identification information according to the above-mentioned target index identification information and the position information of the above-mentioned area to be masked.

[0054] In an embodiment of the present application, obtaining the target index identification information includes: obtaining the hash value of the above-mentioned target document page as the above-mentioned first identification information; obtaining the serial number of the above-mentioned target image area in the above-mentioned target picture as the second identification information. In this embodiment, based on the hash value of the target document page and the serial number of the target image area in the above-mentioned target picture, it is ensured that the subsequent generated target image area carrying the above-mentioned target index identification information is more accurate, and it is convenient to query the target image area according to the target index identification information later, ensuring a relatively high query efficiency.

[0055] Of course, in actual applications, it is not limited to using the hash value of the target document page as the first identification information. The document ID can also be used as the first identification information, or other information that differentiates different documents can be used as the first identification information. Those skilled in the art can select appropriate information as the first identification information according to the actual situation.

[0056] In another embodiment of the present application, obtaining the target index identification information further includes: obtaining the current date and index extension information. The current date is the date when the above-mentioned target document page needs to be filtered currently, and the index extension information is other identification information representing the above-mentioned target document page. This further ensures that the generated target image area carrying the above-mentioned target index identification information is more accurate, and it is convenient to query the target image area according to the target index identification information later, further ensuring a relatively high query efficiency.

[0057] In the actual application process, the above first identification information is the md5 hash value of the target document page. For example, the above hash value md5 can be information such as "bcfdfd0a921f4b1b95cf9023c8033939". The above second identification information is the serial number of the target image area in the above target picture. For example, the serial number areaIndex in the target picture can be "1", "2", and can also include the page number pageNum where it is located (such as "11", "12"), and can also include the type areaType of the shielding area (such as large area big, medium area middle, small area small). Of course, the above second identification information is a combination of areaIndex, pageNum, and areaType such as "15:4:big", or can also be a combination of areaIndex and pageNum such as "3:12"; the above current date can select an appropriate date granularity according to business needs, such as "20210826" with a daily granularity, "202108" with a monthly granularity, or "2021" with an annual granularity; the above index extension information extendId can be a concatenated combination of several index fields. The index fields include but are not limited to the document theme identifier topic (such as economic, medicine, technology, literature, mechanical), the document source identifier source (such as internet www, publication, original), the user group identifier groupId, the user identifier userId, the document language lang (such as English en, Chinese zh, Italian it, Spanish es), and so on. Then extendId can be a combination of topic, source, and lang such as "economic:www:en", or can also be a combination of source, groupId, and lang such as "original:a2s1d4f3h6:zh".

[0058] In the actual application process, the above target index identification information uniqueIndex can use different forms according to actual needs <date> : <md5> : <extendid> : <areaid> ”、" <date> _ <md5> _ <extendid> _ <areaindex>”, "date = <date>&md5= <md5>&extend_id= <extendid>&area_id= <areaid>", etc. Among them, the third combined index form is a typical HTTP request parameter form, which can be directly used to send HTTP requests with query conditions. The advantage of using a combined index is that it can quickly narrow the search scope in massive data in a hierarchical and progressive manner, and locate the target data more quickly. Among them, the md5 field can be used to establish a cache mechanism for the PDF original to avoid wasting storage space for duplicate documents. ExtendId and areaId provide rich scalable spaces at the document level and the shielded area level respectively, which can be used to personalize the screening of the concerned parts from the massive shielded area data, facilitate the evaluation of the effect of shielded area target detection in segmented scenarios, and analyze the distribution of shielded areas in different scenarios. In addition, the combined index is used to save the original image of the shielded area, and the data query scope is gradually narrowed down in four levels, namely date, document content, document features, and shielded area. This can not only speed up the data query efficiency, but also facilitate personalized data analysis and distributed and block storage or backup of data.

[0059] In a specific embodiment of the present application, the target index identification information may include, in sequence, the document ID, the serial number of the target image area in the target image, the current date, and other identification information of the target document page. In this way, by sequentially querying the document ID, the serial number of the target image area in the target image, the current date, and other identification information of the target document page, the query scope of the data can be gradually narrowed, further ensuring high query efficiency.

[0060] In order to ensure that the generated target image area is more accurate, in another embodiment of the present application, the above-mentioned target replacement area carrying the above-mentioned target index identification information is generated according to the above-mentioned target index identification information and the position information of the above-mentioned area to be shielded, including: generating an initial replacement area according to the above-mentioned target index identification information; generating the above-mentioned target replacement area according to the position information of the above-mentioned area to be shielded and the above-mentioned initial image area.

[0061] In a specific embodiment of the present application, Figure 3 As shown, according to the target index identification information, the initial replacement area 101 is generated, which may be a two-dimensional code image. According to the position information of the area to be shielded and the initial replacement area 101, the target replacement area 102 is generated, that is, the scaled two-dimensional code image.

[0062] In actual application, the target index identification information can be used to generate a QR code image for the content, but it is not limited to a QR code image, and can also be a barcode image. Of course, the content can also be changed to a network link, such as "http: / / example.com / pdf_content_mask?unique_index= <uniqueindex>", and requires the user to log in when accessing the link, which is used for permission control during the query stage of the blocked area.

[0063] In the actual application process, the QR code image can be scaled according to the rectangular frame (x0’, y0’, x1’, y1’) of the blocked area in the original page, and the area to be blocked in the PDF page can be replaced.

[0064] In another embodiment of the present application, after determining the position information of the to-be-blocked area according to the above-mentioned target image, the method further includes: obtaining the target replacement area corresponding to the to-be-blocked area in the above-mentioned target image and storing it in the memory. In this solution, storing the target image area in the memory can facilitate subsequent queries and further ensure high query efficiency.

[0065] In the actual application process, the storage medium can be a local disk, a distributed file system, an object storage service, a database, and the above-mentioned target index identification information can be used as a storage path or a record ID of the database.

[0066] In one embodiment of the present application, at least generating a target replacement area according to the position information of the to-be-blocked area, after replacing the to-be-blocked area with the above-mentioned target replacement area, the method further includes: in response to an operation of scanning the above-mentioned target replacement area, determining the corresponding to-be-blocked area according to the corresponding above-mentioned target index identification information, and replacing the above-mentioned target replacement area in the target document with the corresponding to-be-blocked area, which ensures high replacement efficiency.

[0067] In a specific embodiment of the present application, as Figure 4 shown, after replacing the to-be-blocked area, document parsing or manual reading can be performed. Figure 4 It means that after using the QR code to replace the to-be-blocked area, the document after replacing the to-be-blocked area can be normally read, or the document after parsing the document after replacing the to-be-blocked area, such as the translated one, etc., that is, in Figure 4 , xx in the first picture is the text information before translation, and xx in the second picture is the text information after translation.

[0068] In another specific embodiment of the present application, if it is necessary to query the content of the blocked area after manual reading (such as document review), the target index identification information uniqueIndex or network link contained in the QR code image can be parsed, and the picture corresponding to the blocked area can be found through network access, file query or database query.

[0069] In the actual application process, as Figure 5 shown, Figure 5 The xx in it is the translated text information. If the document changes after document automation parsing (such as PDF document translation) and the area to be masked 100 needs to be restored, the processing flow is as follows: Traverse the target replacement area 102 in the PDF document, find the QR code placeholder image in it, parse the target combination index identification information uniqueIndex contained in the QR code image, query the target replacement area 102 corresponding to the area to be masked according to uniqueIndex, and use the target replacement area 102 to replace the area to be masked 100 in the PDF document, that is, the QR code image.

[0070] The embodiment of the present application also provides a filtering device for the content of a document page. It should be noted that the filtering device for the content of the document page in the embodiment of the present application can be used to execute the filtering method for the content of the document page provided by the embodiment of the present application. The following introduces the filtering device for the content of the document page provided by the embodiment of the present application.

[0071] Figure 6 It is a schematic diagram of the filtering device for the content of the document page according to the embodiment of the present application. As Figure 6 shown, the device includes:

[0072] A first acquisition unit 10, configured to acquire a target image, where the target image is an image converted from a target document page, and the target document page includes an area to be masked with content to be masked;

[0073] A determination unit 20, configured to determine the position information of the area to be masked according to the target image;

[0074] A replacement unit 30, configured to generate a target replacement area at least according to the position information of the area to be masked, and use the target replacement area to replace the area to be masked, where the size and shape of the target replacement area match the size and shape of the area to be masked.

[0075] In the filtering device for the content of the above-mentioned document page, the first acquisition unit is used to acquire a target picture, where the target picture is a picture obtained by converting a target document page, and the target document page is a to-be-blocked area including content to be blocked; the determination unit is used to determine the position information of the to-be-blocked area according to the target picture; the replacement unit generates a target replacement area at least according to the position information of the to-be-blocked area, and uses the target replacement area to replace the to-be-blocked area, and the size and shape of the target replacement area match the size and shape of the to-be-blocked area. In this solution, the position information of the to-be-blocked area is determined according to the target picture, which ensures that the content to be blocked can be selectively and intelligently selected. Subsequently, a target replacement area is generated according to the position information of the to-be-blocked area, and the generated target replacement area is used to replace the to-be-blocked area, that is, the content to be blocked is blocked, thus solving the problem that it is difficult to intelligently block document content in the prior art.

[0076] Specifically, the above-mentioned matching means that the size and shape of the target replacement area and the size and shape of the to-be-blocked area can be the same or different; in the case of being different, the size of the target replacement area can be slightly smaller than the size of the to-be-blocked area, or the size of the target replacement area can be slightly larger than the size of the to-be-blocked area, but the difference between the two sizes will not be too large. When the target replacement area is circular, the to-be-blocked area can also be oval, rectangular, etc.

[0077] In the actual application process, the target document page is converted into a target picture, and then the position information of the to-be-blocked area is determined according to the target picture, which ensures that the determined position information of the to-be-blocked area is more accurate. In this application, it is not limited to the method of converting the target document page into a target picture, as long as it can realize the conversion of the target document page into a target picture. In addition, this application does not limit the picture format of the converted target picture. The target picture can be in PNG format or JPG format, which will not be elaborated here one by one. In addition, the size of the target picture can be a fixed size, but the size of the target picture does not have to be restricted. In the case where the size of the target picture is a fixed size, it further ensures that the position information of the to-be-blocked area can be obtained more efficiently.

[0078] In a specific embodiment of the present application, the above-mentioned target document includes, but is not limited to, PDF documents. In the case where the above-mentioned target document is a PDF document, the present application is used to mask visually homogeneous content in the PDF document. Specific scenarios may include: illustrations of a certain category, a certain layout combination of multiple small illustrations, a layout combination of a certain type of illustration and text, text blocks containing specific titles, specific clauses in a standardized document such as a contract document, signatures in a standardized document such as a contract document, and official seals in official documents such as official letters. Additionally, in each page of a scanned PDF, there is only one large image object that is the same size as the page, and it is impossible to replace a certain part of a page. Therefore, this solution is not applicable to the content masking task of scanned PDF documents.

[0079] In an embodiment of the present application, the above-mentioned determination unit further includes a first determination module and a second determination module. The above-mentioned first determination module is used to determine the position information of the target image area corresponding to the area to be masked in the above-mentioned target image according to the above-mentioned target image and the neural network model. The above-mentioned neural network model is obtained by using multiple sets of training data. The above-mentioned multiple sets of training data include: historical pictures and corresponding historical position information, and the above-mentioned historical position information is the position information corresponding to the area to be masked in the above-mentioned historical picture; the above-mentioned second determination module is used to determine the position information of the area to be masked according to the position information of the above-mentioned target image area, the size information of the above-mentioned target image, and the size information of the above-mentioned target document page. In this solution, according to the target image and the neural network model, the position information of the target image area corresponding to the area to be masked in the target image is determined, which ensures that the position information of the target image area can be determined more efficiently and accurately, and further ensures that the determined area to be masked is more accurate.

[0080] In a specific embodiment of the present application, the above-mentioned neural network model includes, but is not limited to, a trained object detection deep learning model such as YOLOv5 or FasterRCNN. This deep learning model requires a certain amount of labeled data to be provided during the training phase. Each data sample includes a historical picture converted to a fixed size (W0, H0) and the corresponding historical position information, that is, the position information of the rectangular frame of the area to be masked.

[0081] In yet another specific embodiment of the present application, as Figure 2 shown, the above-mentioned neural network model needs to provide the converted target picture (W0, H0) during the prediction phase. The model will predict the rectangular frame of the area to be masked 100 in the current target picture, that is, the coordinate points (x0, y0) and (x1, y1) of the two diagonal points.

[0082] In order to ensure that the position information of the target image region can be obtained efficiently and accurately, in another embodiment of the present application, the above-mentioned target image region is a rectangular region, the above-mentioned region to be masked is a rectangular region, the position information of the above-mentioned target image region includes the coordinate information of the first position point and the coordinate information of the second position point, and the above-mentioned first position point and the above-mentioned second position point are located on the same diagonal line of the rectangular region.

[0083] In the actual application process, when the above-mentioned target image region is a rectangular region, the above-mentioned first position point and the above-mentioned second position point can be respectively represented as (x0, y0) and (x1, y1).

[0084] In another embodiment of the present application, the above-mentioned second determination module further includes a first determination sub-module, a second determination sub-module, a first adjustment sub-module, and a second adjustment sub-module. The above-mentioned first determination sub-module is used to determine the height ratio according to the height of the above-mentioned target picture and the height of the above-mentioned target document page; the above-mentioned second determination sub-module is used to determine the width ratio according to the width of the above-mentioned target picture and the width of the above-mentioned target document page; the above-mentioned first adjustment sub-module is used to adjust the coordinate information of the above-mentioned first position point according to the above-mentioned height ratio and the above-mentioned width ratio to obtain the coordinate information of the third position point; the above-mentioned second adjustment sub-module is used to adjust the coordinate information of the above-mentioned second position point according to the above-mentioned height ratio and the above-mentioned width ratio to obtain the coordinate information of the fourth position point. The coordinate information of the above-mentioned third position point and the coordinate information of the above-mentioned fourth position point constitute the position information of the above-mentioned region to be masked. In this solution, the first position point and the second position point are adjusted according to the height ratio and the width ratio, which ensures that the obtained third position point and fourth position point are more accurate, that is, the position information of the above-mentioned region to be masked can be determined more accurately.

[0085] In a specific embodiment of the present application, the size of the above-mentioned target picture is (W0, H0), the size of the above-mentioned target document page is (Wi, Hi), the first position point is (x0, y0), and the second position point is (x1, y1). That is, the height ratio can be Hi / H0 and the width ratio can be Wi / W0. According to the above information, the width of the calculated third position point is x0' = x0 * Wi / W0, the height of the third position point is y0' = y0 * Hi / H0, the width of the calculated fourth position point is x1' = x1 * Wi / W0, and the height of the fourth position point is y1' = y1 * Hi / H0.

[0086] In order to ensure that the target image area can be generated more efficiently and accurately, and to facilitate subsequent querying of the target image area based on the target index information, in another implementation of this application, the above-mentioned replacement unit includes an acquisition module and a generation module. The area in the above-mentioned target picture corresponding to the above-mentioned area to be masked is the target image area. The above-mentioned acquisition module is used to acquire target index identification information, and the above-mentioned target index identification information includes first identification information and second identification information. The above-mentioned first identification information is the unique identification information of the above-mentioned target image area in the above-mentioned target picture, and the above-mentioned second identification information is the unique identification information of the above-mentioned target document. The above-mentioned generation module is used to generate the above-mentioned target replacement area carrying the above-mentioned target index identification information according to the above-mentioned target index identification information and the position information of the above-mentioned area to be masked.

[0087] In an embodiment of this application, the above-mentioned acquisition module further includes a first acquisition sub-module and a second acquisition sub-module. Among them, the above-mentioned first acquisition sub-module is used to acquire the hash value of the above-mentioned target document page as the above-mentioned first identification information. The above-mentioned second acquisition sub-module is used to acquire the serial number of the above-mentioned target image area in the above-mentioned target picture as the second identification information. In this embodiment, based on the hash value of the target document page and the serial number of the target image area in the above-mentioned target picture, it is ensured that the subsequent generated target image area carrying the above-mentioned target index identification information is more accurate, and it is convenient to query the target image area according to the target index identification information subsequently, ensuring higher query efficiency.

[0088] Of course, in actual applications, it is not limited to using the hash value of the target document page as the first identification information. The document ID can also be used as the first identification information, or other information that distinguishes different documents can be used as the first identification information. Those skilled in the art can select appropriate information as the first identification information according to the actual situation.

[0089] In another embodiment of this application, the above-mentioned acquisition module further includes a third acquisition sub-module, which is used to acquire the current date and index extension information. The above-mentioned current date is the date when the above-mentioned target document page needs to be filtered currently, and the above-mentioned index extension information is other identification information representing the above-mentioned target document page. This further ensures that the generated target image area carrying the above-mentioned target index identification information is more accurate, and it is convenient to query the target image area according to the target index identification information subsequently, further ensuring higher query efficiency.

[0090] In the actual application process, the above first identification information is the md5 hash value of the target document page. For example, the above md5 hash value can be information such as "bcfdfd0a921f4b1b95cf9023c8033939". The above second identification information is the serial number of the target image area in the above target picture. For example, the serial number areaIndex in the target picture can be "1", "2", and can also include the page number pageNum where it is located (such as "11", "12"), and can also include the type areaType of the shielding area (such as large area big, medium area middle, small area small). Of course, the above second identification information is a combination of areaIndex, pageNum, and areaType such as "15:4:big", or can also be a combination of areaIndex and pageNum such as "3:12"; the above current date can select an appropriate date granularity according to business needs, such as "20210826" with a daily granularity, "202108" with a monthly granularity, or "2021" with an annual granularity; the above index extension information extendId can be the concatenation combination of several index fields. The index fields include but are not limited to the document theme identifier topic (such as economic, medicine, technology, literature, mechanical), the document source identifier source (such as internet www, publication, original), the user group identifier groupId, the user identifier userId, the document language lang (such as English en, Chinese zh, Italian it, Spanish es), and so on. Then extendId can be a combination of topic, source, and lang such as "economic:www:en", or can also be a combination of source, groupId, and lang such as "original:a2s1d4f3h6:zh".

[0091] In the actual application process, the above target index identification information uniqueIndex can use different forms according to actual needs" <date> : <md5> : <extendid> : <areaid> ”、" <date> _ <md5> _ <extendid> _ <areaindex>”, "date = <date>&md5= <md5>&extend_id= <extendid>&area_id= <areaid>” and other different forms. Among them, the third combined index form is the typical form of HTTP request parameters and can be directly used to send HTTP requests with query conditions. The advantage of using the combined index is that it can hierarchically narrow down the search scope in a large amount of data and locate the target data more quickly. Among them, the md5 field can be used to establish a caching mechanism for the original PDF text to avoid wasting storage space for duplicate documents. The extendId and areaId respectively provide rich expandable spaces at the document level and the shielding area level, which can be used to individually screen and focus on parts from a large amount of shielding area data, facilitating the evaluation of the shielding area target detection effect in a specific scenario and also analyzing the distribution of shielding areas in different scenarios. In addition, the original shielding area image is saved using the combined index method, gradually narrowing down the data query scope in 4 levels: date, document content, document features, and shielding area. This can not only improve the data query efficiency but also facilitate personalized data analysis and distributed and block-based storage or backup of data.

[0092] In a specific embodiment of the present application, the target index identification information may sequentially include the document ID, the serial number of the target image area in the target picture, the current date, and other identification information of the target document page. By sequentially querying the document ID, the serial number of the target image area in the target picture, the current date, and other identification information of the target document page, the data query scope can be gradually narrowed down, further ensuring a relatively high query efficiency.

[0093] In order to ensure that the generated target image area is more accurate, in another embodiment of the present application, the above-mentioned generation module further includes a first generation sub-module and a second generation sub-module. Among them, the first generation sub-module is used to generate an initial replacement area according to the above-mentioned target index identification information; the second generation sub-module is used to generate the above-mentioned target replacement area according to the position information of the area to be shielded and the above-mentioned initial image area.

[0094] In a specific embodiment of the present application, as Figure 3 shown, according to the above-mentioned target index identification information, an initial replacement area 101 is generated, which can be a QR code image. According to the position information of the area to be shielded and the above-mentioned initial replacement area 101, the above-mentioned target replacement area 102, that is, the scaled QR code image, is generated.

[0095] In the actual application process, the above-mentioned target index identification information can be used to generate a QR code picture for the content, but it is not limited to the QR code picture. It can also be a barcode picture. Of course, the content can also be changed to a network link, such as "http: / / example.com / pdf_content_mask?unique_index = <uniqueindex>", and requires the user to log in when accessing the link, which is used for permission control during the shielding area query phase.

[0096] In the actual application process, the QR code image can be scaled according to the rectangular frame (x0’, y0’, x1’, y1’) of the shielding area in the original page, and the area to be shielded in the PDF page can be replaced.

[0097] In another embodiment of the present application, after determining the position information of the area to be shielded according to the above target image, the filtering device further includes a second acquisition unit, which is used to acquire the target replacement area corresponding to the area to be shielded in the target image and store it in the memory. In this solution, storing the target replacement area in the memory can facilitate subsequent queries and further ensure higher query efficiency.

[0098] In the actual application process, the storage medium can be a local disk, a distributed file system, an object storage service, or a database, and the above target index identification information can be used as a storage path or a record ID of the database.

[0099] In one embodiment of the present application, at least the target replacement area is generated according to the position information of the area to be shielded. After replacing the area to be shielded with the target replacement area, the device further includes a response unit, which is used to respond to the operation of scanning the target replacement area, determine the corresponding area to be shielded according to the corresponding target index identification information, and replace the target replacement area in the target document with the corresponding area to be shielded, which ensures higher replacement efficiency.

[0100] In a specific embodiment of the present application, as Figure 4 shown, after replacing the area to be shielded, document parsing or manual reading can be performed. Figure 4 It means that after replacing the area to be shielded with a QR code, the document after replacing the area to be shielded can be normally read, or the document after parsing the document after replacing the area to be shielded, such as the translated one, etc. That is, in Figure 4 , xx in the first picture is the text information before translation, and xx in the second picture is the text information after translation.

[0101] In another specific embodiment of the present application, if the content of the shielding area needs to be queried after manual reading (such as document review), the target index identification information uniqueIndex or network link contained in the QR code image can be parsed, and the picture corresponding to the shielding area can be found through network access, file query, or database query.

[0102] In the actual application process, as Figure 5 shown, Figure 5 The xx in it is the translated text information. If the document changes after document automation parsing (such as PDF document translation), and it is necessary to restore the area to be masked 100, the processing flow is as follows: Traverse the target replacement area 102 in the PDF document, find the QR code placeholder image in it, parse the target combination index identification information uniqueIndex contained in the QR code image, query the target replacement area 102 corresponding to the area to be masked according to uniqueIndex, and use the target replacement area 102 to replace the area to be masked 100 in the PDF document, that is, the QR code image.

[0103] The filtering device for the above document page content includes a processor and a memory. The above first acquisition unit, determination unit, replacement unit, etc. are all stored in the memory as program units, and the corresponding functions are implemented by the processor executing the above program units stored in the memory.

[0104] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and the problem that it is difficult to intelligently mask document content in the prior art can be solved by adjusting the kernel parameters.

[0105] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0106] An embodiment of the present invention provides a storage medium, on which a program is stored, and when the program is executed by a processor, the above filtering method for document page content is implemented.

[0107] An embodiment of the present invention provides a processor, and the above processor is used to run a program, wherein when the above program runs, the above filtering method for document page content is executed.

[0108] An embodiment of the present invention provides a device, the device includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, at least the following steps are implemented:

[0109] Step S101, obtain a target picture, the above target picture is a picture converted from a target document page, and the above target document page includes an area to be masked with content to be masked;

[0110] Step S102, determine the position information of the above area to be masked according to the above target picture;

[0111] Step S103, generate a target replacement area at least according to the position information of the to-be-masked area, and replace the to-be-masked area with the target replacement area, where the size and shape of the target replacement area match those of the to-be-masked area.

[0112] The device in this article can be a server, a PC, a PAD, a mobile phone, etc.

[0113] This application also provides a computer program product, which when executed on a data processing device is adapted to execute a program initialized with at least the following method steps:

[0114] Step S101, obtain a target picture, where the target picture is a picture converted from a target document page, and the target document page includes a to-be-masked area of to-be-masked content;

[0115] Step S102, determine the position information of the to-be-masked area according to the target picture;

[0116] Step S103, generate a target replacement area at least according to the position information of the to-be-masked area, and replace the to-be-masked area with the target replacement area, where the size and shape of the target replacement area match those of the to-be-masked area.

[0117] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0118] In several embodiments provided by this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be electrical or other forms.

[0119] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0120] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0121] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0122] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0123] 1) In the method for filtering the content of the document page of the present application, first, the target document page is transformed to obtain a target picture, then the position information of the area to be blocked is determined according to the above-mentioned target picture, and finally, at least according to the position information of the area to be blocked, a target replacement area is generated, and the generated target replacement area is used to replace the area to be blocked. Among them, the above-mentioned target document page is an area to be blocked including the content to be blocked, and the size and shape of the above-mentioned target replacement area match the size and shape of the area to be blocked. In this solution, the position information of the area to be blocked is determined according to the target picture, which ensures that the content to be blocked can be selectively and intelligently selected. Subsequently, a target replacement area is generated according to the position information of the area to be blocked, and the generated target replacement area is used to replace the area to be blocked, that is, the content to be blocked is blocked, thereby solving the problem that it is difficult to intelligently block the document content in the prior art.

[0124] 2) In the filtering device for the content of the document page of the present application, the first acquisition unit is used to acquire a target picture, where the target picture is a picture converted from a target document page, and the target document page is a to-be-blocked area including content to be blocked; the determination unit is used to determine the position information of the to-be-blocked area according to the target picture; the replacement unit is used to generate a target replacement area at least according to the position information of the to-be-blocked area, and replace the to-be-blocked area with the target replacement area, and the size and shape of the target replacement area match the size and shape of the to-be-blocked area. In this solution, the position information of the to-be-blocked area is determined according to the target picture, which ensures that the content to be blocked can be selectively and intelligently selected. Subsequently, the target replacement area is generated according to the position information of the to-be-blocked area, and the generated target replacement area is used to replace the to-be-blocked area, that is, the content to be blocked is blocked, thus solving the problem that it is difficult to intelligently block the document content in the prior art.

[0125] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.< / uniqueindex> < / areaid> < / extendid> < / date> < / areaindex> < / extendid> < / md5> < / date> < / areaid> < / extendid> < / md5> < / date> < / uniqueindex> < / areaid> < / extendid> < / date> < / areaindex> < / extendid> < / md5> < / date> < / areaid> < / extendid> < / md5> < / date>

Claims

1. A method for filtering the content of a document page, characterized in that, the document page includes multiple pages, including: obtaining a target image, where the target image is an image obtained by converting a target document page, and the target document page includes a to-be-blocked area of to-be-blocked content; determining the position information of the to-be-blocked area according to the target image; generating a target replacement area at least according to the position information of the to-be-blocked area, and replacing the to-be-blocked area with the target replacement area, where the size and shape of the target replacement area match the size and shape of the to-be-blocked area; the area in the target image corresponding to the to-be-blocked area is a target image area, and generating a target replacement area at least according to the position information of the to-be-blocked area includes: obtaining target index identification information, where the target index identification information is used to query the target image area, and the target index identification information includes first identification information and second identification information, the first identification information is the unique identification information of the target image area in the target image, and the second identification information is the unique identification information of the target document; generating the target replacement area carrying the target index identification information according to the target index identification information and the position information of the to-be-blocked area.

2. The method according to claim 1, characterized in that, determining the position information of the to-be-blocked area according to the target image includes: determining the position information of the target image area corresponding to the to-be-blocked area in the target image according to the target image and a neural network model, where the neural network model is obtained by using multiple groups of training data, and the multiple groups of training data include: historical images and corresponding historical position information, and the historical position information is the position information corresponding to the to-be-blocked area in the historical image; determining the position information of the to-be-blocked area according to the position information of the target image area, the size information of the target image, and the size information of the target document page.

3. The method according to claim 2, characterized in that, the target image area is a rectangular area, the to-be-blocked area is a rectangular area, the position information of the target image area includes the coordinate information of a first position point and the coordinate information of a second position point, and the first position point and the second position point are located on the same diagonal of the rectangular area.

4. The method according to claim 3, characterized in that, determining the position information of the to-be-blocked area according to the position information of the target image area, the size information of the target image, and the size information of the target document page includes: determining a height ratio according to the height of the target image and the height of the target document page; determining a width ratio according to the width of the target image and the width of the target document page; adjusting the coordinate information of the first position point according to the height ratio and the width ratio to obtain the coordinate information of a third position point; Adjust the coordinate information of the second position point according to the height ratio and the width ratio to obtain the coordinate information of the fourth position point, and the coordinate information of the third position point and the coordinate information of the fourth position point constitute the position information of the area to be blocked.

5. The method according to claim 1, wherein, obtaining target index identification information includes: obtaining the hash value of the target document page as the first identification information; obtaining the serial number of the target image area in the target picture as the second identification information.

6. The method according to claim 1, wherein, obtaining target index identification information further includes: obtaining the current date and index extension information, where the current date is the date when the target document page needs to be filtered currently, and the index extension information is other identification information characterizing the target document page.

7. The method according to claim 1, wherein, generating the target replacement area carrying the target index identification information according to the target index identification information and the position information of the area to be blocked includes: generating an initial replacement area according to the target index identification information; generating the target replacement area according to the position information of the area to be blocked and the initial replacement area.

8. The method according to claim 1, wherein, after determining the position information of the area to be blocked according to the target picture, the method further includes: obtaining the target replacement area corresponding to the area to be blocked in the target picture and storing it in a memory.

9. The method according to claim 1, wherein, after generating the target replacement area at least according to the position information of the area to be blocked and replacing the area to be blocked with the target replacement area, the method further includes: in response to an operation of scanning the target replacement area, determining the corresponding area to be blocked according to the corresponding target index identification information, and replacing the target replacement area in the target document with the corresponding area to be blocked.

10. A filtering device for document page content, wherein, the document page includes multiple pages, including: a first obtaining unit for obtaining a target picture, where the target picture is a picture converted from a target document page, and the target document page includes an area to be blocked with content to be blocked; a determining unit for determining the position information of the area to be blocked according to the target picture; a replacing unit for generating a target replacement area at least according to the position information of the area to be blocked and replacing the area to be blocked with the target replacement area, where the size and shape of the target replacement area match the size and shape of the area to be blocked; the area corresponding to the area to be blocked in the target picture is a target image area, and generating a target replacement area at least according to the position information of the area to be blocked includes: Obtain target index identification information, which is used to query the target image area. The target index identification information includes first identification information and second identification information. The first identification information is the unique identification information of the target image area in the target picture, and the second identification information is the unique identification information of the target document; Generate the target replacement area carrying the target index identification information according to the target index identification information and the position information of the area to be masked.

11. A computer-readable storage medium, characterized in that, the computer-readable storage medium includes a stored program, wherein the program executes the filtering method according to any one of claims 1 to 9.

12. A processor, characterized in that, the processor is used to run a program, wherein when the program runs, it executes the filtering method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Privacy protection method and device, apparatus and storage medium

    CN110135195A

  • Processing method and device for protecting privacy information in picture

    CN111177757A

  • Method and device for desensitizing document image, electronic equipment and medium

    CN112380566A