File sensitive information processing method and system, computer and storage medium

Through multi-level detection and automated processing, the problems of low efficiency and low accuracy in the blind processing of bid documents have been solved, and the automated and comprehensive blurring of sensitive information has been achieved, ensuring the security and accuracy of bid documents.

CN122020722APending Publication Date: 2026-05-12江西博微新技术有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-10
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The current blinding process for tender documents relies on manual intervention, resulting in low efficiency, low accuracy, and the risk of human leakage, making it difficult to support large-scale blind evaluation.

Method used

Through multi-level detection and automated processing, including header and footer detection, seal detection, and text detection, sensitive information in tender documents is identified and blurred, achieving automatic blind processing.

Benefits of technology

It has achieved automated and comprehensive blind processing of tender documents, improved processing efficiency, avoided human leakage, and ensured the accuracy and security of the processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020722A_ABST
    Figure CN122020722A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information processing, and provides a file sensitive information processing method and system, a computer and a storage medium, and the file sensitive information processing method comprises the steps: obtaining a first task data set, and converting a to-be-processed file into a plurality of to-be-processed file pictures; based on the second task data set, judging whether header and footer detection is started or not, and if the header and footer detection is started, detecting and generating a plurality of area detection results; whether seal detection is started or not is judged based on the third task data set, and if seal detection is started, a plurality of seal detection results are generated through detection; judging whether text detection is started or not based on the fourth task data set, and if the text detection is started, detecting and generating a plurality of sensitive text detection results; and judging whether fuzzy processing is started or not based on the fifth task data set, and if the fuzzy processing is started, generating a fuzzy file. By adopting the method, the problems of low blind state processing efficiency, inaccurate blind state and artificial leakage risk can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information processing technology, and in particular to a method, system, computer, and storage medium for processing sensitive information in documents. Background Technology

[0002] The bidding and tendering process demands increasingly more standardized and efficient management, as well as fairness and transparency. However, traditional bidding and tendering models are susceptible to limitations in expert resources and geographical location, which can lead to biased or favoritism affecting the results. Therefore, an increasing number of bidding and tendering processes are adopting blind evaluation models.

[0003] The blind review model, by reviewing bid documents without the knowledge of the bidding companies, eliminates conflicts of interest, avoids subjective bias, and removes interference from non-technical factors. This allows experts to focus their energy and attention on the quality and technical strength of the bids, enabling them to analyze the feasibility, innovativeness, and advancement of the technical solutions, as well as the reasonableness of the commercial terms, resulting in more professional and accurate evaluations. This ensures fair competition among bidding companies, improves review quality, and enhances the credibility of bidding activities.

[0004] The existing blinding process for tender documents relies heavily on manual labor. Locating sensitive information in a single tender document is time-consuming, requiring page-by-page screening of company names, qualification numbers, project manager information, etc., resulting in a long processing time for a single project's tender documents. Furthermore, different types of tender documents have specific specifications and requirements. When manually classifying documents, a lack of thorough understanding of the rules or inaccurate grasp of the classification standards may affect the professionalism and standardization of the processed documents, making it difficult to support large-scale blind evaluation. In addition, manual desensitization of documents carries the risk of incomplete identity concealment and human leakage. The existing blinding process for documents suffers from problems such as low efficiency, inaccurate blinding, and the risk of human leakage. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide a method, system, computer, and storage medium for processing sensitive information in documents. This invention achieves automatic blind processing of sensitive identifiers, text, and other information in bidding documents by performing multi-level and comprehensive detection of sensitive information and batch blurring. The invention aims to solve the technical problems of low efficiency, inaccurate blinding, and risk of human leakage in existing bidding document processing technologies.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solution: A method for processing sensitive information in a document includes the following steps: Obtain a first task dataset, obtain files to be processed based on the first task dataset, convert the files to be processed into several images of files to be processed, and update the first task dataset to a second task dataset; Based on the second task dataset, determine whether to start header and footer detection. If header and footer detection is started, detect several header and footer regions in several files to be processed, generate several region detection results, and update the second task dataset to the third task dataset. Based on the third task dataset, determine whether to start stamp detection. If stamp detection starts, detect several stamps in several files to be processed, generate several stamp detection results, and update the third task dataset to the fourth task dataset. Based on the fourth task dataset, determine whether to start text detection. If text detection is started, detect several sensitive texts in several files to be processed and generate several sensitive text detection results, and update the fourth task dataset to the fifth task dataset. Based on the fifth task dataset, it is determined whether to start blurring. If blurring starts, based on several region detection results, several seal detection results, and several sensitive text detection results, several regions to be blurred are selected from several files to be processed, and the several regions to be blurred are converted into several blurred regions, so that the several files to be processed are converted into several blurred files, and the several blurred files are merged to generate a blurred file.

[0007] Furthermore, the first task dataset includes a subset of bidding information, a first current page number, and a first current execution time. The subset of bidding information includes file path, file page number, blind review identifier, and bidding sensitive information. The steps of obtaining the file to be processed based on the first task dataset, converting the file to be processed into several file images, and updating the first task dataset to the second task dataset include: Determine whether to perform a blind review based on the blind review indicator. If a blind review is performed, obtain the file to be processed based on the file path. The file to be processed is divided into several pages according to the number of pages in the file; The page to be processed is converted into an image of the file to be processed. The number of pages converted and the conversion time are recorded. The first current page number is updated to the second current page number based on the number of pages converted. The first current execution time is updated to the second current execution time based on the conversion time. The second current page number and the second current execution time are displayed on the user interface. The subset of bidding information, the second current page number, and the second current execution time are combined to form the second task dataset.

[0008] Furthermore, the step of determining whether to start header and footer detection based on the second task dataset, and if header and footer detection is started, detecting several header and footer regions in several of the images to be processed, generating several region detection results, and updating the second task dataset to the third task dataset includes: Compare the second current page number with the document page number. If the second current page number is equal to the document page number, then start header and footer detection. The header and footer areas in the image to be processed are detected, a region detection result is generated, the number of pages detected and the region detection time are recorded, the second current page number is updated to the third current page number based on the number of pages detected, the second current execution time is updated to the third current execution time based on the region detection time, and the third current page number and the third current execution time are displayed on the user interface. The third task dataset is composed of the subset of bidding information, the third current page number, and the third current execution time.

[0009] Furthermore, the step of determining whether to start stamp detection based on the third task dataset, and if stamp detection is started, detecting several stamps in several of the images to be processed, generating several stamp detection results, and updating the third task dataset to the fourth task dataset includes: Compare the third current page number with the document page number. If the third current page number is equal to the document page number, then start the seal detection. An image recognition algorithm is invoked to detect the seal in the image of the file to be processed, a seal detection result is generated, the number of pages detected and the seal detection time are recorded, the third current page number is updated to the fourth current page number based on the number of pages detected, the third current execution time is updated to the fourth current execution time based on the seal detection time, and the fourth current page number and the fourth current execution time are displayed on the user interface. The subset of bidding information, the fourth current page number, and the fourth current execution time are combined to form the fourth task dataset.

[0010] Furthermore, the step of determining whether to start text detection based on the fourth task dataset, and if text detection is started, detecting several sensitive texts in several of the images to be processed, generating several sensitive text detection results, and updating the fourth task dataset to the fifth task dataset includes: Compare the fourth current page number with the file page number. If the fourth current page number is equal to the file page number, then text detection begins. Based on the bidding sensitive information in the bidding information subset, sensitive text in the file image to be processed is detected, sensitive text detection results are generated, the text detection page number and text detection time are recorded, the fourth current page number is updated to the fifth current page number based on the text detection page number, the fourth current execution time is updated to the fifth current execution time based on the text detection time, and the fifth current page number and the fifth current execution time are displayed on the user interface; The fifth task dataset is composed of the subset of bidding information, the fifth current page number, and the fifth current execution time.

[0011] Furthermore, the step of determining whether to start fuzzing based on the fifth task dataset specifically involves: The fifth current page number is compared with the file page number. If the fifth current page number is equal to the file page number, then fuzzing processing begins.

[0012] A file sensitive information processing system, applied to the file sensitive information processing method described in the above technical solution, the system comprising: The acquisition module is used to acquire a first task dataset, acquire files to be processed based on the first task dataset, convert the files to be processed into several images of files to be processed, and update the first task dataset to a second task dataset. The first detection module is used to determine whether to start header and footer detection based on the second task dataset. If header and footer detection is started, it detects several header and footer regions in several files to be processed, generates several region detection results, and updates the second task dataset to the third task dataset. The second detection module is used to determine whether to start stamp detection based on the third task dataset. If stamp detection is started, it detects several stamps in several files to be processed, generates several stamp detection results, and updates the third task dataset to the fourth task dataset. The third detection module is used to determine whether to start text detection based on the fourth task dataset. If text detection is started, it detects several sensitive texts in several files to be processed and generates several sensitive text detection results, and updates the fourth task dataset to the fifth task dataset. The blurring module is used to determine whether to start blurring based on the fifth task dataset. If blurring starts, it selects several regions to be blurred from several files to be processed based on several region detection results, several seal detection results, and several sensitive text detection results. It converts the several regions to be blurred into several blurred regions, so that the several files to be processed are converted into several blurred files. Then, it merges the several blurred files to generate a blurred file.

[0013] A computer includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the file sensitive information processing method as described in the above technical solutions.

[0014] A storage medium storing a computer program thereon, which, when executed by a processor, implements the file sensitive information processing method as described in the above technical solution.

[0015] Compared with existing technologies, the advantages of this invention are as follows: By performing segmentation and conversion, header and footer detection, seal detection, and text detection on the files to be processed, sensitive information in the tender documents requiring blind processing is fully and automatically detected, and the sensitive information is blurred in batches, turning the files to be processed into blurred files. This enables automatic blind processing of batches of tender documents, avoiding human contact with the document content and avoiding the problems of low efficiency, inaccurate blind processing, and human leakage caused by the high dependence on manual processing of tender documents in traditional methods. By identifying, judging, and updating the task dataset step by step, it ensures that no process such as conversion and detection of multi-page documents is missed, and the execution status is displayed intuitively during the processing, making it convenient for users to monitor the automated document processing progress. Attached Figure Description

[0016] Figure 1 This is a flowchart of the document sensitive information processing method in an embodiment of the present invention; Figure 2 This is a structural block diagram of the document sensitive information processing system in an embodiment of the present invention; The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation

[0017] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.

[0018] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0020] Please see Figure 1 The document sensitive information processing method in the first embodiment of the present invention includes the following steps: Step S10: Obtain the first task dataset, obtain the files to be processed based on the first task dataset, convert the files to be processed into several images of files to be processed, and update the first task dataset to the second task dataset; Preferably, the file to be processed is a tender document. Sensitive information affecting the blind evaluation of the tender document includes, but is not limited to, company name, seller of invoice images, buyer of invoice images, various seals, project manager, etc. When processing blind evaluation tender documents on a large scale, the file is decompressed and uploaded to an FTP server. The project name, bid name, package name, bidder name, MD5 code, FTP upload directory and creation time associated with the file are written into a MySQL database. It can be understood that splitting the file to be processed and converting it into several images of the file to be processed is beneficial for the unified and standardized processing of files of different types and formats.

[0021] In step S10, the first task dataset includes a subset of bidding information, a first current page number, and a first current execution time. The subset of bidding information includes file path, file page number, blind review identifier, and bidding sensitive information. Step S10 further includes: S110: Determine whether to perform blind evaluation based on the blind evaluation identifier. If blind evaluation is to be performed, obtain the file to be processed based on the file path. Preferably, a button to switch between blind review and unblind review is provided on the user's unzipped bid document interface. If blind review is not switched, the original unzipping function logic remains unchanged. If blind review is switched, the file will be processed in a blind state.

[0022] S120: Divide the file to be processed into several pages according to the number of pages in the file; S130: Convert the file page to be processed into a file image, record the number of pages converted and the conversion time, update the first current page number to the second current page number based on the number of pages converted, update the first current execution time to the second current execution time based on the conversion time, and display the second current page number and the second current execution time on the user interface; Preferably, several images of the files to be processed are stored in a designated directory. The image conversion process for each page is executed in a loop until all pages are processed. The task status is updated in a timely manner. The file ID is read from the Redis task queue, and it is checked whether there are any tasks to be processed. If there are tasks in the queue, the task status is updated to "processing" and the file is sliced. If there are no tasks in the queue, the queue is queried again after a certain period of time to adjust the program process in a timely manner.

[0023] S140: Combine the tender information subset, the second current page number, and the second current execution time to form a second task dataset.

[0024] Preferably, the first task dataset is stored in the form of several tables. The fields of the tables include, but are not limited to, file ID, project name, round name, bid name, package name, bidder name, bid type, number of pages, file path, whether blind evaluation is required, task type, start time, end time, time taken, number of pages, current page number, current execution time, execution status, and execution client name. The task type is slicing, header and footer detection, stamp detection, text recognition, and fuzzing. The execution status includes unassigned, assigned, executing, and completed, with each status corresponding to a different code. The execution client name consists of the client IP and the client name. In the processing program, the task dataset is queried periodically to identify the task progress status in order to allocate Redis terminals according to the task progress status.

[0025] Step S20: Based on the second task dataset, determine whether to start header and footer detection. If header and footer detection is started, detect several header and footer regions in several files to be processed, generate several region detection results, and update the second task dataset to the third task dataset. Preferably, the task status is updated in a timely manner to enable automatic processing of the next step. If header and footer detection is to begin, the DocLayout-L module is invoked via an HTTP request to perform the detection operation.

[0026] Step S20 includes: S210: Compare the second current page number with the document page number. If the second current page number is equal to the document page number, start header and footer detection. Preferably, if the second current page number is equal to the number of pages in the file, it indicates that the file slicing in the previous process has been completed, and the header and footer detection process can be started automatically.

[0027] S220: Detect the header and footer areas in the image of the file to be processed, generate area detection results, record the number of pages detected and the area detection time, update the second current page number to the third current page number based on the number of pages detected, update the second current execution time to the third current execution time based on the area detection time, and display the third current page number and the third current execution time on the user interface; Preferably, the results of header and footer detection for each page are stored in a directory format, and the area detection results are stored in the form of filename "layout_001.json".

[0028] S230: The tender information subset, the third current page number, and the third current execution time are combined to form a third task dataset.

[0029] Step S30: Based on the third task dataset, determine whether to start stamp detection. If stamp detection is started, detect several stamps in several files to be processed, generate several stamp detection results, and update the third task dataset to the fourth task dataset. Preferably, the seal detection operation is performed by calling the YoLo module via an HTTP request.

[0030] Step S30 includes: S310: Compare the third current page number with the document page number. If the third current page number is equal to the document page number, then start the seal detection. Preferably, if the third current page number is equal to the number of pages in the document, it indicates that the header and footer detection in the previous process has been completed, and the stamp detection process can be started automatically.

[0031] S320: Call the image recognition algorithm to detect the seal in the image of the file to be processed, generate the seal detection result, record the seal detection page number and the seal detection time, update the third current page number to the fourth current page number based on the seal detection page number, update the third current execution time to the fourth current execution time based on the seal detection time, and display the fourth current page number and the fourth current execution time on the user interface; Preferably, the YoLoV8 image recognition algorithm is used, and the generated stamp detection results are stored in a directory, with the stamp detection results stored in the form of the file name "yolo_001.json".

[0032] S330: The subset of bidding information, the fourth current page number, and the fourth current execution time are combined to form the fourth task dataset.

[0033] Step S40: Based on the fourth task dataset, determine whether to start text detection. If text detection is started, detect several sensitive texts in several of the images to be processed, generate several sensitive text detection results, and update the fourth task dataset to the fifth task dataset. Preferably, the text detection operation is performed by calling the OCR recognition module via an HTTP request.

[0034] Step S40 includes: S510: Compare the fourth current page number with the file page number. If the fourth current page number is equal to the file page number, then start text detection. Preferably, if the fourth current page number is equal to the number of pages in the document, it indicates that the stamp detection in the previous process has been completed, and the text detection process can be started automatically.

[0035] S520: Based on the bidding sensitive information in the bidding information subset, detect sensitive text in the file image to be processed, generate sensitive text detection results, record the text detection page number and text detection time, update the fourth current page number to the fifth current page number based on the text detection page number, update the fourth current execution time to the fifth current execution time based on the text detection time, and display the fifth current page number and the fifth current execution time on the user interface; Preferably, the sensitive bidding information includes textual information such as project name, round name, bid name, package name, bidder name, and bid type. Specifically, the OCRv4 recognition algorithm is used, and the detection results of each sensitive text are stored in a directory. The sensitive text detection results are stored in the form of the file name "ocr_001.json".

[0036] S530: The subset of bidding information, the fifth current page number, and the fifth current execution time are combined to form the fifth task dataset.

[0037] Step S50: Based on the fifth task dataset, determine whether to start blurring processing. If blurring processing starts, based on several region detection results, several seal detection results, and several sensitive text detection results, select several regions to be blurred from several files to be processed, convert several regions to be blurred into several blurred regions, so that several files to be processed are converted into several blurred files, and merge the several blurred files to generate a blurred file.

[0038] Step S50 includes: S510: Compare the fifth current page number with the file page number. If the fifth current page number is equal to the file page number, then start fuzzing.

[0039] Understandably, by progressively identifying, judging, and updating the task dataset, it ensures that no process such as the conversion and detection of multi-page, large-volume files is missed. The execution status is also intuitively displayed on the user interface during processing, facilitating user monitoring of the automated file processing progress. The combination of progressive identification and detection of the task dataset with fuzzy processing enables clear and comprehensive blind processing of large amounts of sensitive information, avoiding human access to file content. This avoids the inefficiency, inaccurate blind processing, and potential human leakage problems caused by the high reliance on manual processes in traditional methods for blind processing of bidding documents.

[0040] Please see Figure 2 The second embodiment of the present invention provides a file sensitive information processing system, applied to the file sensitive information processing method as described in the first embodiment, the system comprising: The acquisition module 10 is used to acquire a first task dataset, acquire files to be processed based on the first task dataset, convert the files to be processed into several images of files to be processed, and update the first task dataset to a second task dataset. In the acquisition module 10, the first task dataset includes a subset of bidding information, a first current page number, and a first current execution time. The subset of bidding information includes file path, file page number, blind review identifier, and bidding sensitive information. The acquisition module 10 further includes: The first unit is used to determine whether to perform blind evaluation based on the blind evaluation identifier. If blind evaluation is to be performed, the file to be processed is obtained based on the file path. The second unit is used to divide the file to be processed into several pages to be processed according to the number of pages in the file; The third unit is used to convert the file page to be processed into a file image to be processed, record the number of pages converted and the conversion time, update the first current page number to the second current page number based on the number of pages converted, update the first current execution time to the second current execution time based on the conversion time, and display the second current page number and the second current execution time on the user interface; The fourth unit is used to combine the tender information subset, the second current page number, and the second current execution time into a second task dataset.

[0041] The first detection module 20 is used to determine whether to start header and footer detection based on the second task dataset. If header and footer detection is started, it detects several header and footer regions in several files to be processed, generates several region detection results, and updates the second task dataset to the third task dataset. The first detection module 20 includes: The fifth unit is used to compare the second current page number with the document page number. If the second current page number is equal to the document page number, then header and footer detection begins. The sixth unit is used to detect the header and footer areas in the image of the file to be processed, generate area detection results, record the number of pages detected and the area detection time, update the second current page number to the third current page number based on the number of pages detected, update the second current execution time to the third current execution time based on the area detection time, and display the third current page number and the third current execution time on the user interface. The seventh unit is used to combine the tender information subset, the third current page number, and the third current execution time into a third task dataset.

[0042] The second detection module 30 is used to determine whether to start seal detection based on the third task dataset. If seal detection is started, it detects several seals in several files to be processed, generates several seal detection results, and updates the third task dataset to the fourth task dataset. The second detection module 30 includes: The eighth unit is used to compare the third current page number with the document page number. If the third current page number is equal to the document page number, then the seal detection is started. The ninth unit is used to call an image recognition algorithm to detect the seal in the image of the file to be processed, generate a seal detection result, record the seal detection page number and the seal detection time, update the third current page number to the fourth current page number based on the seal detection page number, update the third current execution time to the fourth current execution time based on the seal detection time, and display the fourth current page number and the fourth current execution time on the user interface. The tenth unit is used to combine the tender information subset, the fourth current page number, and the fourth current execution time into a fourth task dataset.

[0043] The third detection module 40 is used to determine whether to start text detection based on the fourth task dataset. If text detection is started, it detects several sensitive texts in several files to be processed and generates several sensitive text detection results, and updates the fourth task dataset to the fifth task dataset. The third detection module 40 includes: The eleventh unit is used to compare the fourth current page number with the file page number. If the fourth current page number is equal to the file page number, then text detection begins. The twelfth unit is used to detect sensitive text in the image of the file to be processed based on the sensitive information of the bidding information subset, generate sensitive text detection results, record the text detection page number and text detection time, update the fourth current page number to the fifth current page number based on the text detection page number, update the fourth current execution time to the fifth current execution time based on the text detection time, and display the fifth current page number and the fifth current execution time on the user interface; The thirteenth unit is used to combine the tender information subset, the fifth current page number, and the fifth current execution time into a fifth task dataset.

[0044] The blurring module 50 is used to determine whether to start blurring based on the fifth task dataset. If blurring starts, it selects several regions to be blurred from several files to be processed based on several region detection results, several seal detection results and several sensitive text detection results, converts several regions to be blurred into several blurred regions, so that several files to be processed are converted into several blurred files, and merges several blurred files to generate a blurred file.

[0045] The fuzzing module 50 includes: The fourteenth unit is used to compare the fifth current page number with the file page number. If the fifth current page number is equal to the file page number, then fuzzing processing begins.

[0046] The third embodiment of the present invention provides a computer, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the file sensitive information processing method as described in the first embodiment.

[0047] The fourth embodiment of the present invention provides a storage medium on which a computer program is stored, which, when executed by a processor, implements the file sensitive information processing method as described in the first embodiment.

[0048] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0049] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A method for processing sensitive information in documents, characterized in that, Includes the following steps: Obtain a first task dataset, obtain files to be processed based on the first task dataset, convert the files to be processed into several images of files to be processed, and update the first task dataset to a second task dataset; Based on the second task dataset, determine whether to start header and footer detection. If header and footer detection is started, detect several header and footer regions in several files to be processed, generate several region detection results, and update the second task dataset to the third task dataset. Based on the third task dataset, determine whether to start stamp detection. If stamp detection starts, detect several stamps in several files to be processed, generate several stamp detection results, and update the third task dataset to the fourth task dataset. Based on the fourth task dataset, determine whether to start text detection. If text detection is started, detect several sensitive texts in several files to be processed and generate several sensitive text detection results, and update the fourth task dataset to the fifth task dataset. Based on the fifth task dataset, it is determined whether to start blurring. If blurring starts, based on several region detection results, several seal detection results, and several sensitive text detection results, several regions to be blurred are selected from several files to be processed, and the several regions to be blurred are converted into several blurred regions, so that the several files to be processed are converted into several blurred files, and the several blurred files are merged to generate a blurred file.

2. The method for processing sensitive information in documents according to claim 1, characterized in that, The first task dataset includes a subset of bidding information, a first current page number, and a first current execution time. The subset of bidding information includes file path, file page number, blind review identifier, and bidding sensitive information. The steps of obtaining the files to be processed based on the first task dataset, converting the files to be processed into several images of the files to be processed, and updating the first task dataset to the second task dataset include: Determine whether to perform a blind review based on the blind review indicator. If a blind review is performed, obtain the file to be processed based on the file path. The file to be processed is divided into several pages according to the number of pages in the file; The page to be processed is converted into an image of the file to be processed. The number of pages converted and the conversion time are recorded. The first current page number is updated to the second current page number based on the number of pages converted. The first current execution time is updated to the second current execution time based on the conversion time. The second current page number and the second current execution time are displayed on the user interface. The subset of bidding information, the second current page number, and the second current execution time are combined to form the second task dataset.

3. The method for processing sensitive document information according to claim 2, characterized in that, The step of determining whether to start header and footer detection based on the second task dataset, and if header and footer detection is started, detecting several header and footer regions in several images of the files to be processed, generating several region detection results, and updating the second task dataset to the third task dataset includes: Compare the second current page number with the document page number. If the second current page number is equal to the document page number, then start header and footer detection. The header and footer areas in the image to be processed are detected, a region detection result is generated, the number of pages detected and the region detection time are recorded, the second current page number is updated to the third current page number based on the number of pages detected, the second current execution time is updated to the third current execution time based on the region detection time, and the third current page number and the third current execution time are displayed on the user interface. The third task dataset is composed of the subset of bidding information, the third current page number, and the third current execution time.

4. The method for processing sensitive document information according to claim 3, characterized in that, The step of determining whether to start seal detection based on the third task dataset, and if seal detection is started, detecting several seals in several of the images to be processed, generating several seal detection results, and updating the third task dataset to the fourth task dataset includes: Compare the third current page number with the document page number. If the third current page number is equal to the document page number, then start the seal detection. An image recognition algorithm is invoked to detect the seal in the image of the file to be processed, a seal detection result is generated, the number of pages detected and the seal detection time are recorded, the third current page number is updated to the fourth current page number based on the number of pages detected, the third current execution time is updated to the fourth current execution time based on the seal detection time, and the fourth current page number and the fourth current execution time are displayed on the user interface. The subset of bidding information, the fourth current page number, and the fourth current execution time are combined to form the fourth task dataset.

5. The method for processing sensitive document information according to claim 4, characterized in that, The step of determining whether to start text detection based on the fourth task dataset, and if text detection is started, detecting several sensitive texts in several of the images to be processed, generating several sensitive text detection results, and updating the fourth task dataset to the fifth task dataset includes: Compare the fourth current page number with the file page number. If the fourth current page number is equal to the file page number, then text detection begins. Based on the bidding sensitive information in the bidding information subset, sensitive text in the file image to be processed is detected, sensitive text detection results are generated, the text detection page number and text detection time are recorded, the fourth current page number is updated to the fifth current page number based on the text detection page number, the fourth current execution time is updated to the fifth current execution time based on the text detection time, and the fifth current page number and the fifth current execution time are displayed on the user interface; The fifth task dataset is composed of the subset of bidding information, the fifth current page number, and the fifth current execution time.

6. The method for processing sensitive document information according to claim 5, characterized in that, The specific steps for determining whether to start fuzzing based on the fifth task dataset are as follows: The fifth current page number is compared with the file page number. If the fifth current page number is equal to the file page number, then fuzzing processing begins.

7. A document sensitive information processing system, applied to the document sensitive information processing method as described in any one of claims 1 to 6, characterized in that, The system includes: The acquisition module is used to acquire a first task dataset, acquire files to be processed based on the first task dataset, convert the files to be processed into several images of files to be processed, and update the first task dataset to a second task dataset. The first detection module is used to determine whether to start header and footer detection based on the second task dataset. If header and footer detection is started, it detects several header and footer regions in several files to be processed, generates several region detection results, and updates the second task dataset to the third task dataset. The second detection module is used to determine whether to start stamp detection based on the third task dataset. If stamp detection is started, it detects several stamps in several files to be processed, generates several stamp detection results, and updates the third task dataset to the fourth task dataset. The third detection module is used to determine whether to start text detection based on the fourth task dataset. If text detection is started, it detects several sensitive texts in several of the images to be processed, generates several sensitive text detection results, and updates the fourth task dataset to the fifth task dataset. The blurring module is used to determine whether to start blurring based on the fifth task dataset. If blurring starts, it selects several regions to be blurred from several files to be processed based on several region detection results, several seal detection results, and several sensitive text detection results. It converts the several regions to be blurred into several blurred regions, so that the several files to be processed are converted into several blurred files. Then, it merges the several blurred files to generate a blurred file.

8. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the file sensitive information processing method as described in any one of claims 1 to 6.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the file sensitive information processing method as described in any one of claims 1 to 6.