Archived file inspection method based on image processing

Through image processing-based methods, file preprocessing and compliance judgment are carried out, which solves the problem of manual detection difficulties in file archiving, and achieves efficient and accurate compliance judgment and unified detection results.

CN120371775APending Publication Date: 2025-07-25QINGDAO METRO GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510210914.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the process of file archiving, manual inspection is difficult and compliance cannot be accurately judged, resulting in slow progress in archiving work and poor record uniformity, and the inability to effectively carry out multi-factor fusion processing.

Method used

Through image processing methods, file preprocessing is performed to obtain relevant parameters, select appropriate processing methods, and make compliance judgments based on detection results and file parameters, and make secondary judgments through correction results to ensure the accuracy and uniformity of detection results.

Benefits of technology

It realizes efficient document processing and archiving, can accurately judge file compliance, reduce communication time, and improve the accuracy of detection results and data uniformity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371775A_ABST
    Figure CN120371775A_ABST
Patent Text Reader

Abstract

The invention discloses an archived file inspection method based on image processing. The method comprises the following steps: preprocessing a file to obtain a file type, and correspondingly selecting a subsequent processing mode and detection content; performing detection processing on the file, and performing compliance judgment in combination with a detection result and parameter information of the file; correcting the detection result after the detection processing, and performing compliance judgment based on the correction result; according to the method, efficient document processing and archiving can be achieved, multi-element fusion processing is achieved, the uniformity of stored records is high, and the compliance of the document can be accurately judged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital image processing, and particularly relates to compliance detection of PDF and picture-based archival files based on technologies such as digital image processing. More particularly, it relates to a method for inspecting archival files based on image processing. Background Art

[0002] Digital image processing is a method and technology for processing images by a computer to remove noise, enhance, restore, segment, extract features, etc. Image processing technology is widely used in multiple fields. Its core is to analyze, enhance, identify, or reconstruct images through algorithms to solve practical problems. Document recognition refers to the process of automatically recognizing and converting the text, tables, graphics, etc. in paper documents or images into editable and searchable digital information by using computer vision, image processing, and artificial intelligence technologies, or to identify and judge the attributes of documents for convenient subsequent processing. Its core is to extract structured data (such as text, tables, signatures, etc.) in the document through technologies such as optical character recognition (OCR), image segmentation, and text detection, and convert it into a machine-readable format (such as a text file, Excel table, etc.). Document recognition not only includes the recognition of text, but also involves functions such as document layout analysis, table extraction, handwritten character recognition, and multi-language support. Its goal is to achieve the automated processing, storage, retrieval, and analysis of documents, and it is widely used in fields such as office automation, file management, financial bill processing, and medical record digitization.

[0003] When filing files, it is necessary to perform compliance detection on the scanned files, and only the files that meet the requirements of specifications and content can be filed. During manual detection, it is difficult to obtain some parameters. For PDF files, pictures need to be manually extracted for judgment, which will generate a large number of additional files. After the detection, it is difficult to form a problem report with a unified specification, and it is impossible to accurately locate the problems for modification, resulting in slow progress of the filing work and poor unity of the saved records.

[0004] In the prior art, the image recognition of archived documents mainly relies on computer vision and deep learning technologies. Through steps such as image preprocessing, feature extraction, object detection and recognition, combined with a rule engine or a machine learning model for compliance judgment. Specifically, it includes steps such as binarizing, denoising, angle correction, and region segmentation of the image. After preprocessing the image, different categories and contents of parts are formed. When extracting features from the image, it includes operations such as multi-element recognition of drawing lines, contours, text, and graphics by using edge detection, OCR, and shape recognition, etc., so as to finally realize the compliance inspection of whether it meets standards such as size, proportion, and annotation. However, for a large number of current complex document information, although single-element recognition and inspection can be carried out, the elements are not effectively classified, the use of various element data is relatively single, and no multi-element fusion processing is carried out, so that accurate compliance judgment cannot be made. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide an inspection method for archived documents based on image processing, which can achieve high-efficiency document processing and archiving, multi-element fusion processing and high unity of saved records, and can accurately judge the compliance of documents.

[0006] The present invention provides an inspection method for archived documents based on image processing, including the following steps carried out in sequence;

[0007] (1) Obtain the incoming file, preprocess the file to obtain the relevant parameters of the file, and based on the relevant parameters of the file, select the subsequent processing method and the content to be detected after obtaining the file type;

[0008] (2) Carry out detection processing on the file to obtain the detection result of the detected file, combine the detection result with the parameter information of the file itself for compliance judgment. If the judgment result passes, proceed to the next step, otherwise return to step (1), and feedback the non-pass information and the specific problems existing in the file;

[0009] (3) Correct the detection result after the detection processing, and based on the corrected result, carry out compliance judgment. If the judgment result passes, proceed to the next step, otherwise return to step (1), and feedback the non-pass information and the specific problems existing in the file;

[0010] (4) After the detection is completed, return the detection result, return the file specification information and the detection pass flag, and archive the documents that pass the inspection.

[0011] Among them, the file types obtained in step (1) include PDF files and picture files.

[0012] Among them, when the file type of the input file is a picture file, in the step (2), the file is detected and processed to obtain the detection result of the detected file, which specifically includes the following steps:

[0013] (A2.1) Perform picture specification information detection to form a picture type detection result; among them, the picture specification information includes the length and width of the picture, DPI, and compression format;

[0014] (A2.2) Further determine whether the picture is a drawing file. If not, use the picture type detection result to perform compliance judgment and enter step (A2.4); if so, enter step (A2.3);

[0015] (A2.3) Further detect whether the drawing picture contains uncropped white edges, correct the white edges, and then form a picture type detection result for compliance judgment;

[0016] (A2.4) Obtain a judgment result based on the compliance judgment to form a preliminary judgment result.

[0017] Among them, in the step (A2.3), if the background color of the drawing picture is blue, perform grayscale processing and binaryzation processing on the R channel value, and then judge whether there is a continuous white area in the edge part of the picture to judge whether it contains uncropped white edges.

[0018] Among them, in the step (A2.1), the compression format is determined by reading the basic picture data. The picture file format with the compression format of JPEG is JPG, otherwise it is unqualified; the length and width are directly determined by reading the pixels of the picture; the DPI is obtained by respectively obtaining the DPI information in the two directions of the X and Y axes. If there is separate DPI information, this DPI information is directly used for judgment. If the information obtained respectively is a fraction, it is converted into a floating point number before judgment.

[0019] Among them, when the file type of the input file is a PDF file, in the step (2), the file is detected and processed to obtain the detection result of the detected file, which specifically includes the following steps:

[0020] (B2.1) After parsing the data stream of the input PDF file, traverse each page to obtain the rectangular boundary of the page, and obtain the width and height information of the page according to the width and height of the boundary;

[0021] (B2.2) Obtain the page basic information of the PDF file, and then extract the pictures contained in the PDF file from each page and record the number of pictures;

[0022] (B2.3) Determine whether the number of images on the PDF file page meets the compliance criteria. If it meets, proceed to the next step; if the number of images on a certain PDF page is not unique, return non-compliance information and do not perform image information detection.

[0023] (B2.4) Conduct image specification information detection to form an image type detection result. At the same time, read the EXIF data flag in the image to determine whether there is rotation information in the detected image. If there is no rotation information, proceed to the next step; if there is rotation information, return non-compliance information and also return the specific length and width information of the image for comparison with the length and width of the PDF page.

[0024] (B2.5) Combine the basic information of the PDF file page and the returned image detection results to comprehensively judge the compliance of the PDF file; if there is a large difference in the length and width between the image and the PDF on the same page, the page is non-compliant.

[0025] Among them, the image specification information detection in step (B2.4) is carried out in the same way as steps (A2.1)-(A2.4).

[0026] Among them, in step (3), the detection result after detection and processing is corrected, and the compliance judgment is based on the corrected result, which specifically includes the following steps:

[0027] (3.1) Determine the length and width by reading the pixels of the image, and then along the length and width of the image, collect the paired endpoints where the pixel values of the pixel points exceed the pixel threshold in the vertical and parallel directions respectively according to the ratio.

[0028] (3.2) Connect the endpoints in sequence to form an effective pixel area, calculate the effective pixel area of the effective pixel area, and calculate the image area of the image according to the length and width of the image; calculate the area percentage of the effective pixel area in the image area.

[0029] (3.3) Calculate the number of all effective pixel points that exceed the pixel threshold in the effective pixel area, and calculate the number of all pixel points in the image area of the image according to the length and width of the image; then calculate the number percentage of the effective pixel points in the number of all pixel points in the image area.

[0030] (3.4) Set the first and second area thresholds and the first and second pixel number thresholds, where the first area threshold and the first pixel number threshold are respectively greater than the second area threshold and the second pixel number threshold.

[0031] (3.5) Compare the area percentage and the number percentage with the first and second area thresholds and the first and second pixel number thresholds respectively, and determine the compliance of the image according to the comparison results.

[0032] Among them, step (3.5) specifically includes the following steps:

[0033] (3.5.1) When the area percentage is greater than the first area threshold and the quantity percentage is greater than the first pixel quantity threshold, the detection result after detection and processing is not corrected, and it is determined that the picture is compliant.

[0034] (3.5.2) When the area percentage is less than the second area threshold or the quantity percentage is less than the first pixel quantity threshold, it is determined that the picture is non-compliant, and the information of non-pass and the specific problems existing in the file are fed back.

[0035] (3.5.3) In other cases, the detection result after detection and processing is corrected. After meeting the above area threshold and pixel quantity threshold judgment conditions, the subsequent steps are entered.

[0036] The method for inspecting archived files based on image processing according to the present invention, compared with the prior art, can achieve:

[0037] (1) To solve the problems of complex and time-consuming processes for archiving reviewers during the review process, unify the problem feedback format, verify the correctness of the archived file format. During the process, the splitting of PDF and the extraction of file information will not generate additional process files. The detection result format is unified, and the existing problems can be clearly feedback for the uploader to modify, reducing the time consumption of communication between the uploader and the reviewer. At the same time, some key inherent attributes of the file are obtained, which is convenient for problem adjustment and final archiving.

[0038] (2) Use the method of judging by area and pixel quantity to correct and achieve compliance judgment, so as to achieve the purpose of secondary compliance judgment, with higher result accuracy and higher data unity. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a schematic flow chart of the method for inspecting archived files based on image processing;

[0040] Figure 2 is a schematic diagram of logical judgment processing based on file types. DETAILED DESCRIPTION OF THE INVENTION

[0041] The following details the specific implementation of the present invention. It is necessary to point out here that the following implementation is only for further illustration of the present invention and cannot be understood as a limitation on the protection scope of the present invention. Some non-essential improvements and adjustments made by those skilled in the art to the present invention based on the above content of the present invention still fall within the protection scope of the present invention.

[0042] The present invention provides a method for inspecting archived files based on image processing, and its specific implementation is as shown in the attached Figure 1-2 figure, where Figure 1It is a schematic flow diagram of an archived file inspection method based on image processing, Figure 2 and it is a schematic diagram of logical judgment processing based on file types. Next, a specific introduction to the archived file inspection method based on image processing will be given.

[0043] The present invention provides an archived file inspection method based on image processing, and its implementation steps are as shown in the appendix Figure 1 shown. Specifically, the archived file inspection method based on image processing includes the following steps carried out in sequence.

[0044] First, obtain the incoming file, preprocess the file to obtain the relevant parameters of the file, and based on the relevant parameters of the file, obtain the file type and then correspondingly select the subsequent processing method and the content to be detected. Among them, the relevant parameters of the file include the data of the file itself and the file type data, and the class file type can be obtained according to the file type data. The data of the file itself is usually used to describe the attributes, content and characteristics of the file, and these parameters can better understand, manage and process the file. The file types include PDF files and picture files, as shown in the appendix Figure 2 shown as a schematic diagram of the logical judgment processing of file types.

[0045] Secondly, perform detection processing on the file to obtain the detection result of the detected file, and combine the detection result and the parameter information of the file itself to perform compliance judgment. If the judgment result passes, enter the next step; otherwise, return, and at the same time feedback the non-passing information and the specific problems existing in the file. Specifically, the detection processing of the file includes:

[0046] 1. When the incoming file type is a picture file, perform picture specification information detection to form a picture type detection result. Among them, the picture specification information includes the length and width, DPI, and compression format of the picture; when further judging that the picture is a drawing file, it is necessary to further detect whether the drawing picture contains uncropped white edges, and after correcting the white edges, perform compliance judgment. In a preferred method, if the background color of the drawing picture is blue, perform gray-scale processing and binary processing on the R channel value, and then judge whether there is a continuous white area in the edge part of the picture to judge whether it contains uncropped white edges.

[0047] Among them, the compression format is determined by reading the basic data of the picture. Only the picture file format with the compression format of JPEG is JPG, otherwise it is unqualified; the length and width are directly determined by reading the pixels of the picture; the DPI is obtained by respectively obtaining the DPI information in the two directions of the XY axis. If there is a separate DPI information, directly use this DPI information for judgment. If the information obtained separately is a fraction, convert it to a floating point number before making a judgment. If the conditions are met, it is initially judged as compliant.

[0048] It should be noted that if it is a drawing file to be detected, the specification information to be detected includes three basic information: the format, length and width, and DPI (scanning resolution) of the picture. In addition, it is also necessary to detect whether the drawing picture contains uncropped white edges, and perform compliance judgment after correcting the white edges. Among them, in the drawing picture, if the background color is blue, it actually means that the R value in the RGB three channels is relatively low. After gray processing and binarization through the R channel value, the non-drawing part will become very obvious. At this time, by judging the continuous white area existing in the edge part of the picture, it can be judged whether there is a white edge. In practice, it is a comparison diagram of the original picture and the effect picture of the drawing with white edges after gray and binarization processing according to the R channel. It can be seen that the effect after processing is relatively obvious.

[0049] 2. When the input file type is a PDF file, perform PDF file specification detection. Specifically, after parsing the input PDF data stream, traverse each page, obtain the rectangular boundary of the page, obtain the width and height information of the page according to the width and height of the boundary, and perform detection in combination with the subsequent picture information. Specifically, it includes:

[0050] Obtain the page basic information of the PDF file (such as width, height and number of pages), and then extract the pictures contained in the PDF file from each page, and record the number of pictures. Among them, the pictures are saved in memory and no new files will be generated;

[0051] Judge whether the number of pictures in the PDF file page meets the compliance conditions. If it meets, enter the next step; if the number of pictures in a certain PDF page is not unique, return unqualified information and do not perform picture information detection;

[0052] After extracting the pictures from the PDF file, perform detection according to the same picture specification information detection steps as above;

[0053] At the same time, in addition to the same detection content as ordinary pictures, it is also necessary to additionally detect whether the picture has a selection mark. If the picture has rotation information, it is unqualified. Specifically, read the EXIF data flag in the picture, judge whether the detected picture has rotation information. If not, enter the next step; if there is rotation information, return unqualified information and at the same time return the specific length and width information of the picture for comparison with the length and width of the PDF page;

[0054] Then, combining the basic information of the pages of the PDF file and the returned image detection results, comprehensively judge the compliance of the PDF file; if there is a large difference in the length and width between the image and the same page of the PDF, then this page is unqualified. When the detection results are returned, clearly feedback what problems exist in a certain page or the images in the page. For a qualified file, return the qualified information and the page specifications of the file. For an unqualified file, return the unqualified points of the file, specifically what problems exist in a certain image on a certain page of the PDF, so as to ultimately ensure that the results can accurately feedback the content that needs to be modified in the detected file, facilitating subsequent modification.

[0055] On this basis, the present invention further corrects the detection results after detection processing, and makes a compliance judgment based on the corrected results, thereby realizing a secondary judgment of compliance, with higher result accuracy and higher data unity, which is also part of the important inventive content of the present invention. When specifically implemented, correct the detection results after detection processing, and make a compliance judgment based on the corrected results. If the judgment result passes, enter the next step; otherwise, return and feedback the non-pass information and the specific problems existing in the file.

[0056] Specifically, when detecting the picture specification information, the areas shown in the picture, due to problems such as clarity issues, format conversion issues, or information missing, will all affect the detection results of the picture, thus affecting the final compliance judgment. Therefore, on this basis, the present invention further corrects the detection results after the detection process. First, determine the length and width by reading the pixels of the picture, and then, along the length and width of the picture, collect in proportion the paired endpoints whose pixel values of the pixel points exceed the pixel threshold in the vertical and parallel directions respectively. The proportion can be selected according to the pixel situation of the picture. For example, it can be collected every 20 pixels; the pixel threshold can also be selected according to the situation. For example, the endpoints that meet the conditions are those with pixel values exceeding 25, and those lower than 25 continue to be collected in the vertical and parallel directions until a pixel point value exceeding 25 appears, which is used as an endpoint. Then, connect the endpoints in sequence to form an effective pixel area, calculate the effective pixel area of the effective pixel area, and calculate the image area of the picture according to the length and width of the picture; calculate the area percentage of the effective pixel area in the image area; then, calculate the number of all effective pixel points that exceed the pixel threshold in the effective pixel area, and calculate the number of all pixel points in the image area of the picture according to the length and width of the picture; calculate the number percentage of the number of effective pixel points in the number of all pixel points in the image area. After that, set the first and second area thresholds and the first and second pixel number thresholds, where the first area threshold and the first pixel number threshold are respectively greater than the second area threshold and the second pixel number threshold. Compare the area percentage and the number percentage with the first and second area thresholds and the first and second pixel number thresholds respectively. When the area percentage is greater than the first area threshold and the number percentage is greater than the first pixel number threshold, no further correction is made to the detection results after the detection process, and it is determined that the picture is compliant; when the area percentage is less than the second area threshold or the number percentage is less than the first pixel number threshold, it is determined that the picture is non-compliant, and the specific problems existing in the feedback failure information and the file are given; in other cases, the detection results after the detection process are corrected. After meeting the above area threshold and pixel number threshold judgment conditions, then enter the subsequent steps.

[0057] Finally, when the detection is completed, return the detection results, return the file specification information and the detection pass flag, and archive the documents that pass the inspection.

[0058] Although, for purposes of illustration, exemplary embodiments of the present invention have been described, those skilled in the art will understand that various changes in form and detail, such as modifications, additions, and substitutions, can be made without departing from the scope and spirit of the invention disclosed in the appended claims. All such changes should fall within the scope of the appended claims of the present invention, and each part of the product and each step in the method claimed by the present invention can be combined together in any combination. Therefore, the description of the embodiments disclosed in the present invention is not intended to limit the scope of the present invention, but to describe the present invention. Accordingly, the scope of the present invention is not limited by the above embodiments, but is defined by the claims or their equivalents.

Claims

1. An inspection method for archived files based on image processing, characterized in that, It includes the following steps carried out in sequence; (1) Obtain the incoming file, preprocess the file to obtain the relevant parameters of the file, and based on the relevant parameters of the file, select the subsequent processing method and the content to be detected corresponding to the file type; (2) Perform detection processing on the file to obtain the detection result of the detected file, make a compliance judgment by combining the detection result and the parameter information of the file itself. If the judgment result passes, proceed to the next step; otherwise, return to step (1) and feedback the non-passing information and the specific problems existing in the file; (3) Correct the detection result after the detection processing, and make a compliance judgment based on the corrected result. If the judgment result passes, proceed to the next step; otherwise, return to step (1) and feedback the non-passing information and the specific problems existing in the file; (4) After the detection is completed, return the detection result, return the file specification information and the detection pass flag, and archive the documents that pass the inspection.

2. The method according to claim 1, wherein: The file types obtained in step (1) include PDF files and picture files.

3. The method according to claim 2, characterized in that: When the file type of the incoming file is a picture file, in step (2), perform detection processing on the file to obtain the detection result of the detected file, which specifically includes the following steps: (A2.1) Perform picture specification information detection to form a picture type detection result; among them, the picture specification information includes the length and width, DPI, and compression format of the picture; (A2.2) Further determine whether the picture is a drawing file. If not, use the picture type detection result to make a compliance judgment and proceed to step (A2.4); if so, proceed to step (A2.3); (A2.3) Further detect whether the drawing picture contains uncropped white edges, correct the white edges, and then form a picture type detection result and make a compliance judgment; (A2.4) Obtain the judgment result based on the compliance judgment to form a preliminary judgment result.

4. The method according to claim 3, wherein: In step (A2.3), if the background color of the drawing picture is blue, perform gray-scale processing and binary processing on the R channel value, and then judge whether there is a continuous white area in the edge part of the picture to judge whether it contains uncropped white edges.

5. The method according to claim 4, wherein: In step (A2.1), the compression format is determined by reading the basic data of the picture. The picture file format with the compression format of JPEG is JPG, otherwise it is unqualified; the length and width are directly determined by reading the pixels of the picture; the DPI is obtained by respectively obtaining the DPI information in the two directions of the XY axis. If there is a separate DPI information, directly use this DPI information for judgment. If the information obtained respectively is a fraction, convert it to a floating point number before making a judgment.

6. The method according to claim 2, characterized in that: When the file type of the incoming file is a PDF file, in step (2), perform detection processing on the file to obtain the detection result of the detected file, which specifically includes the following steps: (B2.1) Parse the data stream of the incoming PDF file and traverse each page to obtain the rectangular boundary of the page, and obtain the width and height information of the page according to the width and height of the boundary; (B2.2) Obtain the page basic information of the PDF file, and then extract the pictures contained in the PDF file from each page and record the number of pictures; (B2.3) Determine whether the number of images on the PDF file page meets the compliance conditions. If it meets, proceed to the next step; If the number of images on a certain PDF page is not unique, return non-compliance information and do not perform image information detection; (B2.4) Conduct image specification information detection to form image detection results. At the same time, read the EXIF data flag in the image to determine whether there is rotation information in the detected image. If not, proceed to the next step; if there is rotation information, return non-compliance information and also return the specific length and width information of the image for comparison with the length and width of the PDF page; (B2.5) Combine the basic information of the PDF file page and the returned image detection results to comprehensively judge the compliance of the PDF file; if there is a large difference in the length and width between the image and the PDF on the same page, then this page is non-compliant.

7. The method according to claim 6, characterized in that: In step (B2.4), the image specification information detection is carried out in the same way as in steps (A2.1)-(A2.4).

8. The method according to any one of claims 1 to 7, characterized in that: In step (3), the detection results after detection and processing are corrected, and compliance judgment is made based on the corrected results, which specifically includes the following steps: (3.1) Determine the length and width by reading the pixels of the image, and then along the length and width of the image, collect in pairs the endpoints where the pixel values of the pixel points exceed the pixel threshold in the vertical and parallel directions respectively according to the ratio; (3.2) Connect the endpoints in sequence to form an effective pixel area, calculate the effective pixel area of the effective pixel area, and calculate the image area of the image according to the length and width of the image; Calculate the area percentage of the effective pixel area in the image area; (3.3) Calculate the number of all effective pixel points that exceed the pixel threshold in the effective pixel area, and calculate the number of all pixel points in the image area of the image according to the length and width of the image; then calculate the number percentage of the number of effective pixel points in the number of all pixel points in the image area; (3.4) Set the first and second area thresholds and the first and second pixel number thresholds, where the first area threshold and the first pixel number threshold are respectively greater than the second area threshold and the second pixel number threshold; (3.5) Compare the area percentage and the number percentage with the first and second area thresholds and the first and second pixel number thresholds respectively, and determine the compliance of the image according to the comparison results.

9. The method according to claim 8, wherein: Step (3.5) specifically includes the following steps: (3.5.1) When the area percentage is greater than the first area threshold and the number percentage is greater than the first pixel number threshold, no further correction is made to the detection results after detection and processing, and the image is determined to be compliant; (3.5.2) When the area percentage is less than the second area threshold or the number percentage is less than the first pixel number threshold, it is determined that the image is non-compliant, and feedback the non-passing information and the specific problems existing in the file; (3.5.3) In other cases, correct the detection results after detection and processing. After meeting the above area threshold and pixel number threshold judgment conditions, then proceed to the subsequent steps.