PDF (Portable Document Format) document structured loading method based on image recognition

Through the image recognition method, the multi-dimensional feature parameters of the image in the PDF document are collected and the loading engine is adaptively selected, which solves the problem of insufficient perception of image structure in the prior art, and realizes high-quality structural reconstruction and data loading of complex PDF documents.

CN120181041AActive Publication Date: 2025-06-20BEIJING GUANGLIANDA YUNTU DREAM TECH CO LTD

Patent Information

Application Number
CN202510608694.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-20
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The existing structured loading methods of PDF documents lack the ability to sense the structure of the image content and cannot accurately identify the semantic features and layout forms of the image, resulting in problems such as table miscolumns, text misalignment, and text misalignment when processing complex PDF documents, affecting the accuracy of structure reconstruction and data loading quality.

Method used

Using an image recognition-based method, the initial image is acquired by setting a predetermined extraction scale, a preprocessing strategy is introduced for image processing, the feature parameter set of target images is collected in multiple dimensions, and the loading engine classifier is activated to determine the target engine category, and the structured loading engine is adaptively selected for structured image loading.

Benefits of technology

The accuracy of structural reconstruction and data loading quality of complex PDF documents are improved, high-dimensional modeling and adaptive structured loading of image semantic structures are realized, and the accuracy and intelligent adaptability of document structure reconstruction are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181041A_ABST
    Figure CN120181041A_ABST
Patent Text Reader

Abstract

The invention provides a PDF (Portable Document Format) document structured loading method based on image recognition, which relates to the technical field of image data processing, and comprises the following steps: setting a predetermined extraction scale of a document image based on margin information of a PDF document; obtaining an initial image of the PDF document with a predetermined coarse scale in predetermined extraction scales; introducing a preprocessing strategy to process the initial image to obtain a target image; a target feature parameter set of the target image is collected in a multi-dimensional mode, the loading engine classifier is activated to analyze the target feature parameter set, and a target engine category is determined; and carrying out structured loading on the initial image through the target engine category. The technical problem that in the prior art, due to the fact that a loading engine cannot be adaptively selected according to semantic features and layout forms of images in a PDF document, the image type PDF structure restoration effect is poor is solved, and the technical effect of improving complex PDF document structure reconstruction accuracy and data loading quality is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of image data processing, and particularly to a method for structured loading of PDF documents based on image recognition. Background Art

[0002] As a widely used electronic document format, PDF documents play an important role in digital office work, information sharing, and document storage. In order to achieve automated information extraction and systematic management, it is necessary to convert the content of PDF documents into a digitally data format that can be structured loaded, and this process is called the structured loading of PDF documents. Existing methods for structured loading of PDF documents mainly rely on text extraction tools or OCR recognition tools to identify and extract the text areas in the PDF. However, these methods generally lack the ability to perceive the structure of the image content in the document, cannot accurately identify the semantic features and layout forms of the images, and the loading strategy is fixed. As a result, when processing complex PDF documents, especially image-based PDFs, problems such as misaligned tables, misaligned text, and misaligned graphics and text are likely to occur, seriously affecting the accuracy of structure reconstruction and the quality of subsequent data loading. Summary of the Invention

[0003] This application provides a method for structured loading of PDF documents based on image recognition, which solves the technical problem that the prior art cannot adaptively select a loading engine according to the semantic features and layout forms of the images due to the lack of the ability to perceive and classify the structure of the image content in the PDF document, resulting in a poor effect of PDF structure restoration, and achieves the technical effect of improving the accuracy of structure reconstruction and the quality of data loading for complex PDF documents.

[0004] In view of the above problems, this application provides a method for structured loading of PDF documents based on image recognition. The method includes: setting a predetermined extraction scale for the document image based on the margin information of the PDF document; obtaining an initial image of the PDF document at a predetermined coarse scale in the predetermined extraction scale; introducing a preprocessing strategy to process the initial image to obtain a target image; multidimensionally collecting a target feature parameter set of the target image, and activating a loading engine classifier to analyze the target feature parameter set to determine the target engine category; and performing structured loading on the initial image through the target engine category.

[0005] Preferably, introducing a preprocessing strategy to process the initial image to obtain a target image includes: performing Gaussian blur processing on the initial image according to the preprocessing strategy to obtain a blurred image; performing pixelation processing on the blurred image according to the preprocessing strategy to obtain a pixelated blurred image; and using the pixelated blurred image as the target image.

[0006] Preferably, the multi-dimensional collection of the target feature parameter set of the target image includes: obtaining a target layering result by layering the target image, where the target layering result includes a first layer; randomly sampling the first layer to obtain a first lattice set; performing multi-dimensional feature collection on the first lattices in the first lattice set to obtain a first feature parameter set; and constructing the target feature parameter set based on the first feature parameter set.

[0007] Preferably, performing multi-dimensional feature collection on the first lattices in the first lattice set to obtain a first feature parameter set includes: collecting color features of the first lattices to obtain first color feature parameters; collecting texture features of the first lattices to obtain first texture feature parameters; collecting position features of the first lattices to obtain first position feature parameters; and constructing the first feature parameter set based on the first color feature parameters, the first texture feature parameters, and the first position feature parameters.

[0008] Preferably, performing multi-dimensional feature collection on the first lattices in the first lattice set to obtain a first feature parameter set includes: performing discrete cosine transform processing on the first lattices to obtain a first transform result; using the first DC parameter in the first transform result as the first color feature parameter, and using the first AC parameter in the first transform result as the first texture feature parameter.

[0009] Preferably, after the multi-dimensional collection of the target feature parameter set of the target image, it further includes: constructing a first neighborhood of the first lattices with the first lattice set as a constraint, where the first neighborhood includes second lattices; obtaining a second feature parameter set of the second lattices; performing aggregation processing on the first feature parameter set and the second feature parameter set to obtain a first aggregated feature parameter set of the first layer; constructing a target aggregated feature parameter set based on the first aggregated feature parameter set, and replacing the target feature parameter set with the target aggregated feature parameter set.

[0010] Preferably, the activation loading engine classifier analyzes the target feature parameter set to determine the target engine category, including: analyzing the target feature parameter set through the loading engine classifier to obtain the target engine category of the initial image; wherein, the loading engine classifier includes a first classifier and a second classifier; wherein, the first classifier is a support vector machine obtained by supervised training of a first training data set, and the first training data set includes the training color feature parameters, training texture feature parameters of the first training image, and an identifier indicating whether the first training image conforms to the picture constraint; wherein, the second classifier is a support vector machine obtained by supervised training of a second training data set, and the second training data set includes the training position feature parameters of the second training image and an identifier indicating whether the second training image conforms to the distance difference constraint, and the second training image has an image that does not conform to the picture constraint.

[0011] Preferably, the first training data set includes: obtaining a predetermined color texture weight; performing weighted calculation on the training color feature parameters and the training texture feature parameters in combination with the predetermined color texture weight to obtain a training feature value; determining whether the training feature value reaches a predetermined feature value limit to obtain a training judgment result; marking whether the first training image conforms to the picture constraint according to the training judgment result; wherein, if the training feature value reaches the predetermined feature value limit, the first training image conforms to the picture constraint; wherein, if the training feature value does not reach the predetermined feature value limit, the first training image does not conform to the picture constraint.

[0012] Preferably, the second training data set includes: randomly extracting a first lattice position and a second lattice position in the second training image and forming a training lattice pair; calculating the training lattice distance of the training lattice pair and determining whether the training lattice distance conforms to the distance difference constraint.

[0013] Preferably, performing structured loading on the initial image through the target engine category includes: matching a target loading engine in a predetermined loading engine with the target engine category as the activation constraint; performing structured loading on the initial image through the target loading engine; wherein, the predetermined loading engine includes a picture loading engine, a table loading engine, and a text loading engine.

[0014] One or more technical solutions provided in this application have at least the following beneficial effects: By setting a predetermined extraction scale of the document image based on the margin information of the PDF document, the regional perception and field of view control of the initial image extraction are realized, ensuring that the image content under different layout sizes is effectively covered, avoiding information loss or interference from redundant areas, and laying a foundation for structure perception. By obtaining the initial image of the PDF document at a predetermined coarse scale in the predetermined extraction scale, a low-overhead image is quickly generated, which is convenient for the subsequent recognition model to obtain the overall structure layout of the document, improving the image processing efficiency and the macroscopic structure perception ability. By introducing a preprocessing strategy to process the initial image to obtain a target image, the clarity and contrast of the image are improved, providing high-quality image data for structure feature extraction and OCR recognition. By multidimensionally collecting the target feature parameter set of the target image, a multidimensional quantization index describing the image structure, content type, and layout mode is constructed to realize the high-dimensional modeling of the image semantic structure. By analyzing the target feature parameter set through the activation loading engine classifier, the target engine category is determined, and a matching structured loading engine is adaptively selected. Through the target engine category, the initial image is structurally loaded, and structured operations such as text OCR, table disassembling, and graphic recognition are respectively completed, realizing the accurate restoration and unified loading of multiple content types.

[0015] In summary, the present application realizes the high-quality conversion of the PDF document from the original visual information to the semantic structure data by constructing an image recognition-driven process from boundary perception extraction to structure classification loading. It not only ensures the clarity and standardization in the image processing process but also introduces an intelligent perception and classification mechanism for the document structure, which can dynamically select the optimal loading path according to the content forms of different documents, significantly improving the accuracy, integrity, and intelligent adaptation ability of complex documents in structure reconstruction and data loading. Brief Description of the Drawings

[0016] Figure 1 It is a schematic flowchart of the method for structurally loading a PDF document based on image recognition provided by an embodiment of the present application.

[0017] Figure 2 It is a schematic flowchart of processing the initial image to obtain a target image in the method for structurally loading a PDF document based on image recognition provided by an embodiment of the present application.

[0018] Figure 3 It is a schematic flowchart of multidimensionally collecting the target feature parameter set of the target image in the method for structurally loading a PDF document based on image recognition provided by an embodiment of the present application. Detailed Embodiments

[0019] By providing a structured loading method for PDF documents based on image recognition, the embodiments of the present application solve the technical problem in the prior art that due to the lack of the ability to perceive and classify the structure of image content in PDF documents, the loading engine cannot be adaptively selected according to the semantic features and layout forms of images, resulting in poor restoration effect of the structure of image-based PDFs, and achieve the technical effect of improving the accuracy of complex PDF document structure reconstruction and the quality of data loading.

[0020] As Figure 1 shown, the embodiments of the present application provide a structured loading method for PDF documents based on image recognition, and the method includes: Step S1: Set a predetermined extraction scale for the document image based on the margin information of the PDF document.

[0021] Specifically, the margin information of the PDF document refers to the spatial distance between the text, images and the page edge in the PDF, which is used to judge the position and size of the content area. First, analyze the page layout of the PDF file, extract the left, right, top and bottom margins of the page (such as analyzing the border, text box, object coordinates, etc. using a PDF parsing library), and determine the actual content distribution range of the PDF document based on the margin information. Subsequently, set multiple different size ranges of image interception areas according to the margin information to obtain the predetermined extraction scale. By recognizing the PDF page margins, the accurate framing of the content area is realized, providing a scale basis for subsequent image acquisition, ensuring that the image covers all main content areas, and at the same time avoiding intercepting blank areas.

[0022] Step S2: Obtain the initial image of the PDF document with a predetermined coarse scale in the predetermined extraction scale.

[0023] Specifically, the predetermined coarse scale is a relatively large range or a preliminary scale setting in the predetermined extraction scale, which is used to obtain the initial image of the PDF document. It is usually set to the maximum width covering the entire effective content area for subsequent complete image generation. For example, if the page width of a PDF document is 600 pixels, the height is 800 pixels, and the margin is 50 pixels, then the coarse scale in the predetermined extraction scale can be set to a width of 500 pixels and a height of 300 pixels. Taking the coarse scale as the main parameter, call the PDF image rendering module (such as pdf2image, Poppler) to convert the specified page area into an image to obtain the initial image of the PDF document. Coarse scale extraction can retain the global structure of the document, ensure that the generated initial image has complete visual structure information, and enable subsequent recognition algorithms to have more complete context semantics.

[0024] Step S3: Introduce a preprocessing strategy to process the initial image to obtain a target image.

[0025] Specifically, the preprocessing strategy is a set of pre-set operation methods for processing the acquired initial image, such as Gaussian blur, pixelation, etc. The preprocessing strategy is introduced to process the initial image. For example, perform Gaussian blur operation on the initial image to smooth the edges and reduce noise, and then perform pixelation to divide the image into regular grids. The processed image is the target image, which has good local consistency and is suitable for subsequent image feature sampling.

[0026] Step S4: Multidimensionally collect the target feature parameter set of the target image, and activate the loading engine classifier to analyze the target feature parameter set to determine the target engine category.

[0027] Specifically, the target feature parameter set is a data set composed of multidimensional parameters such as color, texture, and position information extracted from the target image. The loading engine classifier is a classification model built based on machine learning algorithms, which can analyze the target feature parameter set and determine the target engine category according to the analysis results, such as a support vector machine model. The target engine category is the engine type used for structured loading of the initial image, such as a picture loading engine, a table loading engine, and a text loading engine, etc.

[0028] After the target image is layered and randomly sampled, feature information such as color, texture, and position coordinates is extracted from it to obtain the target feature parameter set. Then, the target feature parameter set is input into the loading engine classifier for analysis and judgment to obtain the structural loading method suitable for the image area, that is, the target engine category. Exemplarily, if the image contains a large number of equally spaced horizontal lines and numbers, the loading engine classifier determines that the content of this image is suitable for the table loading engine; if the image contains a large amount of text, the loading engine classifier can determine that the content of this image is suitable for the text loading engine. Through the acquisition and classification judgment of multidimensional features, the loading strategy has adaptability and can flexibly match the most suitable structured loading method (target engine category) according to the document content type, improving the loading accuracy and generalization ability.

[0029] Step S5: Structurally load the initial image through the target engine category.

[0030] Specifically, according to the target engine category determined in S4, activate the corresponding loading engine to parse the image content. For example, activate the table engine to detect cells and reconstruct boundaries, and activate the text engine to perform OCR recognition and paragraph segmentation. Each engine has different structural parsing capabilities and data restoration processes, and can achieve accurate structural restoration of PDF content. The classification-driven loading mechanism improves the adaptability and accuracy of the overall PDF document content recognition, especially suitable for complex PDF page structures such as mixed text and graphics, and nested tables and figures.

[0031] Further, as Figure 2As shown, step S3 includes: Step S31: Perform Gaussian blur processing on the initial image according to the preprocessing strategy to obtain a blurred image.

[0032] Step S32: Perform pixelation processing on the blurred image according to the preprocessing strategy to obtain a pixelated blurred image.

[0033] Step S33: Use the pixelated blurred image as the target image.

[0034] Specifically, call the Gaussian filtering algorithm in the image processing tool to perform convolution processing on the initial image, smooth the image details, and reduce noise interference. For example, use the GaussianBlur function of the OpenCV library to dynamically set an appropriate kernel size (such as a 3×3 or 5×5 convolution kernel) and standard deviation according to factors such as the image resolution and content complexity, perform blur processing on the initial image, blur the overly fine local structures, and retain the general structure to obtain a blurred image, making the subsequent lattice extraction more stable.

[0035] Based on the blurred image, divide the entire image into a grid according to a fixed or dynamic size (such as 32×32 pixels), divide the image into several lattice regions to obtain a pixelated blurred image. Each lattice in this image serves as the basic unit for subsequent feature extraction, improving the efficiency of image feature sampling and the regional perception ability.

[0036] After completing the Gaussian blur and pixelation processing, define the finally obtained pixelated blurred image as the target image for performing subsequent layering, sampling, and multi-dimensional feature extraction operations, improving the stability and recognition accuracy of subsequent processing.

[0037] Further, as Figure 3 shown, in step S4 of the embodiment of the present application, multi-dimensionally collect the target feature parameter set of the target image, including: Step S41: Perform layering on the target image to obtain a target layering result, where the target layering result includes a first layer.

[0038] Step S42: Perform random sampling on the first layer to obtain a first lattice set.

[0039] Step S43: Perform multi-dimensional feature collection on the first lattices in the first lattice set to obtain a first feature parameter set.

[0040] Step S44: Construct the target feature parameter set based on the first feature parameter set.

[0041] Specifically, according to criteria such as the gray-scale distribution, edge changes, and texture density of the image, image processing tools such as OpenCV or MATLAB are called to layer the target image. Exemplarily, the image gradient analysis and region merging algorithm of OpenCV is used to perform hierarchical division according to the texture density, and the target image is segmented into multiple different levels. Any one of the layering results will be used as the first layer for subsequent processing operations.

[0042] In the first-layer region, a random function is used to select several lattice blocks to form the first lattice set. For example, there are 600 lattices in the first layer. The NumPy.random.choice() is called for random sampling, and 60 lattices are selected as the first lattice set, effectively compressing the sample scale of feature extraction while maintaining the overall distribution balance of the image. For each first lattice in the first lattice set, an image processing tool is called to collect multi-dimensional features such as color features, texture features, and position features, and the collected features are combined into the first feature parameter set. Then, the collected first feature parameter set is integrated together in the form of a list or matrix, etc., to form the target feature parameter set, realizing the abstract conversion from local features to the global structure, and enabling the image structure content to be input into the classifier in a standardized and machine-readable form, improving the accuracy of classification judgment.

[0043] Further, step S43 includes: Step S431: Collect color features of the first lattice to obtain the first color feature parameter.

[0044] Step S432: Collect texture features of the first lattice to obtain the first texture feature parameter.

[0045] Step S433: Collect position features of the first lattice to obtain the first position feature parameter.

[0046] Step S434: Based on the first color feature parameter, the first texture feature parameter, and the first position feature parameter, form the first feature parameter set.

[0047] Specifically, methods such as RGB average value, color histogram, and DCT (Discrete Cosine Transform) are used to extract the pixel color values of the first lattice region as the first color feature parameter, representing the statistical or transformed features of the current lattice in the color dimension, such as color mean, histogram, or spectral components. Exemplarily, identify the color values (RGB values) of each pixel in the first lattice, and then calculate the RGB average value of all pixels in the first lattice as the first color feature parameter; or construct a color histogram, count the number of pixels in each color interval, and then perform normalization processing to obtain a multi-dimensional vector as the first color feature parameter; or use discrete cosine transform to extract the low-frequency DC component as the color tone.

[0048] Analyze the gray-scale change pattern or texture pattern of the pixels within the first lattice to collect texture features and obtain the first texture feature parameters, so as to quantify the lattice texture complexity, directionality or frequency characteristics. This process can adopt methods such as the gray-level co-occurrence matrix (GLCM) and discrete cosine transform. For example, use the gray-level co-occurrence matrix (GLCM) to calculate parameters such as contrast, energy, and correlation as the first texture feature parameters; or calculate the alternating current component of the discrete cosine transform as the first texture feature parameters.

[0049] Determine the relative position of the first lattice in the target image to obtain the first position feature parameter, which is represented in coordinate form. Exemplarily, taking the upper left corner of the target image as the coordinate origin (0, 0), with the right direction as the positive x-axis direction and the downward direction as the positive y-axis direction, record the coordinates (x, y) of the upper left corner pixel of the first lattice as the first position feature parameter.

[0050] Combine the first color feature parameter, the first texture feature parameter, and the first position feature parameter into a vector or structure in a unified format to obtain the first feature parameter set. Exemplarily, the NumPy concatenated array can be used to form the first feature parameter set in the form of (the first color feature parameter, the first texture feature parameter, x, y), or a JSON dictionary can be constructed to structurally represent the first feature parameter set. This first feature parameter set can comprehensively describe the color, texture, and position features of the first lattice, providing rich basic data for subsequent overall feature analysis of the target image.

[0051] Furthermore, step S43 further includes: Perform discrete cosine transform processing on the first lattice to obtain a first transform result; use the first direct current parameter in the first transform result as the first color feature parameter, and use the first alternating current parameter in the first transform result as the first texture feature parameter.

[0052] Specifically, the discrete cosine transform is a transformation method that converts an image from the spatial domain to the frequency domain, which can decompose the pixel gray-scale information into a set of frequency components. First, convert each first lattice region into a grayscale image to remove color interference, and then use the two-dimensional DCT algorithm to perform spectral analysis on the grayscale image to obtain a transformation matrix F(u, v) of size N×N (such as 8×8 or 16×16), denoted as the first transform result.

[0053] In the obtained first transform result F(u, v), the first direct current parameter is F(0, 0), that is, the value in the upper left corner of the transformation matrix, which reflects the average color or brightness information of the first lattice, and extract it as the first color feature parameter.

[0054] For the first AC parameter, the energy distribution of the transformation matrix can be statistically analyzed, or some high-frequency components among other values of the transformation matrix can be selected as the first texture feature parameter. The high-frequency components in the AC parameter reflect the texture and detail changes in the image.

[0055] The above steps achieve the precise extraction of the color and texture information of the lattice in the frequency domain by introducing the discrete cosine transform. Among them, the DC parameter provides a stable and reliable brightness reference, and the AC parameter quantifies the edge and texture complexity. Compared with the traditional spatial domain analysis method, this method has higher noise resistance and the ability to retain compression features, which helps to improve the feature recognition efficiency and accuracy of the loading engine classifier and provides more accurate semantic support for the structured loading strategy for the image content.

[0056] Further, after collecting the target feature parameter set of the target image in multiple dimensions, it further includes: Step S45: Construct the first neighborhood of the first lattice with the first lattice set as the constraint, where the first neighborhood includes the second lattice.

[0057] Step S46: Obtain the second feature parameter set of the second lattice.

[0058] Step S47: Perform an aggregation process on the first feature parameter set and the second feature parameter set to obtain the first aggregated feature parameter set of the first layer.

[0059] Step S48: Construct the target aggregated feature parameter set based on the first aggregated feature parameter set, and replace the target feature parameter set with the target aggregated feature parameter set.

[0060] Specifically, with the first lattice set as the constraint, according to the image coordinate system, determine the spatial adjacent area of each first lattice to form the first neighborhood (such as up, down, left, right, eight-neighborhood) including several adjacent second lattices. Perform the same feature extraction operation on the second lattices in the first neighborhood as on the first lattice, and extract the second feature parameter set from the color, texture, and position dimensions, including the second DC parameter, the second AC parameter, and the position coordinates.

[0061] Aggregate the features of each first lattice with the features of all second lattices within its neighborhood to obtain the first aggregated feature parameter set of the first layer. The aggregation method can be mean aggregation, weighted aggregation, etc. For mean aggregation, it is necessary to calculate the average value of all feature values under each dimension of the first feature parameter set corresponding to each first lattice and the second feature parameter set to obtain the mean color feature parameter, the mean texture feature parameter, and the mean position feature parameter, which are used as the first aggregated feature parameter set of the first layer. For weighted aggregation, it is necessary to perform weighted averaging on all feature values under each dimension of the first feature parameter set corresponding to each first lattice and the second feature parameter set according to the lattice distance to calculate the first aggregated feature parameter set of the first layer.

[0062] Combine the first aggregated feature parameter set of the first layer according to a structure similar to the previous target feature parameter set to form a new full-image-level feature set, that is, the target aggregated feature parameter set, and replace the original target feature parameter set with it to enhance the global perception ability of the image structure.

[0063] The above steps effectively implement the context modeling of the local structure by introducing the neighborhood aggregation mechanism, improve the integrity and accuracy of the overall feature expression, significantly enhance the robustness, context relevance, and anti-interference ability of feature extraction, provide a more stable and accurate input basis for the loading engine classifier, and improve the reliability of the structured loading classification judgment.

[0064] Further, in step S4, the activation loading engine classifier analyzes the target feature parameter set to determine the target engine category, including: Step S49: Analyze the target feature parameter set through the loading engine classifier to obtain the target engine category of the initial image, where the loading engine classifier includes a first classifier and a second classifier.

[0065] The first classifier is a support vector machine obtained by supervised training on the first training data set, and the first training data set includes the training color feature parameters, training texture feature parameters of the first training image, and the identifier indicating whether the first training image conforms to the picture constraint.

[0066] The second classifier is a support vector machine obtained by supervised training on the second training data set, and the second training data set includes the training position feature parameters of the second training image and the identifier indicating whether the second training image conforms to the distance difference constraint, and the second training image has an image that does not conform to the picture constraint.

[0067] Specifically, the obtained target feature parameter set is input into the loading engine classifier for analysis. The loading engine classifier consists of a first classifier and a second classifier, which can determine the document content format to which the image belongs according to the input feature parameters and output the corresponding target engine category. Both the first classifier and the second classifier are trained based on the support vector machine algorithm.

[0068] The first classifier is supervised and trained based on the first training data set, which includes the training color feature parameters, training texture feature parameters of the first training image, and the identifier indicating whether the first training image conforms to the picture constraint. Among them, the picture constraint is that the complexity of the image area in terms of color and texture features reaches a certain threshold, which is used to determine whether the corresponding document content of the target image is in the picture format.

[0069] After receiving the target feature parameter set, the loading engine classifier activates the first classifier. Based on the color feature parameters and texture feature parameters in the target feature parameter set, it determines whether the target image conforms to the picture constraint, that is, the color is rich and the texture is complex. If the target image conforms to the picture constraint, it directly determines that the corresponding document content of the target image is in the picture format, outputs the corresponding target engine category as the picture loading engine, and the process ends; if the target image does not conform to the picture constraint, it enters the second classifier judgment process.

[0070] The second classifier is supervised and trained based on the second training data set, which includes the training position feature parameters of the second training image and the identifier indicating whether the second training image conforms to the distance difference constraint, and the second training image has an image that does not conform to the picture constraint. Among them, the distance difference constraint is a preset distance range or threshold, which is used to determine whether the corresponding document content of the target image is in the text form or the table form.

[0071] When the first classifier determines that the target image does not conform to the picture constraint, the second classifier is activated. Based on the position features (such as coordinate differences) between different lattice point pairs within the image area, it checks whether these position features satisfy the distance difference constraint. If the target image conforms to the distance difference constraint, that is, the lattice position spacing is regular and the arrangement is neat, it directly determines that the corresponding document content of the target image is a table, outputs the corresponding target engine category as the table loading engine, and the process ends; if the target image does not conform to the distance difference constraint, that is, the lattice position arrangement is disordered, it determines that the corresponding document content of the target image is text, and outputs the corresponding target engine category as the text loading engine.

[0072] The above steps construct a loading engine decision-making scheme based on a multi-classifier collaborative discrimination mechanism, which first performs rough classification on the color and texture complexity, and then makes fine discrimination based on the structural arrangement features, realizing the automatic distinction and classification loading decision of picture-type, table-type, and text-type PDF image content, and significantly improving the accuracy and intelligent level of PDF structured processing.

[0073] Further, step S49 includes: Step S491: Obtain a predetermined color-texture weight.

[0074] Step S492: Perform weighted calculation on the training color feature parameters and the training texture feature parameters by combining the predetermined color-texture weight to obtain a training feature value.

[0075] Step S493: Determine whether the training feature value reaches a predetermined feature value limit to obtain a training judgment result.

[0076] Step S494: Mark whether the first training image conforms to the picture constraint according to the training judgment result; if the training feature value reaches the predetermined feature value limit, the first training image conforms to the picture constraint; if the training feature value does not reach the predetermined feature value limit, the first training image does not conform to the picture constraint.

[0077] Specifically, the predetermined color-texture weight is a weighted ratio of color features and texture features preset in the training stage, which is used to measure the complexity of image content. According to the analysis of the image training set, the importance weights of color and texture features are set, and these weights can be automatically adjusted by expert experience or grid search.

[0078] The color feature value (DC component obtained by discrete cosine transform) and texture feature value (AC component obtained by discrete cosine transform) extracted from each training image are respectively multiplied by the corresponding weights and then summed to obtain the training feature value corresponding to the training image.

[0079] A limit value is preset in advance as the threshold for classification judgment, that is, the predetermined feature value limit. The calculated training feature value is compared with this limit value. If the training feature value reaches the predetermined feature value limit, a mark indicating that the first training image conforms to the picture constraint is added, for example, adding a "picture" label to it; if the training feature value does not reach the predetermined feature value limit, a mark indicating that the first training image does not conform to the picture constraint is added. For example, adding a "non-picture" label to it, and the marked training images are formed into a first training data group for training the first classifier.

[0080] The above steps provide high-quality supervised samples for the training of the first classifier. Through the weighted fusion of color and texture features and the threshold determination mechanism, the classifier can more effectively distinguish picture-type document content from non-picture-type document content, enhancing the adaptability and discrimination ability of the loading engine classifier when processing complex PDF content.

[0081] Further, step S49 also includes: Step S495: Randomly extract the first lattice position and the second lattice position in the second training image, and form a training lattice pair.

[0082] Step S496: Calculate the training lattice distance of the training lattice pair, and determine whether the training lattice distance meets the distance difference constraint.

[0083] Specifically, the second training image contains images that do not meet the picture constraints. For each second training image, through a random number generation algorithm, two coordinate points are randomly determined within the position feature parameters, which are used as the first lattice position and the second lattice position respectively, and a training lattice pair is formed, obtaining multiple lattice pairs in several training images.

[0084] Use the Euclidean distance formula to calculate the distances between these training lattice pairs, obtaining multiple training lattice distances. Compare the calculated training lattice distances with the distance difference constraint to determine whether they meet the distance difference constraint. If the training lattice distance of the second training image meets the distance difference constraint, add a "table" label to the second training image. If the training lattice distance of the second training image does not meet the distance difference constraint, add a "text" label to the second training image. Form these labeled second training images into a second training data set for training the second classifier.

[0085] The above steps establish a distance feature model for judging the structural stability of the image by extracting the position relationship of the lattice pairs in the image, thereby providing a key training basis for the second classifier. Combining the previous color texture judgment mechanism enables the classifier to further distinguish between tables and texts, achieving higher-precision structured recognition and processing.

[0086] Further, step S5 includes: Match the target loading engine in the predetermined loading engines with the target engine category as the activation constraint; perform structured loading on the initial image through the target loading engine; where the predetermined loading engines include a picture loading engine, a table loading engine, and a text loading engine.

[0087] Specifically, according to the target engine category output by the loading engine classifier, a search and match is performed among the predetermined loading engines including the image loading engine, the table loading engine, and the text loading engine, and the corresponding target loading engine is selected and activated, and the initial image is structurally loaded using the target loading engine. The image loading engine reads the metadata of the image file (such as resolution, color mode, etc.), and then loads the pixel data of the image into the memory in a predetermined format (such as RGB or CMYK, etc.) to construct an image data structure that is convenient for operation. The table loading engine first detects the table lines, then uses a grid segmentation algorithm to construct cells, and converts the content in each cell into a structured form. The text loading engine extracts text through an OCR tool, and then combines natural language processing (NLP) to segment, identify titles, text, lists, etc., and reconstructs a text document with a clear structure, such as a JSON structure, an HTML paragraph, etc.

[0088] By activating the matching structured loading engine in the above steps, the content in the PDF image can be accurately parsed according to its actual semantic structure, ensuring that the image type, table type, and text type content are respectively processed in the most appropriate way, so as to achieve highly restored graphic restoration and information extraction, and improve the accuracy of document structured reconstruction and data loading quality.

[0089] In summary, the method for structurally loading a PDF document based on image recognition provided by the embodiments of the present application has the following beneficial effects: By setting a predetermined extraction scale for the document image based on the margin information of the PDF document, the regional perception and field of view control of the initial image extraction are realized, ensuring that the image content under different layout sizes is effectively covered, avoiding information loss or interference from redundant regions, and laying a foundation for structure perception. By obtaining the initial image of the PDF document with a predetermined coarse scale in the predetermined extraction scale, a low-overhead image is quickly generated, which is convenient for the subsequent recognition model to obtain the overall structure layout of the document, and improves the image processing efficiency and macroscopic structure perception ability. By introducing a preprocessing strategy to process the initial image to obtain a target image, the image clarity and contrast are improved, providing high-quality image data for structure feature extraction and OCR recognition. By multi-dimensionally collecting the target feature parameter set of the target image, a multi-dimensional quantization index describing the image structure, content type, and layout mode is constructed to realize high-dimensional modeling of the image semantic structure. By activating the loading engine classifier to analyze the target feature parameter set, the target engine category is determined, and a matching structured loading engine is adaptively selected. Through the target engine category, the initial image is structurally loaded, and structured operations such as text OCR, table disassembling, and graphic recognition are respectively completed to realize the accurate restoration and unified loading of multiple content types.

[0090] Overall, the embodiments of the present application implement a high-quality transformation of PDF documents from original visual information to semantic structure data by constructing an image recognition-driven process from boundary-aware extraction to structure classification and loading, which not only ensures the clarity and standardization in the image processing process, but also introduces an intelligent perception and classification mechanism for document structures, enabling dynamic selection of the optimal loading path according to the content forms of different documents, and significantly improving the accuracy, integrity, and intelligent adaptation ability of complex documents in terms of structure reconstruction and data loading.

[0091] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A structured PDF document loading method based on image recognition, characterized in that: include: Setting a predetermined extraction scale of the document image based on margin information of the PDF document; Acquire an initial image of the PDF document at a predetermined coarse scale in the predetermined extraction scale; Introducing a preprocessing strategy to process the initial image to obtain a target image; Collecting a target feature parameter set of the target image in multiple dimensions, and activating a loading engine classifier to analyze the target feature parameter set to determine a target engine category; Performing structured loading of the initial image through the target engine category; Wherein, introducing a preprocessing strategy to process the initial image to obtain a target image includes: Performing Gaussian blur processing on the initial image according to the preprocessing strategy to obtain a blurred image; Performing lattice processing on the blurred image according to the preprocessing strategy to obtain a lattice blurred image; The lattice blurred image is used as the target image.

2. The method for loading PDF documents in a structured manner based on image recognition according to claim 1, characterized in that: The target feature parameter set of the target image is collected in multiple dimensions, including: Layering the target image to obtain a target layering result, wherein the target layering result includes a first layer; Randomly sampling the first layer to obtain a first lattice set; Performing multi-dimensional feature collection on the first lattice in the first lattice set to obtain a first feature parameter set; The target feature parameter set is constructed based on the first feature parameter set.

3. The method for loading PDF documents in a structured manner based on image recognition according to claim 2, characterized in that: Performing multi-dimensional feature collection on the first lattice in the first lattice set to obtain a first feature parameter set includes: Collecting color characteristics of the first lattice to obtain first color characteristic parameters; Collecting texture features of the first lattice to obtain first texture feature parameters; Collecting position characteristics of the first lattice to obtain first position characteristic parameters; The first feature parameter set is formed based on the first color feature parameter, the first texture feature parameter, and the first position feature parameter.

4. The method for loading PDF documents in a structured manner based on image recognition according to claim 3 is characterized in that: Performing multi-dimensional feature collection on the first lattice in the first lattice set to obtain a first feature parameter set includes: Performing discrete cosine transform processing on the first lattice to obtain a first transformation result; The first DC parameter in the first transformation result is used as the first color feature parameter, and the first AC parameter in the first transformation result is used as the first texture feature parameter.

5. The method for loading PDF documents in a structured manner based on image recognition according to claim 2, characterized in that: After collecting the target feature parameter set of the target image in multiple dimensions, the method further includes: assembling a first neighborhood of the first lattice using the first lattice set as a constraint, wherein the first neighborhood includes a second lattice; Acquire a second characteristic parameter set of the second lattice; Aggregating the first feature parameter set and the second feature parameter set to obtain a first aggregated feature parameter set of the first layer; A target aggregated feature parameter set is established based on the first aggregated feature parameter set, and the target feature parameter set is replaced by the target aggregated feature parameter set.

6. The method for loading PDF documents in a structured manner based on image recognition according to claim 1, characterized in that: Activate the loading engine classifier to analyze the target feature parameter set to determine the target engine category, including: Analyzing the target feature parameter set by the loading engine classifier to obtain the target engine category of the initial image; Wherein, the loading engine classifier includes a first classifier and a second classifier; The first classifier is a support vector machine obtained by supervised training of a first training data group, wherein the first training data group includes training color feature parameters, training texture feature parameters of a first training image, and an identification of whether the first training image meets the image constraint; Among them, the second classifier is a support vector machine obtained by supervised training of the second training data group, the second training data group includes the training position feature parameters of the second training image and an identification of whether the second training image meets the distance difference constraint, and the second training image has an image that does not meet the picture constraint.

7. The method for loading PDF documents in a structured manner based on image recognition according to claim 6, characterized in that: The first training data set includes: Get the predetermined color texture weight; Performing weighted calculation on the training color feature parameter and the training texture feature parameter in combination with the predetermined color texture weight to obtain a training feature value; Determine whether the training characteristic value reaches a predetermined characteristic value limit, and obtain a training determination result; Marking whether the first training image meets the picture constraint according to the training judgment result; Wherein, if the training feature value reaches a predetermined feature value limit, then the first training image meets the picture constraint; If the training feature value does not reach a predetermined feature value limit, the first training image does not meet the picture constraint.

8. The method for loading PDF documents in a structured manner based on image recognition according to claim 6, characterized in that: The second training data set includes: randomly extracting a first lattice position and a second lattice position in the second training image and forming a training lattice pair; The training lattice distance of the training lattice pair is calculated, and it is determined whether the training lattice distance meets the distance difference constraint.

9. The method for loading PDF documents in a structured manner based on image recognition according to claim 1, characterized in that: The initial image is structured loaded through the target engine category, including: Matching a target loading engine among predetermined loading engines with the target engine category as an activation constraint; Performing structured loading on the initial image through the target loading engine; Wherein, the predetermined loading engine includes a picture loading engine, a table loading engine and a text loading engine.

Citation Information

Patent Citations

  • Bloody image detection classifier implementing method, bloody image detection method and bloody image detection system

    CN105760885A

  • Image segmentation method based on weak supervision multi-kernel classification optimization combination

    CN110930413A

  • Document AI system based on deep learning

    CN118470730A

  • Textile texture classification method and device, electronic equipment and storage medium

    CN118762213A

  • Charging method of intelligent robot

    CN119864922A

Cited By

  • Intelligent retrieval method and system for association of images and contents in PDF (Portable Document Format) document

    CN120407819A

  • Intelligent analysis method and device for unstructured PDF document, equipment and medium

    CN120747992A