Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

38 results about "Document image processing" patented technology

Seal removing and document repairing method based on quantum state cooperative regulation and control

The invention discloses a seal removing and document repairing method based on quantum state collaborative regulation and control, and relates to the field of document image processing and quantum computing cross technology, the method comprises the following steps: obtaining to-be-processed information, and carrying out quantum-classical feature collaborative preparation; the character stroke continuity is guaranteed through quantum entangled state modeling and entanglement degree constraint iteration, texture decoupling is achieved through quantum wavelet transform and Gram-Schmidt orthogonalization, and adaptive filling is conducted in combination with a quantum generative adversarial network; dynamic quantum phase adjustment is used for counteracting superposition interference of stamps with different transparency, and quantum neural network noise reduction and multi-scale quantum Fourier sharpening are used for optimizing image quality; checking the repair result, if the repair result does not reach the standard, returning to the edge sharpening link to perform decoupling and filling the edge sharpening link to readjust the parameter; and for special scenes such as inclination, multi-color overprinting and ultra-thin frames, quantum rotation correction, color channel separation and boundary annihilation operator processing are used, finally, high-precision, high-naturalness and high-adaptability restoration of seal removal is achieved, and high fidelity of results is guaranteed.
Owner:SICHUAN JISU POWER TECH CO LTD

Document image processing method and device, storage medium and electronic equipment

ActiveCN121095972ABiological modelsEngineeringDocument image processing
The embodiment of the invention discloses a document image processing method and device, a storage medium and electronic equipment, and the method comprises the steps: obtaining an initial document image, carrying out the image segmentation processing through a document image enhancement model, and obtaining the data of each initial image block, performing image enhancement processing on the initial image block data to obtain reference enhanced image block data, and determining first reference enhanced image block data of a quality inspection failure type and second reference enhanced image block data of a quality inspection success type, and performing local supplementary enhancement processing on the first reference enhanced image block data to obtain third reference enhanced image block data, and performing sliding window fusion processing based on the second reference enhanced image block data and the third reference enhanced image block data to obtain a target enhanced document image. And performing character recognition processing to obtain a target character sequence corresponding to the initial document image and a target character confidence coefficient corresponding to the target character sequence. Therefore, the character recognition accuracy of the initial document image is improved.
Owner:BEIJING QIHOOD TECHNOLOGY CO LTD

PDF document intelligent identification and content extraction method based on deep learning

The invention discloses a PDF document intelligent identification and content extraction method based on deep learning, and relates to the technical field of artificial intelligence, deep learning, computer vision and document image processing, and the method comprises the steps: obtaining a positioning table region of each table in a PDF whole page image; obtaining a basic grid structure; cells with cross-row or cross-column structures are obtained; performing consistency detection and repair on the cells with the cross-row or cross-column structure by using a structure verification network to obtain a repaired table structure; and performing text recognition on each logic cell in the repaired table structure, and binding row and column position information corresponding to each logic cell to obtain table content which can be output in a preset structured format. According to the method, various forms of PDF tables such as scanners and pictures can be effectively processed, different table styles, fonts and backgrounds are adapted, the requirement on the quality of the input image is reduced, and high-precision table recognition and content extraction are ensured.
Owner:ZHONGSHAOXUAN TECHNOLOGY GROUP CO LTD

Document layout analysis and reconstruction method based on vector topology and conflict arbitration

The invention discloses a document layout analysis and reconstruction method based on vector topology and conflict arbitration, and relates to the technical field of document image processing, in particular to a method for carrying out accurate identification, classification and structured reconstruction on layout elements of an unstructured electronic document containing a complex vector chart. Constructing a page vector topological graph by using the bottom vector instruction and the connected topological structure; on the basis, multi-dimensional geometric features, statistical features and text mode features are introduced to jointly participate in conflict arbitration between the table and the complex graph; and applying a forced spatial exclusive constraint to text extraction by utilizing an arbitration result, and carrying out semantic classification and rearrangement in combination with a style reference and a spatial proximity relationship. According to the document layout analysis and reconstruction method, the distinguishing capacity of a table and a complex graph is improved, the attribution judgment precision of characters and a main body text in the graph is improved, and self-adaptive recognition of a title and a main body style is achieved under different layout styles.
Owner:SICHUAN ENRISING INFORMATION TECH CO LTD

Document image processing method and device, electronic device, and storage medium

The disclosed embodiments relate to a document image processing method and apparatus, electronic device, and storage medium, and relate to the field of image processing technology. The document image processing method includes: obtaining a document image to be processed, performing document edge detection on the document image to obtain document edges; performing straight line fitting on the document edges to determine a set of straight lines; determining four vertices of the document in the document image to be processed based on the set of straight lines; and performing a fill operation and perspective transformation on the document image to be processed based on the four vertices to obtain a document correction result. The disclosed technical solution can improve the accuracy of document correction.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

PDF (Portable Document Format) document formula detection method for improving lightweight network and enhancing multi-scale features

The invention discloses an improved lightweight network and multi-scale feature enhanced PDF document formula detection method, and belongs to the technical field of document image processing and computer vision. The invention aims to solve the problems of high detection omission ratio, inaccurate positioning frame and large calculation redundancy when formulas in a PDF (Portable Document Format) complex layout are varied in size and densely mixed with text charts in the existing method. The method is realized through the following steps: firstly, inputting a PDF document image, and carrying out efficient multi-scale hierarchical feature extraction by using a YOLOv8s network improved based on a Faster Net trunk and SMSC multi-scale convolution; the exact bounding box of the formula and its category information (such as inline formula or independent formula blocks) are then directly predicted and output by a decoder. According to the method, by optimizing the network structure, the precision and robustness of formula detection in the complex format document are remarkably improved, meanwhile, the detection efficiency is guaranteed, and a reliable basis is provided for subsequent formula recognition and document understanding tasks.
Owner:SOUTHWEAT UNIV OF SCI & TECH +1

Document image processing method, electronic equipment and storage medium

The invention provides a document image processing method, electronic equipment and a storage medium, and the method comprises the steps: inputting a document image into a trained unified model, and obtaining a plurality of different types of recognition results corresponding to each target region in the document image; for each target area, generating a prompt text corresponding to the target area according to a plurality of different types of recognition results corresponding to the target area; and inputting the prompt text of each target area and the received user question into a trained multi-modal model to obtain a document understanding result corresponding to the user question output by the multi-modal model. The method and the device are used for fully fusing information of multiple modes, performing targeted understanding on complex document images and improving the accuracy, the flexibility and the practicability of document image processing.
Owner:CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD

Document image processing method and document image processing device

This application provides a document image processing method and device. The method comprises: first, obtaining an initial document image and determining whether the initial document image contains shadows; then, if shadows exist in the initial document image, determining the normal vectors of each pixel in the initial document image; then, inputting the initial document image and the corresponding normal vectors of each pixel into at least three de-shadowing models to obtain at least three preliminary document images, each de-shadowing model having different constraints and corresponding loss functions; finally, fusing the at least three preliminary document images using a target loss function to obtain a target document image, where the target loss function is the weighted sum of the loss functions corresponding to the at least three de-shadowing models. This method improves the de-shadowing effect of the target document image and enhances its robustness, thereby resolving the problem of poor image de-shadowing effect in the prior art.
Owner:中国邮政储蓄银行股份有限公司

Reverse synthesis driven multi-language ancient book high-fidelity label generation method and system

The invention provides a reverse synthesis driven multi-language ancient book high-fidelity label generation method and system, and belongs to the technical field of document image processing, and the method comprises the steps: S1, constructing an OCR corpus covering multiple languages; s2, performing typesetting modeling and visual rendering by using a multilingual structure modeling algorithm based on specific writing rules of all languages; s3, constructing a degradation model for a degradation sample, and realizing accurate adjustment of degradation type, intensity and distribution through multi-parameter mapping and combined control; s4, performing automatic quality evaluation on the degraded sample, and performing dynamic self-correction and closed-loop optimization according to an evaluation result to obtain a final degraded sample; and S5, performing semantic-level labeling on the final degraded sample by using a lightweight target detection and OCR recognition network to generate a final standardized labeling result. The method disclosed by the invention provides a solid technical basis for constructing a high-quality and multi-language historical literature database.
Owner:MINZU UNIVERSITY OF CHINA

Image processing method and device, equipment and storage medium

The embodiment of the invention provides an image processing method and device, equipment and a storage medium. The method comprises the following steps: acquiring an original document image; determining a tripartite graph corresponding to the original document image; wherein the tripartite graph comprises a foreground graph, a background graph and an unknown region graph; based on the tripartite graph and the original document image, determining a mask graph corresponding to the original document image; and processing the mask image to obtain a target document image. According to the embodiment of the invention, by determining the tripartite graph corresponding to the original document image, determining the mask graph corresponding to the original document image based on the tripartite graph and the original document image, and then processing the mask graph to obtain the target document image, the original document image can be effectively enhanced, the image quality is improved, and the user experience is improved. And thus, the efficiency and accuracy of target document image processing are improved.
Owner:AGRICULTURAL BANK OF CHINA

Geology field literature scatter diagram data extraction method, storage medium and device

The invention belongs to the technical field of document image processing and data mining, and particularly discloses a geoscience field literature scatter diagram data extraction method, a storage medium and equipment, and the method comprises the steps: classifying pictures extracted from geoscience field literature PDF into conventional scatter diagram pictures and table type scatter diagram pictures; positioning an independent scatter diagram region in the conventional scatter diagram picture by adopting a target detection model, segmenting the independent scatter diagram region to obtain a first type of independent scatter diagrams, extracting cells in the table type scatter diagram picture by adopting a cell segmentation method based on morphological operation and line segment intersection detection, and obtaining a second type of independent scatter diagrams; the cells are spliced with the coordinate axis area of the picture to obtain a second type of independent scatter diagrams; and inputting the first type of independent scatter diagrams and the second type of independent scatter diagrams into a scatter diagram data extraction model to obtain coordinate axis scale lines, then converting pixel coordinates of scatter points into actual data values, and outputting scatter diagram data. According to the method, the scatter diagram data can be automatically and accurately extracted from the geoscience literature.
Owner:CHINA UNIV OF GEOSCIENCES (WUHAN)

Intelligent document detection method and system based on OCR and ES

The invention relates to the technical field of document image processing and character recognition, and discloses an intelligent document detection method and system based on OCR and ES. The method comprises the following steps of: preprocessing an input blurred image, extracting and enhancing edge gradient information, gradually optimizing character edge information in the image through a series of processing such as gradient amplitude and direction calculation, edge continuity judgment, breakpoint connection, contour refinement, gradient reconstruction, dynamic tracking and multi-source fusion, and obtaining the character edge information in the image. And finally, a clear character contour gradient map is generated. According to the method, low-quality, fuzzy or low-contrast document images can be effectively processed, the extraction accuracy and integrity of character edges are remarkably improved, and the method is suitable for document digitization and OCR recognition tasks under the complex background.
Owner:STATE GRID GANSU ELECTRIC POWER CO LANZHOU POWER SUPPLY CO

Document image processing method and device

The embodiment of the invention provides a document image processing method and device, and the method comprises the steps: obtaining an original document image, and determining a document region image from the original document image; performing local extraction processing on the document region image to obtain a first local image adjacent to the boundary and a second local image related to the first local image; and based on the similarity between the first local image and the second local image, removing the first local image from the document region image to obtain a target document image. According to the scheme, the local image corresponding to the black edge in the document region image can be accurately identified by using the content similarity, the local image is removed, the effective image content is reserved, the image quality is improved, and the accuracy and stability of a downstream task are guaranteed.
Owner:ZHUHAI KINGSOFT OFFICE SOFTWARE +2

Card and document image processing method, system and device and medium

The invention discloses a card and document image processing method, system and device and a medium, and relates to the technical field of image processing. Aiming at the problems of perspective distortion, inclined rotation and incorrect direction of a card image, the method comprises the following steps: acquiring a card or document image, and preprocessing the card or document image to generate a standard image; constructing a lightweight multi-task model comprising a lightweight convolutional neural network backbone, a corner regression network branch and a direction classification network branch, and training the lightweight multi-task model; a multi-scale shared feature map is extracted through a lightweight convolutional neural network backbone, then two-dimensional coordinates of four corners are output through a corner regression network branch, and an image direction category is output through a direction classification network branch; constructing a perspective transformation matrix by using the two-dimensional coordinates of the four angular points, and performing perspective distortion correction on the original image; and rotating and outputting the corrected image according to the direction category. According to the invention, the image processing efficiency and accuracy can be improved.
Owner:INSPUR SOFTWARE CO LTD

Document image processing method and related device

The embodiment of the invention provides a document image processing method and a related device, and the method comprises the steps: carrying out the OCR recognition of a document image, and obtaining an OCR text; the document image and the OCR text are input into a VL large model, a structured result of the document image is obtained, and the structured result comprises a classification result and a key field of the document image; based on a preset rule in a rule base, verifying the structured result to obtain a verification result; the preset rule is used for indicating a mapping relationship between the classification result and a target field; and determining a processing result of the document image based on the verification result. According to the document image processing method and device, the classification result and the key field of the document image can be obtained at the same time based on the VL large model capable of processing the image data and the text data at the same time, on this basis, the classification result and the key field are verified based on the preset rule, and the accuracy of document image processing is improved.
Owner:太保科技有限公司

Document image processing method and device, training sample generation method and device

The present disclosure provides a document image processing method and apparatus, and a training sample generation method and apparatus. The method comprises: determining an initial character region containing characters in a document image to be processed; optimizing the initial character region to determine the character boundaries, and determining an optimized target character region based on the character boundaries; removing the target character region from the document image to be processed, and generating a lighting image based on the document image to be processed without the target character region, wherein the lighting image is used to reflect ambient lighting information. This generates a lighting image containing real lighting information, which serves as the basis for training sample generation, resolving the current difficulty in obtaining lighting information.
Owner:BEIJING XIAOMI PINECONE ELECTRONICS CO LTD

Intelligent Document Detection Method and System Based on OCR and ES

The application relates to the technical field of document image processing and character recognition, and discloses a document intelligent detection method and system based on OCR and ES. The method comprises the following steps: pre-processing an input fuzzy image, extracting and enhancing edge gradient information, performing a series of processing such as gradient amplitude and direction calculation, edge continuity judgment, breakpoint connection, contour thinning, gradient reconstruction, dynamic tracking and multi-source fusion, gradually optimizing the character edge information in the image, and finally generating a clear character contour gradient graph. The method can effectively process low-quality, fuzzy or low-contrast document images, significantly improves the extraction accuracy and integrity of the character edge, and is suitable for document digitization and OCR recognition tasks in a complex background.
Owner:STATE GRID GANSU ELECTRIC POWER CO LANZHOU POWER SUPPLY CO

A spliced document image batch cutting method and system based on a split line detection and quality check

This invention relates to the field of document image processing, proposing a method and system for batch cropping of stitched images based on segmentation line detection and quality verification. The method acquires an image formed by stitching multiple pages of content row by row and column, initializes the ultra-high resolution image, performs grayscale conversion and noise reduction, performs edge detection, and enhances linear structures using directional morphological operations; based on the enhanced results, candidate line segments are obtained through line detection, and segmentation line coordinates are filtered by directional constraints and boundary exclusion, and similar segmentation lines are merged, deduplicated, and sorted; cropping regions are constructed based on ordered segmentation lines to generate candidate sub-images, and invalid slices are filtered through minimum size / area thresholds and geometric consistency verification, outputting a verification list when necessary; the numbers of sub-images that pass verification are saved, and anomaly information is recorded. Through the above scheme, the cropping consistency and processing stability of automatic splitting of stitched document images are improved, erroneous cutting is reduced, and the traceability of results is enhanced.
Owner:GUANGZHOU CLOUD COMPUTING POWER NETWORK TECH CO LTD

Multi-level text correction method and system based on document layout analysis

The application relates to the technical field of application of artificial intelligence in document image processing, and discloses a multi-level text correction method and system based on document layout analysis, which comprises the following steps: combining a multi-scale self-similarity feature algorithm and a directional frequency domain peak value feature algorithm to distinguish the type of an image to be corrected, and performing adaptive pretreatment to obtain a standardized image; extracting a text connected domain in the standardized image; using an unsupervised clustering technology to cluster each symbol in the text connected domain to obtain a plurality of word clusters; merging the word clusters to form text blocks; obtaining the minimum circumscribed quadrilateral of each text block to obtain a corresponding text box; respectively obtaining the center point coordinates of each text box; determining whether two text boxes are the same line of text; performing horizontal alignment and tilt correction on the text boxes; performing morphological regularization processing on the characters in the rotated text boxes; and outputting the corrected image and structured JSON data. The application has strong adaptability.
Owner:SICHUAN ENRISING INFORMATION TECH CO LTD

Document image processing method and device, medium, equipment and model training method

ActiveCN115376135Bwill not be disturbedvisual inspectionInstrumentsPattern recognitionEdge maps
The present disclosure relates to a document image processing method, device, medium, equipment and model training method. The processing method comprises: performing image segmentation on the document image through a hue saturation value color model to generate a three-channel space image, filtering a background area of the document image to generate a binary edge image, splicing and fusing the three-channel space image and the binary edge image to generate target input data, performing flip recognition on the target input data according to an image flip recognition model, determining a target flip category of the document image, and performing reverse flip on the document image according to the target flip category to obtain a target image. Thus, the text features in the document image are recognized through the image flip recognition model, and the document image is flipped according to the flip angle, so that the uploaded document image is not disturbed by the shooting angle, and the user can intuitively view the information in the document image.
Owner:BEIJING DINGSHIXINGJIAOYU CONSULTATION CO LTD

Image splitting and sorting method and system based on horse riding subscription A3 document

This invention provides a method and system for image segmentation and sorting of saddle-stitched A3 documents, belonging to the field of document image processing technology. The method includes determining the source type of the scanned image by extracting feature parameters, calculating a one-dimensional horizontal gradient distribution curve based on this, adaptively determining the optimal segmentation line, and precisely cropping the image to obtain sub-images. Then, based on the source type, the sub-images are subjected to interleaved sorting or order-preserving algorithms to restore the original page order. Finally, unique naming identifiers are generated for the ordered sub-images based on filename semantics and position indexes. This invention achieves automated, high-precision segmentation and sorting of saddle-stitched documents from mixed scan sources, improving processing efficiency and accuracy.
Owner:BEIJING GUOYANG HAICHUANG TECHNOLOGY CO LTD

Multilayer structure information inference method and device for complex document image

The invention discloses a multi-level structure information inference method and device for a complex document image, and relates to the technical field of complex document image processing. The method comprises the following steps: according to a first layout element set, performing page range division by using a minimum coverage rectangular frame calculation method to obtain page structure information; on the basis of a geometric expansion-intersection test, according to the page structure information, a connected subgraph solving method is used for carrying out column dividing range division, and column dividing structure information is obtained; according to the column structure information, performing similar element combination by using an element alignment method to obtain row structure information; according to the line structure information, performing coverage element splitting by using a central point calculation method to obtain in-line structure information; and identifying inter-line and inter-column element relationships, and constructing multi-level structure information of the complex document image. The method is a real-time and efficient multi-level result information inference method oriented to complex document images.
Owner:UNIV OF SCI & TECH BEIJING

Document image processing including tokenization of non-textual semantic elements

A method of document image processing comprises, based on at least a document page image, generating a plurality of semantic tokens that includes a plurality of word tokens and a plurality of special tokens. Each special token among the plurality of special tokens represents a non-textual semantic element of the document image, and generating the plurality of semantic tokens includes predicting, for each special token among the plurality of special tokens, a token type of the special token. The method also comprises generating, for each semantic token among the plurality of semantic tokens, a corresponding semantic token embedding among a plurality of semantic token embeddings; and applying a trained model to process an input that is based on the plurality of semantic token embeddings and a plurality of visual token embeddings based on at least the document page image to generate a semantic processing result.
Owner:IRON MOUNTAIN INC

Estimation result output apparatus, program, and estimation result output system

To provide an estimation result output apparatus, a program, and an estimation result output system configured to allow a user to efficiently confirm or correct an estimation result.SOLUTION: A document image processing server 1 includes: a document image acquisition unit 11 which receives an input of a target document image; a character string acquisition unit 12 which acquires a character string from the target document image received by the document image acquisition unit 11; a relation estimation unit 13 which estimates attribute information of the character string and correspondence relation from the position and content of the character string acquired by the character string acquisition unit 12; a result image generation unit 14 which generates an estimation result image obtained by presenting the attribute information and correspondence relation estimated by the relation estimation unit 13 on the target document image; and a correction image output unit 15 which outputs a correction image which is a correctable image of the estimation result image generated by the result image generation unit 14.SELECTED DRAWING: Figure 1
Owner:DAI NIPPON PRINTING CO LTD

Multi-level text correction method and system based on document layout analysis

The invention relates to the technical field of application of artificial intelligence in document image processing, and discloses a multi-level text correction method and system based on document layout analysis. The method comprises the steps that the type of an image to be corrected is judged by combining a multi-scale self-similarity feature algorithm and a directional frequency domain peak value feature algorithm; adaptive preprocessing is carried out to obtain a standardized image; extracting a text connected domain in the standardized image, clustering each symbol in the text connected domain by using an unsupervised clustering technology to obtain a plurality of word clusters, combining the word clusters to form text blocks, and obtaining a minimum circumscribed quadrangle of each text block to obtain a corresponding textbox; respectively acquiring a center point coordinate of each textbox, and judging whether the two textboxes are texts in the same line or not; performing horizontal alignment and tilt correction on the textbox; performing form regularization processing on the characters in the rotated textbox; and outputting the corrected image and the structured JSON data. The method has extremely high adaptability.
Owner:SICHUAN ENRISING INFORMATION TECH CO LTD

Round seal text recognition method based on color filtering and coordinate transformation

The invention provides a circular seal text recognition method based on color filtering and coordinate transformation, and relates to the technical field of document image processing. The circular seal text recognition method based on color filtering and coordinate transformation comprises the steps that through nonlinear transformation from RGB to HSV color space, a self-adaptive mask is constructed by means of the specific chromaticity distribution characteristic of a seal, and a seal foreground and a black clutter and noise background are separated; calculating a zero moment and a first moment based on the geometric moment of the binarized image, accurately deducing the center of mass coordinate of the seal, and determining an expansion axis; establishing a reversible mapping relation between a Cartesian coordinate system and a polar coordinate system, and using a bilinear interpolation algorithm to unfold the annularly distributed text region into a rectangular image strip with significant linear features; and finally, performing text extraction on the expanded image. According to the method, the identification accuracy of the circular seal can be remarkably improved, and the method has the technical effects of high noise resistance, high calculation efficiency and high robustness.
Owner:HUA DATA TECH (SHANGHAI) CO LTD

Document image processing method, device and electronic equipment

The present invention relates to the field of image processing technology, and discloses a method, device, and electronic device for processing document images. The document image processing method comprises: copying an original document image to obtain a document copy image; processing the document copy image through a mean blur algorithm to obtain a blurred image; and dividing the original document image and the blurred image to obtain a calculated image. The document image processing method of the present invention is mainly aimed at the image characteristics of the document image itself. Document images are generally simple in color, and users pay more attention to the document content in the document image. After a simple image division operation, pixels with small grayscale values ​​are smaller, and pixels with large grayscale values ​​are larger, thereby widening the grayscale value gap between the text foreground part and the white background part of the document image, and the document content in the document image can be extracted to obtain a document image with a black and white effect. The algorithm is simpler, the implementation difficulty is lower, and the document presentation effect is better.
Owner:BEIJING BAIGEFEICHI TECH LLC

Method and system for generating high-fidelity multilingual ancient book annotation driven by reverse synthesis

The application provides a multi-lingual ancient book high-fidelity marking generation method and system driven by reverse synthesis, and belongs to the technical field of document image processing, and comprises the following steps: S1, constructing an OCR corpus covering multiple languages; S2, based on the writing rules specific to each language, using a multi-lingual structure modeling algorithm to perform layout modeling and visual rendering; S3, constructing a degradation model for degraded samples, and precisely adjusting the degradation type, intensity and distribution through multi-parameter mapping and combination control; S4, automatically evaluating the quality of the degraded samples, and dynamically self-correcting and closed-loop optimizing according to the evaluation results to obtain the final degraded samples; S5, using a lightweight target detection and OCR identification network to perform semantic-level marking on the final degraded samples to generate the final standardized marking results. The method provides a solid technical foundation for constructing a high-quality, multi-lingual historical literature database.
Owner:MINZU UNIVERSITY OF CHINA

Document image seal removing method based on frequency perception and residual diffusion

The invention discloses a document image seal removing method based on frequency perception and residual diffusion, and belongs to the technical field of document image processing. Comprising; constructing a cascade processing model of a feature level frequency sensing network and a frequency domain constraint diffusion refiner, and taking a'seal-clean 'paired document image as a training set; inputting the image with the seal into a feature-level frequency sensing network, executing fast Fourier transform and amplitude spectrum adjustment by using an embedded frequency domain processing unit, and inhibiting the periodic frequency response of the seal to generate a rough prediction image; constructing a noise adding process based on a Markov chain, splicing a rough prediction image and a noisy residual error as conditional guidance, and predicting a high-frequency residual error image through a frequency domain constraint diffusion refiner; and superposing the predicted high-frequency residual error and the rough prediction image pixel by pixel, and reconstructing to obtain a seal-free document image. Through a frequency domain feature decoupling and generative residual error repair mechanism, the problem of content fuzziness caused by insufficient feature discrimination and regression repair in the prior art is effectively solved.
Owner:TIANJIN UNIV OF SCI & TECH

A deep learning-based PDF document intelligent recognition and content extraction method

The application discloses a kind of based on deep learning's PDF document intelligent identification and content extraction method, it is related to artificial intelligence, deep learning, computer vision and document image processing technical field, including: obtaining the positioning table area of each table in PDF whole page image;Obtain basic grid structure;Obtain cell with cross row or cross column structure;With the consistency detection and repair of cell with cross row or cross column structure using structure checking network, obtain the table structure after repair;Text recognition is carried out to each logical cell in the table structure after repair, and the row and column position information corresponding to each logical cell is bound, to obtain the table content that can be output as pre-set structured format.The application can effectively process scanned copy, picture and other various forms of PDF table, adapt to different table style, font and background, reduce the requirement to input image quality, ensure high-precision table recognition and content extraction.
Owner:ZHONGSHAOXUAN TECHNOLOGY GROUP CO LTD