Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

20 results about "Document image processing" patented technology

Document layout analysis and reconstruction method based on vector topology and conflict arbitration

The invention discloses a document layout analysis and reconstruction method based on vector topology and conflict arbitration, and relates to the technical field of document image processing, in particular to a method for carrying out accurate identification, classification and structured reconstruction on layout elements of an unstructured electronic document containing a complex vector chart. Constructing a page vector topological graph by using the bottom vector instruction and the connected topological structure; on the basis, multi-dimensional geometric features, statistical features and text mode features are introduced to jointly participate in conflict arbitration between the table and the complex graph; and applying a forced spatial exclusive constraint to text extraction by utilizing an arbitration result, and carrying out semantic classification and rearrangement in combination with a style reference and a spatial proximity relationship. According to the document layout analysis and reconstruction method, the distinguishing capacity of a table and a complex graph is improved, the attribution judgment precision of characters and a main body text in the graph is improved, and self-adaptive recognition of a title and a main body style is achieved under different layout styles.
Owner:SICHUAN ENRISING INFORMATION TECH CO LTD

PDF (Portable Document Format) document formula detection method for improving lightweight network and enhancing multi-scale features

The invention discloses an improved lightweight network and multi-scale feature enhanced PDF document formula detection method, and belongs to the technical field of document image processing and computer vision. The invention aims to solve the problems of high detection omission ratio, inaccurate positioning frame and large calculation redundancy when formulas in a PDF (Portable Document Format) complex layout are varied in size and densely mixed with text charts in the existing method. The method is realized through the following steps: firstly, inputting a PDF document image, and carrying out efficient multi-scale hierarchical feature extraction by using a YOLOv8s network improved based on a Faster Net trunk and SMSC multi-scale convolution; the exact bounding box of the formula and its category information (such as inline formula or independent formula blocks) are then directly predicted and output by a decoder. According to the method, by optimizing the network structure, the precision and robustness of formula detection in the complex format document are remarkably improved, meanwhile, the detection efficiency is guaranteed, and a reliable basis is provided for subsequent formula recognition and document understanding tasks.
Owner:SOUTHWEAT UNIV OF SCI & TECH +1

Reverse synthesis driven multi-language ancient book high-fidelity label generation method and system

The invention provides a reverse synthesis driven multi-language ancient book high-fidelity label generation method and system, and belongs to the technical field of document image processing, and the method comprises the steps: S1, constructing an OCR corpus covering multiple languages; s2, performing typesetting modeling and visual rendering by using a multilingual structure modeling algorithm based on specific writing rules of all languages; s3, constructing a degradation model for a degradation sample, and realizing accurate adjustment of degradation type, intensity and distribution through multi-parameter mapping and combined control; s4, performing automatic quality evaluation on the degraded sample, and performing dynamic self-correction and closed-loop optimization according to an evaluation result to obtain a final degraded sample; and S5, performing semantic-level labeling on the final degraded sample by using a lightweight target detection and OCR recognition network to generate a final standardized labeling result. The method disclosed by the invention provides a solid technical basis for constructing a high-quality and multi-language historical literature database.
Owner:MINZU UNIVERSITY OF CHINA

Geology field literature scatter diagram data extraction method, storage medium and device

The invention belongs to the technical field of document image processing and data mining, and particularly discloses a geoscience field literature scatter diagram data extraction method, a storage medium and equipment, and the method comprises the steps: classifying pictures extracted from geoscience field literature PDF into conventional scatter diagram pictures and table type scatter diagram pictures; positioning an independent scatter diagram region in the conventional scatter diagram picture by adopting a target detection model, segmenting the independent scatter diagram region to obtain a first type of independent scatter diagrams, extracting cells in the table type scatter diagram picture by adopting a cell segmentation method based on morphological operation and line segment intersection detection, and obtaining a second type of independent scatter diagrams; the cells are spliced with the coordinate axis area of the picture to obtain a second type of independent scatter diagrams; and inputting the first type of independent scatter diagrams and the second type of independent scatter diagrams into a scatter diagram data extraction model to obtain coordinate axis scale lines, then converting pixel coordinates of scatter points into actual data values, and outputting scatter diagram data. According to the method, the scatter diagram data can be automatically and accurately extracted from the geoscience literature.
Owner:CHINA UNIV OF GEOSCIENCES (WUHAN)

Document image processing method and device

The embodiment of the invention provides a document image processing method and device, and the method comprises the steps: obtaining an original document image, and determining a document region image from the original document image; performing local extraction processing on the document region image to obtain a first local image adjacent to the boundary and a second local image related to the first local image; and based on the similarity between the first local image and the second local image, removing the first local image from the document region image to obtain a target document image. According to the scheme, the local image corresponding to the black edge in the document region image can be accurately identified by using the content similarity, the local image is removed, the effective image content is reserved, the image quality is improved, and the accuracy and stability of a downstream task are guaranteed.
Owner:ZHUHAI KINGSOFT OFFICE SOFTWARE +2

Card and document image processing method, system and device and medium

The invention discloses a card and document image processing method, system and device and a medium, and relates to the technical field of image processing. Aiming at the problems of perspective distortion, inclined rotation and incorrect direction of a card image, the method comprises the following steps: acquiring a card or document image, and preprocessing the card or document image to generate a standard image; constructing a lightweight multi-task model comprising a lightweight convolutional neural network backbone, a corner regression network branch and a direction classification network branch, and training the lightweight multi-task model; a multi-scale shared feature map is extracted through a lightweight convolutional neural network backbone, then two-dimensional coordinates of four corners are output through a corner regression network branch, and an image direction category is output through a direction classification network branch; constructing a perspective transformation matrix by using the two-dimensional coordinates of the four angular points, and performing perspective distortion correction on the original image; and rotating and outputting the corrected image according to the direction category. According to the invention, the image processing efficiency and accuracy can be improved.
Owner:INSPUR SOFTWARE CO LTD

Document image processing method and related device

The embodiment of the invention provides a document image processing method and a related device, and the method comprises the steps: carrying out the OCR recognition of a document image, and obtaining an OCR text; the document image and the OCR text are input into a VL large model, a structured result of the document image is obtained, and the structured result comprises a classification result and a key field of the document image; based on a preset rule in a rule base, verifying the structured result to obtain a verification result; the preset rule is used for indicating a mapping relationship between the classification result and a target field; and determining a processing result of the document image based on the verification result. According to the document image processing method and device, the classification result and the key field of the document image can be obtained at the same time based on the VL large model capable of processing the image data and the text data at the same time, on this basis, the classification result and the key field are verified based on the preset rule, and the accuracy of document image processing is improved.
Owner:太保科技有限公司

Intelligent Document Detection Method and System Based on OCR and ES

The application relates to the technical field of document image processing and character recognition, and discloses a document intelligent detection method and system based on OCR and ES. The method comprises the following steps: pre-processing an input fuzzy image, extracting and enhancing edge gradient information, performing a series of processing such as gradient amplitude and direction calculation, edge continuity judgment, breakpoint connection, contour thinning, gradient reconstruction, dynamic tracking and multi-source fusion, gradually optimizing the character edge information in the image, and finally generating a clear character contour gradient graph. The method can effectively process low-quality, fuzzy or low-contrast document images, significantly improves the extraction accuracy and integrity of the character edge, and is suitable for document digitization and OCR recognition tasks in a complex background.
Owner:STATE GRID GANSU ELECTRIC POWER CO LANZHOU POWER SUPPLY CO

A spliced document image batch cutting method and system based on a split line detection and quality check

This invention relates to the field of document image processing, proposing a method and system for batch cropping of stitched images based on segmentation line detection and quality verification. The method acquires an image formed by stitching multiple pages of content row by row and column, initializes the ultra-high resolution image, performs grayscale conversion and noise reduction, performs edge detection, and enhances linear structures using directional morphological operations; based on the enhanced results, candidate line segments are obtained through line detection, and segmentation line coordinates are filtered by directional constraints and boundary exclusion, and similar segmentation lines are merged, deduplicated, and sorted; cropping regions are constructed based on ordered segmentation lines to generate candidate sub-images, and invalid slices are filtered through minimum size / area thresholds and geometric consistency verification, outputting a verification list when necessary; the numbers of sub-images that pass verification are saved, and anomaly information is recorded. Through the above scheme, the cropping consistency and processing stability of automatic splitting of stitched document images are improved, erroneous cutting is reduced, and the traceability of results is enhanced.
Owner:GUANGZHOU CLOUD COMPUTING POWER NETWORK TECH CO LTD

Multi-level text correction method and system based on document layout analysis

The application relates to the technical field of application of artificial intelligence in document image processing, and discloses a multi-level text correction method and system based on document layout analysis, which comprises the following steps: combining a multi-scale self-similarity feature algorithm and a directional frequency domain peak value feature algorithm to distinguish the type of an image to be corrected, and performing adaptive pretreatment to obtain a standardized image; extracting a text connected domain in the standardized image; using an unsupervised clustering technology to cluster each symbol in the text connected domain to obtain a plurality of word clusters; merging the word clusters to form text blocks; obtaining the minimum circumscribed quadrilateral of each text block to obtain a corresponding text box; respectively obtaining the center point coordinates of each text box; determining whether two text boxes are the same line of text; performing horizontal alignment and tilt correction on the text boxes; performing morphological regularization processing on the characters in the rotated text boxes; and outputting the corrected image and structured JSON data. The application has strong adaptability.
Owner:SICHUAN ENRISING INFORMATION TECH CO LTD

Document image processing method and device, medium, equipment and model training method

ActiveCN115376135Bwill not be disturbedvisual inspectionInstrumentsPattern recognitionEdge maps
The present disclosure relates to a document image processing method, device, medium, equipment and model training method. The processing method comprises: performing image segmentation on the document image through a hue saturation value color model to generate a three-channel space image, filtering a background area of the document image to generate a binary edge image, splicing and fusing the three-channel space image and the binary edge image to generate target input data, performing flip recognition on the target input data according to an image flip recognition model, determining a target flip category of the document image, and performing reverse flip on the document image according to the target flip category to obtain a target image. Thus, the text features in the document image are recognized through the image flip recognition model, and the document image is flipped according to the flip angle, so that the uploaded document image is not disturbed by the shooting angle, and the user can intuitively view the information in the document image.
Owner:BEIJING DINGSHIXINGJIAOYU CONSULTATION CO LTD

Image splitting and sorting method and system based on horse riding subscription A3 document

This invention provides a method and system for image segmentation and sorting of saddle-stitched A3 documents, belonging to the field of document image processing technology. The method includes determining the source type of the scanned image by extracting feature parameters, calculating a one-dimensional horizontal gradient distribution curve based on this, adaptively determining the optimal segmentation line, and precisely cropping the image to obtain sub-images. Then, based on the source type, the sub-images are subjected to interleaved sorting or order-preserving algorithms to restore the original page order. Finally, unique naming identifiers are generated for the ordered sub-images based on filename semantics and position indexes. This invention achieves automated, high-precision segmentation and sorting of saddle-stitched documents from mixed scan sources, improving processing efficiency and accuracy.
Owner:BEIJING GUOYANG HAICHUANG TECHNOLOGY CO LTD

Document image processing including tokenization of non-textual semantic elements

A method of document image processing comprises, based on at least a document page image, generating a plurality of semantic tokens that includes a plurality of word tokens and a plurality of special tokens. Each special token among the plurality of special tokens represents a non-textual semantic element of the document image, and generating the plurality of semantic tokens includes predicting, for each special token among the plurality of special tokens, a token type of the special token. The method also comprises generating, for each semantic token among the plurality of semantic tokens, a corresponding semantic token embedding among a plurality of semantic token embeddings; and applying a trained model to process an input that is based on the plurality of semantic token embeddings and a plurality of visual token embeddings based on at least the document page image to generate a semantic processing result.
Owner:IRON MOUNTAIN INC

Round seal text recognition method based on color filtering and coordinate transformation

The invention provides a circular seal text recognition method based on color filtering and coordinate transformation, and relates to the technical field of document image processing. The circular seal text recognition method based on color filtering and coordinate transformation comprises the steps that through nonlinear transformation from RGB to HSV color space, a self-adaptive mask is constructed by means of the specific chromaticity distribution characteristic of a seal, and a seal foreground and a black clutter and noise background are separated; calculating a zero moment and a first moment based on the geometric moment of the binarized image, accurately deducing the center of mass coordinate of the seal, and determining an expansion axis; establishing a reversible mapping relation between a Cartesian coordinate system and a polar coordinate system, and using a bilinear interpolation algorithm to unfold the annularly distributed text region into a rectangular image strip with significant linear features; and finally, performing text extraction on the expanded image. According to the method, the identification accuracy of the circular seal can be remarkably improved, and the method has the technical effects of high noise resistance, high calculation efficiency and high robustness.
Owner:HUA DATA TECH (SHANGHAI) CO LTD

Method and system for generating high-fidelity multilingual ancient book annotation driven by reverse synthesis

The application provides a multi-lingual ancient book high-fidelity marking generation method and system driven by reverse synthesis, and belongs to the technical field of document image processing, and comprises the following steps: S1, constructing an OCR corpus covering multiple languages; S2, based on the writing rules specific to each language, using a multi-lingual structure modeling algorithm to perform layout modeling and visual rendering; S3, constructing a degradation model for degraded samples, and precisely adjusting the degradation type, intensity and distribution through multi-parameter mapping and combination control; S4, automatically evaluating the quality of the degraded samples, and dynamically self-correcting and closed-loop optimizing according to the evaluation results to obtain the final degraded samples; S5, using a lightweight target detection and OCR identification network to perform semantic-level marking on the final degraded samples to generate the final standardized marking results. The method provides a solid technical foundation for constructing a high-quality, multi-lingual historical literature database.
Owner:MINZU UNIVERSITY OF CHINA

Document image seal removing method based on frequency perception and residual diffusion

The invention discloses a document image seal removing method based on frequency perception and residual diffusion, and belongs to the technical field of document image processing. Comprising; constructing a cascade processing model of a feature level frequency sensing network and a frequency domain constraint diffusion refiner, and taking a'seal-clean 'paired document image as a training set; inputting the image with the seal into a feature-level frequency sensing network, executing fast Fourier transform and amplitude spectrum adjustment by using an embedded frequency domain processing unit, and inhibiting the periodic frequency response of the seal to generate a rough prediction image; constructing a noise adding process based on a Markov chain, splicing a rough prediction image and a noisy residual error as conditional guidance, and predicting a high-frequency residual error image through a frequency domain constraint diffusion refiner; and superposing the predicted high-frequency residual error and the rough prediction image pixel by pixel, and reconstructing to obtain a seal-free document image. Through a frequency domain feature decoupling and generative residual error repair mechanism, the problem of content fuzziness caused by insufficient feature discrimination and regression repair in the prior art is effectively solved.
Owner:TIANJIN UNIV OF SCI & TECH

A multi-level structure information inference method and device for complex document images

The application discloses a multi-level structure information inference method and device for a complex document image, and relates to the technical field of complex document image processing. The method comprises the following steps: performing page range division by using a minimum covering rectangular frame calculation method according to a first set of page elements, and obtaining page structure information; performing column range division by using a connected subgraph solving method according to the page structure information based on geometric expansion-intersection testing, and obtaining column structure information; performing same-element merging by using an element alignment method according to the column structure information, and obtaining line structure information; performing covering element splitting by using a center point calculation method according to the line structure information, and obtaining line-internal structure information; and recognizing the element relationship between lines and columns, and constructing multi-level structure information of the complex document image. The application is a real-time and efficient multi-level result information inference method for a complex document image.
Owner:UNIV OF SCI & TECH BEIJING

A method, system, apparatus and storage medium for scanning document correction

The application discloses a kind of scanning document correction method, system, device and storage medium, wherein method includes: obtaining document image, the document image is segmented processing, obtains segmentation mask chart;Boundary line segment detection is carried out to the segmentation mask chart, obtains multiple boundary line segments;The boundary line segment is identified to obtain the type of the boundary line segment;According to the boundary line segment after identification, four boundaries of document are selected respectively one characteristic line segment;According to the characteristic line segment, affine transformation correction is carried out, and the document image after correction is obtained.The application can process various boundaries, corner missing conditions by processing line segment, and can also process documents with folding at corner points, with good applicability.In addition, the application only uses affine transformation for correction, without introducing additional deformation.The application can be widely applied in document image processing technical field.
Owner:SOUTH CHINA UNIV OF TECH

Document image processing method and device, electronic equipment and readable storage medium

The invention provides a document image processing method and device, electronic equipment and a readable storage medium, and the method comprises the steps: obtaining a to-be-processed document image with a shadow; inputting the document image with the shadow into the trained dynamic multi-region background modeling model to obtain a background prediction image of the document image with the shadow; and combining the document image with the shadow and the background prediction image, and inputting a combination result into the trained shadow diffusion removal model to obtain a shadow-free document image corresponding to the document image with the shadow. Therefore, the shadow removal process of the trained shadow removal diffusion model is guided through the accurate background information extracted by the trained dynamic multi-region background modeling model, and the accuracy of the shadow removal of the document image is improved.
Owner:PICC INFORMATION TECH CO LTD

Document image processing method and apparatus, storage medium, and electronic device

ActiveCN121095972BBiological modelsComputer graphics (images)Document image processing
This application discloses a document image processing method, apparatus, storage medium, and electronic device. The method includes: acquiring an initial document image; performing image block processing using a document image enhancement model to obtain initial image block data; performing image enhancement processing on each initial image block data to obtain reference enhanced image block data; determining first reference enhanced image block data of the quality check failure type and second reference enhanced image block data of the quality check success type; performing local supplementary enhancement processing on the first reference enhanced image block data to obtain third reference enhanced image block data; performing sliding window fusion processing on the second and third reference enhanced image block data to obtain a target enhanced document image; and performing character recognition processing to obtain a target character sequence corresponding to the initial document image and a target character confidence level corresponding to the target character sequence. Thus, this application improves the character recognition accuracy of the initial document image.
Owner:BEIJING QIHOOD TECHNOLOGY CO LTD