Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

125 results about "Document recognition" patented technology

Intelligent document recognition is a new technology that promises to transform the way businesses handle document processing. An Intelligent document recognition system analyzes the content of the document that it receives, and looks for certain keywords that match its database of business terms.

Knowledge question and answer method based on multiple agents and heterogeneous data sources

The invention discloses a knowledge question-answering method based on multiple agents and heterogeneous data sources, which comprises the following steps of: collecting an original document, identifying data, processing the format of the data, and respectively realizing the construction of a vector database and a graph database; obtaining a standard question and answer data set, and generating expansion questions by utilizing partial sub-graphs of the knowledge graph to complete construction of a basic question pool; obtaining user questions, and performing preliminary retrieval in the basic question pool; for problems needing deep retrieval, according to the types and characteristics of the problems, the routing agent flexibly calls required tools or distributes tasks to different retrieval agents; and the answer agent integrates the obtained retrieval results to generate a final answer. According to the method, a multi-agent cooperation mechanism is utilized, the intention of the user can be accurately recognized, dynamic information retrieval is carried out in different types of data sources, and meanwhile efficient and rapid intelligent question and answer services are provided by constructing the basic question pool.
Owner:ZHEJIANG UNIV

Training method, device and equipment of archive file intelligent identification large model and medium

The invention discloses a training method, device and equipment for a large intelligent recognition model of archive files and a medium, and relates to the technical field of document recognition. The training method comprises the following steps: constructing a self-supervised diffusion model for first-stage training: carrying out random mask processing on image samples to generate mask image samples, respectively inputting the mask image samples into an image encoder to extract high-dimensional information, and further enhancing the discrimination of an attention map by using a token selection module, the weight of task related parameters is dynamically adjusted through an attention refocusing mechanism, the perceptual ability of the model to a task target is improved, a text encoder embedded by an empty text is combined to serve as condition input of a diffusion model, and the encoder is optimized by using generation feedback of the diffusion model; and constructing a second-stage fine-tuning Qwen-vl large model: freezing the image encoder trained in the first stage, and finely tuning the Qwen-vl large model by using a small number of samples. According to the method, the visual reasoning and fine-grained sensing capabilities of a large archive identification model in a complex scene are realized, and the generalization and precision of archive identification are improved.
Owner:CHENGDU TECH UNIV

Vehicle after-sales maintenance document analysis and maintenance knowledge acquisition system and method based on AD-RAG

The invention relates to an AD-RAG-based vehicle after-sales maintenance document analysis and maintenance knowledge acquisition system and method, and the system comprises a document analysis module which carries out the layout analysis and document recognition of after-sales maintenance document information, and obtains structural data; the knowledge organization module is used for establishing an association relationship between the text data and the structural relationship in the structured data to obtain a maintenance knowledge graph; the knowledge retrieval enhancement module is used for analyzing and expanding the vehicle fault query request based on an AD-RAG model to obtain an expanded query request, and querying in a maintenance knowledge graph based on the expanded query request to obtain document fragment information; and the task decomposition and reasoning optimization module analyzes and decomposes the extended query request to obtain at least one fault sub-problem sequence, and performs fault reasoning through a preset fault reasoning model based on the fault sub-problem sequence and the document fragment information to obtain a fault maintenance suggestion result. The retrieval requirement can be better met, and the obtaining efficiency of the maintenance knowledge can be improved.
Owner:AIDONG SUPER AI

Zero sample template inference and document structured recognition method and device

The invention relates to the cross technical field of computer vision and natural language processing, and particularly provides a zero sample template inference and document structured recognition method and device, and the method comprises the following steps: S1, document image collection and preprocessing; s2, performing layout sensing partitioning and position coding; s3, priori or example information is constructed and injected; s4, performing cross-modal fusion and expression construction; s5, performing automatic format analysis and field slot filling; s6, generating a cue word-free extraction instruction; s7, performing field area parallel character recognition; s8, performing semantic verification and result standardization; and S9, outputting the structured field-value data. Compared with the prior art, the method has the advantages that the dependence of a traditional method on template making, cue word writing and large-scale sample training can be avoided, the flexibility, accuracy and online speed of document structured recognition are remarkably improved, and the method has good intelligence and rapid adaptation capacity and is suitable for diversified document recognition scenes.
Owner:INSPUR SOFTWARE CO LTD

Medical document intelligent identification method and system based on OCR (Optical Character Recognition)

The invention discloses an OCR-based medical document intelligent identification method and system, and relates to the technical field of document identification, and the method comprises the steps: collecting a medical document image through a mobile terminal, generating a binary image based on the collected image through U-Net in combination with multi-scale feature fusion and an attention mechanism, and cutting the binary image; based on the cut binary image, text information is extracted through an OCR model, and a structured field is extracted according to the typesetting rule and geometric distribution of the medical document; and performing deterministic rule judgment and risk assessment on the structured data. According to the method, the image calculation complexity is reduced through standard graying processing, the text region feature extraction precision is improved by fusing a U-Net structure of a CBAM attention mechanism, effective fusion and noise suppression of multi-scale features are realized in combination with Attention Gate, text direction correction is realized in combination with Hough transform, and the text detection robustness and recognition accuracy are improved.
Owner:SHALLBRIGHT HEALTHTECH CO LTD

Image recognition method and recognition system

The invention relates to the field of image recognition, discloses an image recognition method and an image recognition system, and systematically solves the core pain point of traditional document recognition through optical-topology fusion processing and a dynamic resource allocation mechanism. The image recognition method is composed of an acquisition module, a grid module, a phase module, a setting module and a distribution module, pixel brightness is calculated and coded based on document RGB data, and a two-dimensional coding matrix is generated; constructing a geometric correction grid, forming a nonlinear constraint field, and enhancing the anti-deformation capability of the image; generating spiral phase light waves in the constraint field by using a spatial light modulator, and generating a time-varying phase map; determining a local topology index by detecting the number of phase jump times; and extracting the closed boundary region as a character block, and generating an analysis result file. The system breaks through traditional limitation, improves image geometric correction precision, feature extraction sensitivity and boundary judgment accuracy, efficiently completes document analysis, and is suitable for scenes such as document digitization and information retrieval.
Owner:JIANGSU GUANGGUANG INFORMATION SYSTEM CO LTD

Verifiable large model retrieval enhancement generation system and method based on evidence chain

The invention relates to the technical field of natural language processing, in particular to a verifiable large model retrieval enhancement generation system and method based on an evidence chain, and the method comprises the steps: receiving an initial query, recognizing the fuzziness and information gap of the initial query in combination with an associated retrieval document, and generating a supplementary query set; based on the initial query, the supplementary query and the corresponding retrieval document, generating candidate answers with references and verifying the information supportability of the candidate answers; for the candidate answers passing the verification, extracting support information and constructing a hierarchical attribution mapping relation; and integrating the information to form to-be-evaluated information, and if the current verified to-be-evaluated information meets a preset sufficiency condition, integrating the generated preliminary answer and the to-be-evaluated information to synthesize a target answer. The method effectively overcomes the defects that a traditional RAG system is fragmented in information integration, has one-sided fuzzy query and answer, is low in attribution efficiency and excessively depends on retrieval content, and has the advantages of answer comprehensiveness, verifiability and deployment lightweighting.
Owner:JIANGNAN UNIV +2

A general method and system for structured document recognition based on deep learning

This invention relates to the field of image recognition technology, and more particularly to a general-purpose structured document recognition method and system based on deep learning. The method includes: acquiring document image information and preprocessing the acquired images; inputting a standardized document image into a text detection network to locate text instances in the image and saving the location information to a file; extracting text images from the document image based on the detected text instance location coordinates and inputting them into a text recognition network for recognition, saving the recognition result after the location information; using a key text extraction network to classify the text entities in the recognition result, removing non-key types, and then saving the classification result after the recognition result; and performing structured processing on the key text extraction result and displaying it. This invention can extract key text information from complex document images, enabling intelligent document reading, and is applicable to various types of documents.
Owner:HUAZHONG UNIV OF SCI & TECH

Document data entry method, electronic device, storage medium and program product

The invention discloses a document data entry method, electronic equipment, a storage medium and a program product, and relates to the technical field of document recognition, and the document data entry method comprises the following steps: obtaining a to-be-recognized document containing at least one to-be-recognized page; the document to be recognized is recognized through the optical character recognition technology, a first recognition result and a target page are obtained, and the target page is a page to be recognized containing non-text elements; identifying the target page through the target large language model to obtain a second identification result; and inputting the first recognition result and the second recognition result into the target form according to the similarity between the form field of the preset target form and the first recognition result and the second recognition result, and obtaining the input target form. Through cooperative work of the OCR and the large language model, synchronous and efficient recognition of text and non-text information is achieved, and the accuracy and the automation level of document recognition and document data entry are improved in combination with an intelligent matching entry mechanism.
Owner:YILINYUN (SHENZHEN) TECH CO LTD

Complex document recognition method based on layout analysis and OCR (optical character recognition) and medium

The invention discloses a complex document recognition method based on layout analysis and OCR (optical character recognition) and a medium. Obtaining a to-be-recognized document image in real time, and performing preprocessing operation through a preprocessing module to obtain a to-be-recognized standard document image; processing the to-be-identified standard document image through a layout analysis module, and obtaining at least one target area image, a target area image type and an area vision matrix feature in combination with a layout analysis strategy method; performing text analysis on each target area image through a multi-mode OCR engine module according to the type of each target area image to obtain each text semantic analysis result and an area semantic matrix feature; obtaining a current visual text fusion feature through a multi-modal attention fusion module; and generating and feeding back a complex document identification output result according to the current visual text fusion feature. The problem that the accuracy is poor due to the fact that irregular layouts cannot be processed is solved, and the accuracy and flexibility of complex layout recognition are improved.
Owner:ZHEJIANG BAORONG MEDIA TECH (ZHEJIANG) CO LTD

Document analysis inspection-free method and system, electronic equipment and storage medium

The invention relates to the technical field of document processing, and discloses a document analysis inspection-free method and system, electronic equipment and a storage medium, and the method comprises the following steps: inputting a document, identifying layout elements, and adding element-level inspection-free marks to the layout elements meeting a preset credible condition; extracting text, image and path information, and recombining the text, image and path information into structured memory data according to a reading sequence; executing multi-source information fusion and rule matching, and adding a structure-level inspection-free mark to a chapter structure division result meeting a preset credible condition; attribute labeling: adding an attribute-level inspection-free mark to the text block meeting a preset credible condition; and executing post-processing operation, adding a processing-level inspection-free mark to a post-processing result meeting a preset credible condition, and storing the post-processing result as a standard structured document with the inspection-free mark. According to the method, identification errors caused by prejudice or limitation of a single model are avoided, the identification accuracy and the processing efficiency can be improved, manual intervention is reduced, and the automation level is improved.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

Software development-oriented security processing method and device, equipment and medium

The invention relates to the technical field of data security, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a security processing method, device and equipment oriented to software development and a medium. Obtaining an architecture design document to identify potential safety hazards and generate design improvement suggestions; generating a code based on the business logic description and the design improvement suggestion, and completing security detection and repair to obtain a processed code and a code repair record; performing security test on the processed code and recording a test result; collecting data of exception identification, hidden danger identification, code detection and repair and security test to update the security knowledge base; and generating a security analysis report based on the demand exception list, the design improvement suggestion, the code repair record and the test result. According to the invention, through a security identification and restoration process from demand to test, early discovery of security problems, linkage processing and knowledge self-updating are realized.
Owner:PING AN TECH (SHENZHEN) CO LTD

Visual language large model-based document identification and structuring method

The invention discloses a document recognition and structuring method based on a visual language large model. The method comprises the following steps that S1, an input access layer receives a PDF or image document, and a page preprocessing and rendering layer unifies the resolution ratio and geometric parameters and performs denoising; s2, the heterogeneous recognition layer calls multiple recognizers in parallel on the same page to generate candidate results; s3, the alignment and fusion layer completes spatial alignment and text consistency evaluation of the candidate segments in a unified coordinate system to form a single main result; according to the method, the heterogeneous recognizers are called in parallel, a comprehensive scoring mechanism of space alignment, text consistency and model reliability is combined, fusion judgment is carried out on multi-source candidate results, segmentation errors and recognition deviation of a single model in complex scenes of nesting tables, cross-column titles and scanning noise can be avoided, and the recognition accuracy of the multi-source candidate results is improved. Stable output can still be kept in a multi-template and high-noise environment, and semantic consistency of recognition results is improved.
Owner:SHENZHEN SHENGWEI THREAD TECHNOLOGY CO LTD

Document identification method and device, equipment, storage medium and computer program product

The invention discloses a document identification method and device, equipment, a storage medium and a computer program product. The method comprises the steps of obtaining target document text information and task condition information; the target document text information is task type document text information; inputting the target document text information and the task condition information into a pre-trained document recognition model to obtain a document recognition result output by the document recognition model; the document recognition model is obtained by performing target fine tuning on a large language model, and the target fine tuning comprises low-rank adaptive processing and regularization processing.
Owner:CHINA MOBILE COMM LTD RES INST +1

Multi-modal mixed document OCR (Optical Character Recognition) and structured extraction method

The invention relates to the technical field of content extraction, in particular to a multi-modal hybrid document OCR (Optical Character Recognition) and structured extraction method, which comprises the following steps of: acquiring an image text region bounding box and classifying a style, extracting a font or stroke sequence to generate a character positioning structure, dividing paragraph and sentence groups to classify semantic fields, and calculating a field matching relationship to generate structural mapping. According to the method, logic mapping is constructed through character two-dimensional coordinate sorting and paragraph contours, complete reconstruction of a page structure is enhanced, semantic field categories are extracted by using syntactic density of sentence paragraph division and inter-paragraph features, the accuracy of field classification is improved, and the method has the advantages of being simple in structure, convenient to operate and high in practicability. A field mapping relation is established through Jaccard similarity and part-of-speech consistency analysis between a head word and a standard field keyword, a field path index and a structure node link are clarified, field semantic affiliation and structure position output are unified, and the document structure reduction degree and field extraction accuracy are improved.
Owner:HANGZHOU JINGSHENG HANGXING TECH CO LTD

Integrated test method and device and storage medium

The embodiment of the invention provides an integrated testing method and device and a storage medium. In the embodiment of the invention, the AI large model, the MCP protocol and the multiple test frameworks of the execution layer are fused, so that an integrated test process with the cooperation of multiple test types is realized; the AI large model analyzes the test requirement document, recognizes multiple test requirements, is responsible for generating multiple test cases from the multiple test requirements, further converts the generated multiple test cases into multiple types of test scripts, calls a corresponding test framework to execute the test script corresponding to each test type through the MCP service of the execution layer, and performs the test on the test scripts. An integrated process from test case generation and script conversion to a test process is realized, and collaborative execution of multiple types of test scripts is realized. In addition, the integrated test framework further comprises a scheduling layer which is used for responding to the output of the AI large model, calling the MCP service of the execution layer, unloading the function of scheduling and managing the MCP service of the execution layer from the AI large model, and improving the reasoning efficiency of the AI large model.
Owner:BEIJING 58 INFORMATION TTECH CO LTD

Method and device for improving RAG recall effect

The invention provides a method and device for improving an RAG recall effect, and belongs to the technical field of computers, and the method comprises the following steps: file input and typesetting structure analysis: identifying a typesetting unit for an input file, extracting text content in the typesetting unit, and generating associated data of typesetting and content; constructing a two-dimensional relation graph: performing semantic segmentation based on the typesetting units, and extracting a logic relation of the typesetting units; constructing a two-dimensional relation graph of the semantic relation and the typesetting relation; and multi-dimensional information fusion recall: receiving user query and performing semantic analysis, recalling similar semantic slices from a semantic community and a typesetting community, executing double-graph cross validation, dynamically adjusting weights, and calculating and obtaining a final recall result. According to the method, the knowledge base construction mode of the RAG is optimized from the perspective of typesetting, the multi-dimensional relation between text semantics and typesetting logic is fused, information association in the knowledge base is more comprehensive, and the retrieval recall can be based on the semantic similarity and the typesetting logic at the same time, so that the recall effect is remarkably improved.
Owner:KYLIN CORP

Ocr-based medical document intelligent recognition method and system

The application discloses an OCR-based medical document intelligent identification method and system, relates to the technical field of document identification, and comprises the following steps: collecting a medical document image through a mobile terminal, generating a binary image based on the collected image through a U-Net combined with multi-scale feature fusion and an attention mechanism, and performing clipping; based on the clipped binary image, extracting text information through an OCR model, and extracting structured fields according to the layout rules and geometric distribution of the medical document; and performing deterministic rule judgment and risk assessment on the structured data. The application reduces the image calculation complexity through standard gray scale processing, improves the text region feature extraction accuracy through the U-Net structure combined with the CBAM attention mechanism, realizes effective fusion of multi-scale features and noise suppression in combination with the Attention Gate, realizes text direction correction in combination with the Hough transform, and improves the text detection robustness and recognition accuracy.
Owner:SHALLBRIGHT HEALTHTECH CO LTD

PDF (Portable Document Format) document content identification method and device, equipment and storage medium

The invention provides a PDF (Portable Document Format) document content identification method and device, equipment and a storage medium. The method comprises the following steps: acquiring an access link of an unanalyzed document; downloading the target PDF document from the object storage service according to the access link; when the content region corresponding to the page type in the target PDF document is a text region, dividing the content region into a text region, a table region and an image region according to the page type corresponding to each content region; and when the content region is a text region, extracting a native text character sequence from the text region, and performing similarity calculation on the native text character sequence to generate a semantic coherent paragraph text. According to the method, the sentences with similar semantics in the text region are automatically divided into the same text block based on the cosine similarity, so that the text content with coherent and complete semantics is analyzed from the document, the problem that text paragraphs are broken after document recognition is solved, and the semantic coherence of the document content is effectively improved.
Owner:SHENZHEN ISSMART SCI & TECH CO LTD

Intelligent OCR (Optical Character Recognition) data extraction method, equipment and medium

The invention discloses an intelligent OCR data extraction method and device and a medium, and belongs to the technical field of information processing. The method comprises the steps of receiving a to-be-processed PDF document uploaded by a user, and performing initialization processing on the to-be-processed PDF document; converting the to-be-processed PDF document into a high-resolution image, removing a watermark based on a color threshold value, and intercepting an ROI region corresponding to key information according to a preset coordinate; calling a Paddle OCR (Optical Character Recognition) model to recognize the ROI, and returning an original recognition result containing textbox coordinates, a recognition text and confidence; and post-processing the original identification result, and outputting a JSON format identification result. According to the method, watermark removal and ROI accurate positioning are achieved, interference factors in the to-be-processed PDF document recognition process are reduced, the key information recognition accuracy is remarkably improved, meanwhile, high-concurrency document processing is conducted through a multi-thread parallel framework, document processing efficiency and processing performance are improved, and document processing time consumption is reduced.
Owner:INSPUR ZHUOSHU BIG DATA IND DEV CO LTD

An evidence chain-based verifiable large model retrieval enhancement generation system and method

The application relates to the technical field of natural language processing, in particular to an evidence chain-based verifiable large model retrieval enhancement generation system and method, which comprises the following steps: receiving an initial query, identifying the ambiguity and information gap of the initial query in combination with associated retrieval documents, and generating a supplementary query set; generating a cited candidate answer based on the initial query, the supplementary query and the corresponding retrieval document, and verifying the information supportability of the cited candidate answer; extracting supportable information from the verified candidate answer and constructing a hierarchical attribution mapping relationship; integrating the above information to form to-be-evaluated information; if the to-be-evaluated information currently verified satisfies a preset sufficiency condition, integrating a generated preliminary answer and the to-be-evaluated information to synthesize a target answer. The application effectively solves the defects of the traditional RAG system, such as information integration fragmentation, one-sided query answering, inefficient attribution and excessive dependence on retrieval content, and has the advantages of comprehensive answer, verifiability and lightweight deployment.
Owner:JIANGNAN UNIV +2

Document identification method and device

The invention discloses a document identification method and device, and relates to the field of artificial intelligence. The method comprises the following steps: acquiring a to-be-identified document; performing optical character recognition on the document to obtain a first recognition result; performing multi-mode identification on the document to obtain a second identification result; and performing text synthesis on the first recognition result and the second recognition result to output a structured recognition result. According to the method, the OCR and the multi-modal recognition are combined, compared with single OCR, the recognition accuracy is higher, and the structured recognition result can be output through text synthesis.
Owner:CHINALCO DIGITAL (CHENGDU) TECHNOLOGY CO LTD

Foreign affairs certificate recognition method and system based on multi-modal fusion

This application relates to the field of image recognition technology, specifically to a method and system for recognizing diplomatic documents based on multimodal fusion. The method includes: using semantic segmentation to obtain text recognition regions and portrait regions from diplomatic document images; performing edge detection and rectangle fitting on the text recognition regions to determine the text ratio of each fitted rectangle and the text adhesion degree of each image block; matching pixels on both sides of the portrait region, analyzing the symmetry of each matching point with respect to the facial symmetry line of the portrait region, and determining a local enhancement factor; determining the image gain requirement based on the pixel ratio of the text recognition region to the portrait region in each image block, and correcting the contrast cropping threshold of the image enhancement algorithm to enhance the diplomatic document image; and recognizing the text content and portrait content in the enhanced diplomatic document image. This application aims to enhance diplomatic document images and increase the reliability of diplomatic document recognition.
Owner:BEIJING DEXUN AVIATION SERVICE CO LTD

Unstructured document recognition method and system

This invention relates to the field of document recognition technology, and more particularly to a method and system for recognizing unstructured documents. The method includes: acquiring a PDF document to be recognized; segmenting the PDF document to form a row dataset; dividing the row dataset into regions to obtain text regions and table regions; and using a recognition method to recognize data in the text regions and table regions respectively, obtaining data in the text regions and data in the table regions respectively. The purpose of this invention is to solve the problem of low accuracy in recognizing text and tables in unstructured documents using existing technologies.
Owner:CRRC QINGDAO SIFANG CO LTD

A PDF document content recognition method, device, equipment and storage medium

The application provides a PDF document content recognition method and device, equipment and a storage medium, the method comprises the following steps: obtaining an access link of an unanalyzed document; downloading a target PDF document from an object storage service according to the access link; when the content area corresponding to the page type in the target PDF document is a text area, the content area is divided into a text area, a table area and an image area according to the page type corresponding to each content area; when the content area is a text area, the original text character sequence is extracted from the text area, and the semantic coherent paragraph text is generated by similarity calculation on the original text character sequence. Based on the cosine similarity, the application automatically divides the sentences with similar semantics in the text area into the same text block, so that the document analysis obtains semantic coherent and complete text content, solves the problem of broken text paragraphs after document recognition, and effectively improves the semantic coherence of the document content.
Owner:SHENZHEN ISSMART SCI & TECH CO LTD

A script classification method, apparatus and electronic device

Embodiments of this application disclose a script classification method, apparatus, and electronic device. The method includes: constructing a script knowledge base, a script intent base, and a script category base based on internal documents and script datasets; obtaining relevant documents from the script knowledge base, relevant intents from the script intent base, and relevant candidate classes from the category base based on the script to be classified; constructing a thought chain based on the relevant documents, relevant intents, and candidate classes; and inputting the thought chain into a target large language model for analysis to obtain the target category corresponding to the script to be classified. This application achieves intelligent classification of scripts generated by enterprise operations and maintenance. When a script is input, some candidate classes are provided, and an automated intent description is generated for the input script. The reasoning ability of the large language model itself and relevant documents are used to identify the intent description associated with the input script, thereby improving both the efficiency and accuracy of script classification.
Owner:WEBANK (CHINA)

Document identification and checking method and device based on large model

The invention provides a document recognition and checking method and device based on a large model. The method comprises the steps that a document to be recognized is obtained, the document is converted into a picture format, the obtained document picture is input into a pre-trained visual large model, and recognized information is output. Inputting information recognized by the visual large model into a language large model, writing field information needing to be extracted through a cue word project, guiding the language large model to perform keyword extraction on the input information, and outputting the field information needing to be extracted in a structured mode; configuring a checking rule, including constructing a checking script and a checking rule knowledge base; inputting the output data into a large checking model, constructing a dynamic double-engine retrieval mechanism by the large checking model in combination with an RAG retrieval enhancement generation technology, and checking the input information; and generating a checking result, and performing visual display. According to the invention, the accuracy and robustness of recognition of various documents can be improved.
Owner:TUGUAN (TIANJIN) DIGITAL TECH CO LTD

Document identification method and computer program product

PendingCN121808227ASolve the problem of insufficient recognition accuracyeasy to handleNatural language data processingDocument recognitionRapid processing
The invention discloses a document recognition method and a computer program product, and relates to the technical field of data processing.The document recognition method comprises the steps that the document complexity of a target document is recognized, and a text analysis strategy matched with the document complexity is determined; the text analysis strategy comprises at least one of a lightweight model analysis strategy, a large model analysis strategy and a lightweight collaborative large model analysis strategy; and analyzing the target document according to the text analysis strategy to obtain a document identification result of the target document. According to the method, the problem that the recognition precision of the mixed elements in the complex document is insufficient is effectively solved, low-complexity content can be rapidly processed locally by establishing a complexity evaluation and strategy matching mechanism, high-complexity content can obtain sufficient computing resources, efficiency and precision of medium-complexity content are balanced through a collaborative mechanism, and the recognition precision of the mixed elements in the complex document is improved. Therefore, the accuracy and the real-time performance of recognizing the complex document are greatly improved, and meanwhile, the resource utilization rate of the edge end and the cloud end is also greatly improved.
Owner:HANGZHOU ROBAM APPLIANCES CO LTD

Customs declaration document identification method and device

This invention provides a method and apparatus for recognizing customs declaration documents, relating to the field of artificial intelligence technology. The method includes: acquiring a target image of a customs declaration document to be recognized; inputting the target image and prompt words into a visual language model to obtain a first recognition result output by the visual language model; the prompt words are generated based on preset business requirements; the visual language model is obtained by performing low-rank adaptive fine-tuning of a CogVLM40B model; performing optical character recognition on the target image to obtain a second recognition result; and correcting the first recognition result based on the second recognition result to obtain a target recognition result corresponding to the customs declaration document. The customs declaration document recognition method provided by this invention improves the efficiency and accuracy of customs declaration document recognition by integrating the semantic understanding advantages of visual language models and the character accuracy advantages of optical character recognition through automated collaborative verification.
Owner:SINOTRANS +1