Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1003 results about "Optical character recognition" patented technology

Optical character recognition or optical character reader (OCR) is the electronic or mechanical conversion of images of typed, handwritten or printed text into machine-encoded text, whether from a scanned document, a photo of a document, a scene-photo (for example the text on signs and billboards in a landscape photo) or from subtitle text superimposed on an image (for example from a television broadcast).

Operational research course knowledge graph construction method based on multi-source data fusion

The invention provides an operational research course knowledge graph construction method based on multi-source data fusion. The method comprises the following steps: firstly, discussing logical association of courses, majors and students, collecting data such as textbooks, exercises and teaching programs by taking operational research knowledge points as a core and utilizing technologies such as OCR (Optical Character Recognition) and crawlers, and carrying out preprocessing and manual labeling; then, an improved deep learning model is adopted for entity recognition and relation extraction, BERT + BiLSTM + CRF is adopted for entity recognition, and a dynamic context pooling enhancement model is fused to improve the capture ability of a complex knowledge boundary; bERT + BiLSTM is adopted for relation extraction, a multi-head attention mechanism is combined, and hidden logical relation mining is enhanced. And finally, constructing a multi-level knowledge network which takes knowledge points as nodes and logic relations as edges, and embedding the multi-level knowledge network into a Neo4j graph database for visualization. The map can optimize a teaching path, provides personalized learning recommendation, and is widely applied to the fields of wisdom education, knowledge retrieval and the like.
Owner:KUNMING UNIV OF SCI & TECH

Engineering document index consistency proofreading method and system based on multi-modal large model

The invention relates to an engineering document index consistency proofreading method and system based on a multi-modal large model, and the method comprises the steps: Q1. OCR detection and recognition: carrying out the optical character recognition and format analysis of a source document, converting an uploaded PDF document into a processable text message in a Markdown format, and carrying out the format discrimination of a table, a formula and a plain text; and Q2, table and formula processing: adopting a hierarchical processing strategy, intelligently selecting an optimal processing mode according to the complexity of the table, and converting table information into a descriptive long text through a language large model and cue words. According to the method, accurate, reliable and efficient document index checking service can be provided for a user, the quality and efficiency of professional document processing are remarkably improved, the efficiency and quality of knowledge graph construction are remarkably improved, a knowledge verification system capable of being evolved continuously is established, and the method is suitable for popularization and application. And a reliable technical support is provided for knowledge management and professional decision-making in a complex field.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Multi-modal document content cross-platform analysis system

The invention provides a multi-modal document content cross-platform analysis system, and relates to the technical field of data processing, and the system comprises a document preprocessing module which is used for receiving multi-source heterogeneous document input data, and generating a preprocessed document through format conversion, page segmentation, noise reduction and optical character recognition; the feature extraction module is used for extracting three types of features, including spatial layout features, semantic features and logic structure features, based on the preprocessed document; according to the method, the multi-source heterogeneous document is cooperatively processed, accurate extraction, calibration and association of features in cross-platform analysis are realized, finally structured data are generated, and the accuracy, consistency and cross-platform applicability of multi-modal document analysis are improved.
Owner:XIAMEN CITIZEN DATA SERVICE CO LTD +1

Multi-dimensional data intelligent retrieval matching method and system for graphic and text features

The invention provides an intelligent retrieval matching method and system for multidimensional data of image-text features, and relates to the technical field of icon image retrieval. Comprising the following steps: extracting image features, text content and semantic features of an icon image by using a convolutional neural network, an image segmentation attention mechanism network, a converter optical character recognition model and a bidirectional semantic understanding model, constructing the extracted features into heterogeneous feature tensors, and performing singular value decomposition to obtain icon feature fingerprint vectors; and constructing a multi-level index based on locality sensitive hashing, realizing rapid retrieval, calculating visual, text and semantic similarities in combination with a deep metric learning model, weighting according to variances and discrimination coefficients of similarity features to obtain a comprehensive similarity score, and outputting a retrieval result with the highest similarity.
Owner:BEIJING YIZHUANG TECHNOLOGY INNOVATION CO LTD

Automated and semi-automated extraction of data from tables and graphs in scientific literature

A system and method for automatically extracting data from tables and graphs in scientific literature, particularly in life sciences and healthcare, is presented. The invention employs a hybrid approach combining computer vision, natural language processing (NLP), and a large language model (LLM)-based data extraction module. A neural network enhances accuracy by providing contextual information. The system performs table structure detection, optical character recognition (OCR), Vision Transformers (ViTs) for text recognition, and graph-to- table conversion. An LLM refines the extracted data using advanced prompt engineering. A user interface enables data review, validation, and iterative refinement. The system employs a novel two-stage approach: de-rendering tables and graphs into a machine-readable format, followed by interpretation and user-defined mapping. A knowledge graph enhances extraction by resolving ambiguities and inferring relationships. Designed for scalability and continuous improvement, the invention significantly enhances data extraction efficiency and accuracy, accelerating research and knowledge discovery.
Owner:EVIDENCE PRIME SP ZOO

Circuit netlist generation method supporting deep learning model multi-stage reasoning

The invention discloses a circuit netlist generation method supporting deep learning model multi-stage reasoning, and belongs to the technical field of electronic circuit design and automation, and the method comprises the steps: obtaining a circuit diagram image through optical scanning, carrying out the preprocessing of the image, carrying out the recognition of a component through a deep learning model YOLOv5, and determining the boundary frame coordinates and class information. And further performing port fine identification on the component, and determining the position and the direction of the port. And eliminating character interference by using an optical character recognition technology, performing wire recognition, including Hough straight line detection and jumper processing, and determining starting point and ending point coordinates of the wire. And matching the identification results of the components, the ports and the wires, constructing a topological structure of a circuit diagram, finally generating a circuit netlist, and performing error elimination and JSON format output. According to the method, automatic netlist generation of the analog circuit diagram can be realized, the analog circuit diagram file is converted into the netlist file with device function annotations, and the efficiency and quality of circuit design and simulation are improved.
Owner:HUNAN UNIV OF SCI & TECH

Multi-algorithm collaborative intelligent PDF (Portable Document Format) document analysis system

The invention belongs to the technical field of document analysis, and particularly relates to a multi-algorithm collaborative PDF document intelligent analysis system which comprises a preprocessing module used for recognizing the type of a PDF document, performing layout correction on a scanning document PDF and uniformly adjusting the scanning document PDF into a vertical storage format; the layout analysis module is used for identifying page element categories based on a target detection algorithm, and processing element superposition and separation problems through a merging de-duplication or priority discarding strategy; and the text processing module is used for analyzing all places with characters in the PDF based on the OCR, outputting the characters and corresponding textbox coordinates, and dividing the contents of the text blocks in combination with layout analysis. The system can extract characters of the PDF of the scanned copy, distinguish element types such as titles, texts, tables and formulas, solve the problem of confusion of cross-page tables, column texts and characters with similar shapes, and can completely reserve the content structure of the document.
Owner:BEIJING HUAYUN WORLD TECH CO LTD

Methods for Automatically Generating a Training Dataset for Training an Optical Recognition Model for Reading Street Signs

Various embodiments include methods for generating image datasets for training an artificial intelligence machine learning (AI / ML) optical character recognition (OCR) model. Image processing may be performed on a plurality of roadway images to identify street signs within the images and generate a dataset of sign images categorized into sign variants of the same shape, color, pictogram, and characters. An OCR model may process sign images to obtain OCR results for images of each sign variant. An aggregation process may be performed on the OCR results for all sign images within each sign variant to identify a ground truth OCR result for each sign variant. The ground truth OCR result may be used to automatically label all sign images of each sign variant to produce an OCR model training dataset. The produced training dataset may then be used to retrain the initial AI / ML OCR model and / or train other AI / ML OCR models.
Owner:QUALCOMM INC

Contract auditing method, device and equipment based on large model and storage medium

The invention discloses a contract auditing method and device based on a large model, equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: analyzing a target contract file based on an optical character recognition technology and a natural language processing technology, and generating a target structured analysis fragment corresponding to the target contract file; constructing a contract relation knowledge graph based on the target structured analysis fragment by using the target large model, and generating a target structured data table; verifying each contract term in the target structured data table based on a dynamic rule engine, and generating an early warning report based on a verification result; performing risk assessment on contract terms in the target structured data table based on the target domain knowledge base, the target large model and the contract relation knowledge graph to generate risk data; and generating a revision suggestion based on the early warning report and the risk data by using the target large model, and revising the target contract file based on the revision suggestion. According to the invention, the automation level and accuracy of contract auditing can be improved.
Owner:INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD

RAG-based voucher classification method, medium and equipment

The invention relates to an RAG-based voucher classification method, a medium and equipment, and the method comprises the steps: receiving digital image data of a to-be-classified voucher, carrying out the multi-modal optical character recognition processing to generate a structured OCR result, extracting a text semantic feature vector and a visual layout feature vector based on the structured OCR result, carrying out the fusion of the text semantic feature vector and the visual layout feature vector to generate a multi-modal query vector, and carrying out the classification of the to-be-classified voucher. Similar samples and semantic similarity scores and category metadata thereof are obtained through approximate nearest neighbor retrieval, after an initial candidate category list is generated, key field values are extracted for each candidate category, evidence credibility scores are calculated, comprehensive confidence scores are generated by fusing the semantic similarity scores and the evidence credibility scores, reordering is conducted, and a candidate category list is obtained. And finally, selecting a classification decision path according to the score distribution, and outputting a classification result and an interpretability report. The accuracy and robustness of voucher classification are effectively improved, and complex voucher scenes with changeable formats and fuzzy semantics can be processed; and the interpretability and reliability of the classification decision are enhanced.
Owner:FUJIAN BOSS SOFTWARE

Multi-mode intelligent auditing method for supply chain purchase bid invitation project

The invention provides a supply chain purchase bid invitation project multi-mode intelligent auditing method, and belongs to the technical field of electronic purchase. Comprising the following steps: defining a review rule system comprising rule contents, review points and review logic; in the bidding stage, bidding files and quotation information uploaded by suppliers are received, and after bid opening, bid invitation files, bidding files of entered suppliers and corresponding review rule systems are input into a multi-mode large model server; performing unstructured data analysis on the bidding document and the bidding document by using a multi-modal large model in combination with an OCR (Optical Character Recognition) technology, and identifying and matching response contents in the bidding document through intention; pushing the extracted content to a corresponding business auditing agent, and performing judgment according to an auditing point and auditing logic; and summarizing the review results, and generating an intelligent review report containing the number of conformity items, the number of non-conformity items, conclusions of the review points and judgment bases. The examination efficiency and accuracy are improved, and human errors and compliance risks are reduced.
Owner:INSPUR GENERSOFT CO LTD

Geological domain named entity recognition and classification method based on thinking chain and hybrid experts

The invention discloses a geological domain named entity accurate recognition and classification method based on thinking chain enhancement and hybrid expert architecture, which is characterized by comprising the following steps: firstly, extracting geological document text data through an OCR (Optical Character Recognition) technology, and extracting structured entity data by utilizing a locally deployed large language model; then calling a local large model to generate a diversified sentence pattern template according to language styles in the geological field, filling the template with the extracted professional entities, and constructing an instruction fine tuning data set; further constructing a thinking chain (CoT) enhanced data set on the basis, and explicitly simulating an expert reasoning process; efficient fine tuning is carried out on the large model by innovatively combining a low-rank adaptation (DoRA) technology and a hybrid expert (MoE) architecture, the DoRA technology carries out dimension reduction decomposition and orthogonal transformation on weight matrixes of a decoder layer and a multi-layer perceptron, and the MoE architecture constructs a plurality of special sub-networks to enhance the multi-task processing capability; and finally, performing entity extraction on the geological document by using the fine-tuned model, outputting an identification result containing a reasoning process, and filtering and perfecting the result through a rule matching mechanism. According to the method, the problems of fuzzy boundary, indefinite semantics, difficulty in classification and the like of the named entities in the geological field are effectively solved, the recognition and classification accuracy of the named entities in the geological field is remarkably improved, and key technical support is provided for downstream applications such as geological resource exploration and mineral evaluation.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

Intelligent analysis method and device for unstructured PDF document, equipment and medium

The invention discloses an intelligent analysis method and device for an unstructured PDF document, equipment and a medium, and relates to the field of document analysis, the method comprises the steps of obtaining a to-be-analyzed PDF document, analyzing page elements in the PDF document, and generating a document metadata dictionary; if the PDF document does not contain the extractable text, converting the PDF document into an image and performing optical character recognition to generate first structured data; if the PDF document contains the extractable text, judging whether the PDF document contains a table or not; if the PDF document does not contain the table, a PDFMiner is adopted to extract the text, and second structured data is generated; if the PDF document contains the table, performing multi-modal feature extraction and feature fusion on the PDF document according to the document metadata dictionary to obtain multi-modal fusion features, and generating third structured data according to the multi-modal fusion features; according to the method and the device, the analysis precision and efficiency of the PDF document are improved.
Owner:LU ZE TECH CO LTD

Systems and methods for intelligent real-time KYC identity verification using government issued documents and biometric matching

According to various embodiments, a system and method for verifying a user's identity using both document-based and biometric data is disclosed. The system may prompt a user to upload an image of a government-issued identification document and extract user information from the image using optical character recognition (OCR). The system may also extract the embedded face image from the ID and prompt the user to take a real-time selfie while performing one or more randomized actions or poses. A facial recognition engine may compare the extracted ID image to the live selfie to determine a similarity score. Based on this match, along with optional document authenticity checks, the system may confirm the user's identity in real-time for use cases such as account access, onboarding, and regulatory KYC compliance.
Owner:CELLIGENCE INTERNATIONAL LLC

Graph retrieval enhancement generation method based on telecom specification multi-mode knowledge graph

The invention discloses a graph retrieval enhancement generation method based on a telecom specification multi-modal knowledge graph, and belongs to the technical field of multi-modal knowledge graphs. The method comprises the following steps of: firstly, analyzing titles, texts, pictures and tables in a PDF (Portable Document Format) by virtue of layout analysis, OCR (Optical Character Recognition) and LLM tools through a telecom standard analysis agent; calling a telecom specification multi-modal extraction agent, and extracting chart names and description information from pictures and tables by using a multi-modal large model; the method comprises the following steps: constructing a title hierarchical structure and extracting a triple from a text block, and constructing a multi-modal, multi-hierarchy and multi-granularity telecommunication specification knowledge graph comprising a title, the text block, a picture, a table and the triple; and finally, calling a telecommunication specification multi-granularity retrieval agent, and deeply combining the knowledge graph with the large model, thereby solving the problem that the traditional RAG cannot perform multi-hop reasoning and relieving the illusion problem possibly occurring in the question and answer process of the large model.
Owner:KEDADUOCHUANG CLOUD NETWORK TECH CO LTD

Mobile check deposit

Methods and systems for remote check deposit are disclosed. A check for deposit is processed without the need for a server to receive any image of the check initially. Instead, optical character recognition (OCR) data is received at the server from a mobile device. Verification processing for the check is then performed using the OCR data. If the verification process is successful, a confirmation notification is sent to the mobile device. Subsequently, after sending the confirmation notification, a check image is received, from which the OCR data was determined. The check is, in turn, processed for deposit using the received check image.
Owner:US BANK NATIONAL ASSOCIATION

Device for time-based tracking and cost optimization in construction projects

A device for time-based tracking and cost optimization in construction projects, the device comprising the following: a robust housing suitable for use on construction sites; a processing unit located inside the housing, configured to perform real-time time-stamping, data acquisition and preprocessing tasks; a multimodal sensor unit that is operationally coupled with the processing unit, wherein the sensor unit comprises at least a motion sensor, an RFID reader, sensors for environmental conditions and a vision module with optical character recognition; a real-time clock module that is operationally connected to the processing unit to provide time synchronization for all sensor data streams; a wireless communication module that supports the Wi-Fi, LoRa and LTE protocols and is configured for transmitting time-stamped data to a central project server; a storage module that is operationally coupled with the processing unit to locally buffer time series data of construction activity during offline operation; a housing-mounted, touchscreen-based human-machine interface configured to allow site personnel to enter activity updates and confirm the status of construction tasks; a cost optimization engine running on the central server, the engine being configured to receive time-synchronized sensor data from multiple such devices and dynamically calculate time-cost trade-offs using a predictive planning technique that incorporates the principles of the critical path and the power value; furthermore, the device is configured to be integrated into a digital twin environment of the building under construction in order to provide real-time visualization of progress and to generate suggestions for resource reallocation based on a time-cost-benefit analysis.
Owner:1XL INFRA & REAL ESTATE DEVELOPMENT LLC +2

Document table extraction method and device, equipment and medium

The invention discloses a document table extraction method and device, equipment and a medium, and relates to the technical field of computer information processing. The extraction method comprises the following steps: performing OCR (Optical Character Recognition) on a to-be-processed document table image to obtain a text block; performing visual feature coding on the document table image to obtain deep visual features; performing semantic feature coding on the text sequence of the text block to obtain a semantic feature vector; performing spatial feature coding on the bounding box of the text block to obtain a spatial feature vector; performing feature fusion processing on the deep visual features, the semantic feature vectors and the spatial feature vectors to obtain multi-modal guide features; and performing structured decoding processing on the multi-modal guide features to obtain structured representation of the table. According to the method, the text and the position information pre-recognized by the OCR are fused with the visual features of the document table, so that the visual features are guided to be expressed again and are actively aligned to the logic structure defined by the prior information, and the extraction accuracy of the table logic structure is improved.
Owner:SICHUAN ENRISING INFORMATION TECH CO LTD

Cancer early warning management method and system based on physical examination report

The invention discloses a cancer early warning management method and system based on a physical examination report, and belongs to the field of medicines.The method comprises the steps that health data of individuals are collected, and unstructured texts in the physical examination report are processed through combination of optical character recognition and a natural language processing technology; calculating a cancer risk based on a single physical examination report, evaluating a specific cancer by using a cancer risk scoring system, and analyzing whether an index exceeds a normal range or not through a rule engine; based on historical data trend prediction risks, index change rates are calculated, and an abnormal trend is analyzed and predicted in combination with a time sequence; multi-modal deep learning is utilized to predict individual cancer risks, and a Bayesian network is combined to perform joint analysis on multiple factors to generate a health management scheme; further examination is arranged according to the early warning level, and a screening strategy is provided for specific cancers. According to the method, the OCR + NLP technology is adopted, and the data quality is improved. And in combination with a multi-modal deep learning model, abnormal changes are found in advance, and early warning is realized.
Owner:TONGXIANG MATERNAL & CHILD HEALTH HOSPITAL (TONGXIANG MATERNAL & CHILD HEALTH & FAMILY PLANNING SERVICE CENT)

Nuclear power safety report data extraction method and system based on multi-modal feature fusion

The invention provides a nuclear power safety report data extraction method and system based on multi-modal feature fusion, and belongs to the technical field of data processing, and the method comprises the steps: obtaining a nuclear power safety report document; carrying out OCR (Optical Character Recognition) and paragraph segmentation on the report document to obtain first text data; performing text semantic understanding on the first text data by adopting a dynamic template matching and semantic driving extraction mechanism to obtain a text feature vector; segmenting pages of the report document and identifying fonts and fonts to generate first format data; adopting a cross-page table reconstruction algorithm to detect table cells of the first format data and splicing a cross-page table to obtain format feature vectors; performing feature alignment on the text feature vector and the format feature vector, and inputting the text feature vector and the format feature vector into a large pre-training language model to obtain a semantic vector; and fusing the text feature vector, the format feature vector and the semantic vector to form a composite expression unit, and intelligently extracting structured data based on the composite expression unit. The semantic entity and structural relationship in the nuclear power report can be effectively identified.
Owner:SHANGHAI NUCLEAR ENGINEERING RESEARCH & DESIGN INSTITUTE CO LTD

System and methods for document processing for data extraction and matching

System and methods are disclosed for matching extracted text data based on one or more similarity scores. The method may include receiving one or more documents from a plurality of data sources, utilizing an optical character recognition algorithm for extracting text data from the one or more documents, comparing, utilizing a fuzzy matching algorithm, the extracted text data to reference dataset(s) to determine one or more matches between the extracted text data and at least one of the reference dataset(s), wherein the one or more matches are based on at least one similarity score, inputting the determined one or more matches and the at least one similarity score into a trained machine-learning model to refine the one or more matches, and outputting a representation of the refined one or more matches and the at least one similarity score to a graphical user interface of a device.
Owner:STATE FARM MUTAL AUTOMOBILE INSURANCE COMPANY

Document interpretation and report generation method and device, equipment and medium

The invention relates to the technical field of natural language processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a document interpretation and report generation method, device, equipment and medium, which comprises the following steps: receiving an original document set to generate a structured document object, executing optical character recognition on an image content set to generate a recognition text set, the recognition text set and the text content set are combined into a unified text sequence, element item extraction is executed based on the interpretation template parameter set to generate an interpretation element set, a retrieval enhancement context is retrieved and generated from the domain knowledge base, and the unified text sequence, the interpretation template parameter set and the retrieval enhancement context are input into a language model to generate an interpretation result. And generating report content based on the historical report template set. According to the method, automatic closed loop of document interpretation and report generation is realized through multi-modal unified processing and semantic enhanced reasoning, the efficiency is improved, and the manual dependence and compliance risk are reduced.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Multi-modal information analysis and scheme reminding method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a multi-modal information analysis and scheme reminding method, device, equipment and medium, which comprises the steps of receiving input data and converting the input data into multi-modal data, executing optical character recognition and image classification recognition to generate a recognition result, and sending the recognition result to a server; a natural language processing model is used for analyzing fuzzy description to generate an analysis result, a knowledge base is inquired, a knowledge graph is combined to generate an association result, the analysis result, the association result and user feature data are fused to generate an execution scheme, the execution scheme is compared with an abnormal list, supervision confirmation is triggered, and a compliance instruction is generated. And personalized reminding contents are generated. The information analysis integrity is improved through multi-modal recognition, natural language processing and the knowledge graph, supervision confirmation and personalized reminding are introduced, and intelligent, compliant and reliable reminding management is achieved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Financial robot invoice element identification method based on semantic extraction

The invention discloses a financial robot invoice element identification method based on semantic extraction, and the method comprises the following steps: S1, obtaining an original invoice image, and carrying out the image preprocessing; s2, executing optical character recognition operation, and extracting invoice text information; s3, inputting a semantic potential model; s4, constructing a semantic kernel vector set according to preset invoice element categories; s5, generating a potential tensor field based on the semantic kernel vector set; s6, performing iterative semantic migration operation on the character units in the potential tensor field to form a semantic clustering region; s7, calculating a comprehensive confidence score, and outputting an invoice element recognition result; and S8, performing field legality verification on the invoice element identification result, and submitting the invoice element identification result to a financial robot system after verification is passed to drive related business processes. According to the method, semantic potential modeling and context coding technologies are fused, invoice elements are accurately extracted, and the method has the advantages of being clear in structure, high in robustness and high in adaptability.
Owner:LIANYUNGANG GUOTU INFORMATION TECHNOLOGY CO LTD

OCR (optical character recognition) method and system for high-precision table data structuring

The invention provides an OCR (Optical Character Recognition) method and system for high-precision table data structuring. The method comprises the following steps: converting an original image into a grayscale image and preprocessing the grayscale image; extracting a table edge structure line and filling a fracture part; detecting longitudinal and transverse straight lines in the preprocessed image, calculating intersection points, and determining a table row-column structure; dividing a cell region and positioning to generate a cell coordinate matrix; pixels in the cells are divided into a frame influence area and an effective data area, and the frame influence area executes neighborhood mean filtering and weighted fusion operation; performing end-to-end detection on characters and symbols in the effective data area, and outputting an OCR recognition result with coordinates; and dynamically generating a table structure template based on the cell coordinate matrix and an OCR recognition result, speculating a strategy matching field type through a rule, processing and merging cell missing data based on adjacent cell information, and outputting structured data. The reliability and accuracy of the OCR technology are improved, and the requirement of automatic information processing for high-precision data extraction is met.
Owner:WUHAN UNIV

Mobile check deposit

Methods and systems for remote check deposit are disclosed. A check for deposit is processed without the need for a server to receive any image of the check initially. Instead, optical character recognition (OCR) data is received at the server from a mobile device. Verification processing for the check is then performed using the OCR data. If the verification process is successful, a confirmation notification is sent to the mobile device. Subsequently, after sending the confirmation notification, a check image is received, from which the OCR data was determined. The check is, in turn, processed for deposit using the received check image.
Owner:US BANK NATIONAL ASSOCIATION

Medical guide document analysis and content extraction system and method based on natural language processing and computer vision technology

The invention relates to a medical guide document analysis and content extraction system based on natural language processing and computer vision technology, the system comprises a layout detection model, a formula detection model, a table recognition model, a formula recognition model and an OCR (optical character recognition) model, the layout detection model is used for positioning different elements in a document; the formula detection model is used for positioning a formula in a document; and the table identification model is used for detecting table boundaries, segmenting table cells and analyzing a complex structure. The invention further provides a medical guide analysis and content extraction method applying the medical guide document analysis and content extraction system based on the natural language processing and computer vision technology. According to the invention, efficient analysis and accurate extraction of the medical guide document are realized. The system can accurately identify the text, table, formula and picture information in the document, improves the medical information processing efficiency, enables the acquisition of medical knowledge to be more convenient and efficient, and facilitates the improvement of the medical service quality and research efficiency.
Owner:SHUGUANG HOSPITAL AFFILIATED WITH SHANGHAI UNIV OF T C M

Processing method for medical report structured information extraction and privacy protection

The invention discloses a processing method for structural information extraction and privacy protection of a medical report, which relates to the field of medical report processing and can efficiently and accurately extract text information in a table by integrating advanced image preprocessing, dynamic layout analysis, OCR (optical character recognition) technologies and a dual-channel privacy detection model. Establishing a coordinate mapping relation between the image description paragraph and the corresponding image; the generated JSON structure containing the text, table and image-text mapping relation ensures the integrity and accuracy of the data; the problems of data dislocation, context splitting, privacy disclosure and the like when the medical document with complex typesetting is processed in the prior art are solved; besides, an improved Transform framework is constructed on the basis of a Qwen model, a logical reasoning chain is injected, meanwhile, a case reading large model is obtained by integrating clinical diagnosis rules and multi-modal contrast learning strategy training, the model is utilized to generate a high-quality structured abstract, and the abstract generation accuracy is greatly improved.
Owner:SHANGHAI HENGFANG HEALTH TECHNOLOGY CO LTD

Intelligent ship repair quotation method and system based on multi-modal analysis and large language model

The invention relates to an intelligent ship repair quotation method and system based on multi-modal analysis and a large language model, and the method comprises the steps: firstly receiving a ship repair inquiry sheet file uploaded by a user, and carrying out the feature enhancement of image data through a DocEnTR model, and obtaining a binary image; automatically extracting unstructured text data, semi-structured table data and text information in the binarized image by adopting an OCR (Optical Character Recognition) technology, and performing multi-modal analysis on image features of the binarized image and the extracted text information by adopting a large language model to generate an initial data stream; synchronously generating a standard semantic tag and a non-standard content tag by adopting semantic association of the large language model, automatically generating a dynamic cue word according to the standard semantic tag, inputting the dynamic cue word into the large language model, and outputting a standardized engineering description; performing manual auditing on the non-standard content mark to correct the engineering item field so as to generate structured engineering data; and finally, calculating the total price based on the structured engineering data and the material quantity parameters, and further generating a final quotation list to complete ship repair quotation.
Owner:COSCO SHIPPING HEAVY IND CO LTD +1