Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1123results about "Character recognition" patented technology

Detecting triggering conditions for video game help sessions

The disclosed concepts relate to automatically identifying conditions in a video game to trigger a help session. When a help session is triggered, another video game player or machine learning model can temporarily take over for the current video game player until an ending condition is reached. Help session triggering can be designated by evaluation of prior gameplay data of other video game players to identify in-game conditions that may tend to cause user disengagement, such as in-game conditions that are associated with difficult in-game goals or negative in-game consequences.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Generating unsupervised adversarial examples for machine learning

A trained machine learning model and a training dataset used to train the trained machine learning model can be received. Based on the training dataset, unsupervised adversarial examples can be generated. Robustness of the trained machine learning model can be determined using the generated unsupervised adversarial examples. The training dataset can be augmented with the generated unsupervised adversarial examples. The trained machine learning model can be retrained using the augmented training dataset.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION +1

Real-time document image evaluation

Disclosed herein are system, apparatus, device, method and / or computer program product embodiments for determining, in a remote deposit system, whether a deposit attempt is illegitimate (e.g. fraudulent). Whether the deposit attempt is illegitimate may be assessed based on one or more of the following processes: comparing location data to a location parameter determined from past deposits, comparing an image capture location with a deposit location, and analyzing image-of-image characteristics obtained through image processing to identify whether an image associated with the deposit attempt is an image of an image. In some embodiments, a remote deposit status related to acceptance of the deposit attempt may be provided in real-time
Owner:CAPITAL ONE SERVICES LLC

Generating event commentary in videos using ai models

Disclosed are apparatuses, systems, and techniques for automatically generating commentary to videos that capture sporting activities, computer games, artistic events, political rallies, security-sensitive scenes, and / or any other actions. The techniques include processing a video segment that includes a plurality of video frames, to obtain a description of one or more objects pictured in the video segment and generating, using the obtained description, a prompt for a language model (LM). The techniques further include causing the LM to process the prompt to generate a commentary about an action performed by the one or more objects over a time interval associated with the plurality of video frames.
Owner:NVIDIA CORP

System and method for trade finance operations and sanctions screening process

PendingUS20250378488A1Digital data information retrievalFinanceRisk ControlDocument representation
The present invention discloses a system and method for processing trade finance documents and performing automated compliance screening. The system comprises a computing device, and a database for storing trade finance documents. The system processes documents using OCR to extract text and positional data, generating structured document representations via a layout-aware AI model. An AI classifier module categorizes documents based on content, layout, and domain-specific roles, while a semantic verification module aligns document data with master Letter of Credit templates. A rule management module validates compliance against international trade standards, and a financial crime risk control module performs real-time checks against external sanctions, vessel, and dual-use goods databases. The system further determines and reports discrepancies, anomalies, and compliance issues. The system supports heterogeneous layouts, multi-language documents, and integration with banking APIs.
Owner:CLEARTRADE AI INC

Context aware document augmentation and synthesis

A method includes obtaining a document structure from a data repository. The document structure includes multiple structured sections. A table is detected in a first structured section. A table representation of the table is processed by a general large language model (LLM) to generate a natural language description of the table. An image is detected in the first structured section. The image is processed by an image-processing LLM to generate a natural language description of the image. A form is detected in the first structured section. The form is processed by the general LLM to generate a natural language description of the form. The natural language descriptions of the table, image and form are inserted into the first structured section to obtain a modified first structured section. A modified document structure including the modified first structured section is outputted.
Owner:INTUIT INC

Methods and systems for generating textual outputs from images

Embodiments of the present disclosure provide systems and methods for performing text extraction from an image including textual data. The method performed by a processor includes extracting machine-readable textual data from the image. The machine-readable textual data includes one or more words. The method includes comparing each of the one or more words with a dataset including a domain lexicon database and a language dictionary database to determine a first set of words and a second set of words. The first set of words is words successfully matching with words available in the dataset, and the second set of words is words with no successful match with words available in the dataset. Further, the method includes splitting at least one word of the second set of words into two or more words to determine a third set of words and generating a textual output associated with the image.
Owner:MAERSK AS

Information extraction from unstructured documents with hybrid retrieval augmentation using multi-modal language models

A system for extracting a number of data elements from one or more data sources. Image-based documents are indexed using optical character recognition and a text embedding model to convert the document text to vector embeddings. Relevant portions of the document are identified by comparing the vector embedding of the documents to a vector embedding of a prompt or a request to extract information. The relevant text is mapped to a corresponding page of the documents. The page may be provided to a multi-modal language model for information extraction. The multi-modal language model can process contextual information included in the layout, figures, markings, etc. of the document to extract the information. The system populates an ontological data store based on the response from the language model. Extraction accuracy is improved without significant increases in computations performed by the system.
Owner:AMERICAN INTERNATIONAL GROUP INC

Personalized document field prediction based on learning from user feedback

Particular embodiments relate to personalized document field prediction based on user behavior and feature generation. Specifically, various embodiments have the technical effect of improved accuracy with respect to field / entity value prediction (e.g., predicting that the amount due is X via a Gradient Boosting Model) relative to document processing technologies by learning through user behavior data or feedback (e.g., through continuous reinforcement learning from human feedback (RLHF)). This is at least partially because of the technical solution of accessing or generating unique features from one or more documents previously used by a user.
Owner:BILL OPERATIONS LLC

Character recognition model training method and apparatus, character recognition method and apparatus, device and storage medium

The present disclosure provides a character recognition model training method and apparatus, a character recognition method and apparatus, a device and a medium, relating to the technical field of artificial intelligence, and specifically to the technical fields of deep learning, image processing and computer vision, which can be applied to scenarios such as character detection and recognition technology. The specific implementing solution is: partitioning an untagged training sample into at least two sub-sample images; dividing the at least two sub-sample images into a first training set and a second training set; where the first training set includes a first sub-sample image with a visible attribute, and the second training set includes a second sub-sample image with an invisible attribute; performing self-supervised training on a to-be-trained encoder by taking the second training set as a tag of the first training set, to obtain a target encoder.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Extracting images and determining their meaning for semantic image retrieval and training a transformer-based multi-modal large language model to generate domain-aware images based on image meanings

The disclosure relates to systems and methods automatically extracting an image and related image components, computationally determining an understanding of the image, and generating mathematical vector embeddings via sentence encoders based on the computationally determined understanding. The mathematical vector embeddings may be used for semantic image retrieval that enables image searching based on a semantic understanding of input images and / or input text. The mathematical vector embeddings may be used for training and executing generative Artificial Intelligence (AI) models to create new content that includes retrieved images and / or generate new images.
Owner:ROHIRRIM INC

Business data automatic reconciliation method and device and computer equipment

The invention relates to a business data automatic reconciliation method and device and computer equipment, and the method comprises the steps: automatically obtaining multi-source data, positioning unstructured data, carrying out the data enhancement, generating data with rich features, and supporting the subsequent accurate automatic reconciliation; the data processing efficiency is improved through cluster cleaning and large model auxiliary mapping; according to the technical scheme, the transaction tail difference and the time delay are automatically processed based on the preset rule and the fuzzy matching, the subtle difference between the transaction tail difference and the time delay can be accurately recognized and processed, the accuracy of the account checking result is ensured, and therefore efficient and accurate business data automatic account checking can be achieved.
Owner:IND CONSUMER FINANCE CO LTD

Controllable visual text generation with adapter-enhanced diffusion models

A method, apparatus, non-transitory computer readable medium, and system for image generation include obtaining a text content image and a text style image. The text content image is encoded to obtain content guidance information and the text style image is encoded to obtain style guidance information. Then a synthesized image is generated based on the content guidance information and the style guidance information. The synthesized image includes text from the text content image having a text style from the text style image.
Owner:ADOBE INC

Method and system of converting unstructured digital documents to a structure format using a secure API

In one aspect, a computerized method for document extraction workflow for unstructured documents includes the steps of implementing a text mining operation on a set of digital documents the incoming documents. This is done by defining a document type of each digital document. Based on the document type, the method defines a set of data dictionaries to extract any data from each digital document. The method uses the defined set of data dictionaries to extract any data from each digital document.
Owner:YERRAMSETTY VENKATA SAI RAMAN +1

A method of visual recognition and tracking of objects using virtual sensor

This invention relates to a method for visually recognising and tracking objects and events in 3D space with high accuracy, reduced computational load and fast setup, using virtual sensors that are designed to ensure high reliability and reduce costs by maximising their potential for accurate object recognition for given visual input data. The main task of a virtual sensor is to provide cost-effective and reliable data and replace the use of humans and / or hardware sensors to detect and recognise objects or events that are important to a given entity, such as an industrial enterprise and various business entities. The solution is well suited for industrial applications, especially for tracking larger objects of known shape or appearance, such as in large industrial halls, warehouses without fixed racking systems, container docks, train docks, car parks, etc. It is particularly suitable for tracking coils, metal pieces and larger building components. The method provides a reliable source of i data for obtaining a digital twin of the monitored objects, which represents real-time information about each stock keeping unit (SKU) in the warehouse (WH), bringing significant benefits to WH managers and thus saving human labour spent on searching for materials, improving management flow, reducing overall equipment effectiveness (OEE) of vehicles, improving throughput, quality, etc.
Owner:INOVEC TECHNOLOGY SRO

Neural network based determination of evidence relevant for answering natural language questions

A system makes evidence-based decision for actions associated with a user. The system receives information describing an action associated with the user. The system receives documents associated with the user and questions associated with the action. For each of the plurality of questions the system performs the following steps. The system evaluates sentences from the document in relation to the question using an evidence extraction model. The system classifies the evidence using an evidence classification model to determine whether the evidence refutes or supports a decision based on the question. The system makes a decision regarding the action based on the classifications of evidence sentences in relation to each of the plurality of questions. The evidence extraction model and the evidence classification model are trained neural networks.
Owner:HUMANA INC

Aircraft in-flight entertainment devices installation verification by machine vision-assisted identification

Systems, devices, and methods for aircraft in-flight entertainment (IFE) devices installation verification are disclosed herein. In some implementations, a verification system includes a computing device that can operably connect to multiple IFE devices and a reader device. Each IFE device can report its own stored identity to the computing device, and the computing device can determine the position of each IFE device based on a network topology-based network address assigned to each IFE device. The reader device can obtain and communicate, to the computing device, actual identities and expected installation positions of the IFE devices. The verification system can then compare, for each position, the actual identity, which indicates which IFE device should be installed in that position, against the reported stored identity, which indicates which IFE device has been installed in that position. The verification system can display, on a graphical user interface, the comparison.
Owner:PANASONIC AVIONICS CORP

Methods and system for image-based analysis for intelligent item identification and utilization

Techniques described herein are directed to image-based analysis for intelligent item identification and utilization. Image data corresponding to an image of a list of items may be analyzed to determine what the items on the list are, and thereafter interaction data and machine-trained models may be utilized to determine merchant offerings for the items to display to a user. Various payment options may be presented to the user and utilized in association with a payment service.
Owner:AFTERPAY PTY LTD

Information auditing system and method based on multi-modal content identification

The invention discloses an information auditing system and method based on multi-modal content recognition, multi-modal content comprises text content, image content, video content and short link content, and the system comprises a rapid auditing module, a full-modal auditing module and a self-adaptive feedback module which are mutually coupled, the rapid auditing module rapidly audits the to-be-audited information; and the full-modal auditing module performs deep analysis on the to-be-audited information, and the adaptive feedback module is used for training the lightweight visual language large model and / or the visual language large model according to the to-be-audited information and a final auditing result corresponding to the to-be-audited information. The system adopts a three-layer auditing architecture, the first layer is multi-modal rapid filtering (rapid auditing module), the second layer is multi-modal deep analysis (full-modal auditing module), the third layer is a self-adaptive feedback mechanism (self-adaptive feedback module), and a'rapid-accurate-stable 'closed-loop auditing process is integrally formed.
Owner:BEIJING HONGLIAN 95 INFORMATION IND CO LTD

Augmented reality typography personalization system

Disclosed are augmented reality (AR) personalization systems to enable a user to edit and personalize presentations of real-world typography in real-time. The AR personalization system captures an image depicting a physical location via a camera coupled to a client device. For example, the client device may include a mobile device that includes a camera configured to record and display images (e.g., photos, videos) in real-time. The AR personalization system causes display of the image at the client device, and scans the image to detect occurrences of typography within the image (e.g., signs, billboards, posters, graffiti).
Owner:SNAP INC

Guided content capture

A computing system can process incident information corresponding to a vehicle incident involving a vehicle of a user. The system can initiate a guided content capture process with a user to capture images of a vehicle of the user. The system generates a sequential set of vehicle outlines of the vehicle on an image interface presented on a computing device of the user to implement the guided content capture process. The system performs computer vision on image data captured by a camera of the computing device to determine when the vehicle is aligned with a particular vehicle outline of the sequential set of vehicle outlines.
Owner:ASSURED INSURANCE TECH INC

Two-layered image compression for text content

Coding an image that includes text content and a background is disclosed. Text portions are identified in the image. The text portions are extracted from the image to obtain a background image, where the background image includes holes corresponding to respective areas of the text portions within the image. A filled-in background image is obtained based on the background image. The filled-in background image is encoded into a compressed bitstream using a block-based encoder. The text portions is also encoded into the compressed bitstream. Encoding the text portions includes encoding respective high quality text binarization upscaled binary maps.
Owner:GOOGLE LLC

Image processing apparatus and image processing method for identifying a classification type of a document image

The present application is to obtain a character recognition result by performing character recognition processing on a document image and identify a classification type of the document image based on a character string included in the character recognition result and a predefined condition. The condition for identifying classification types that are hints for expense items is defined in advance.
Owner:CANON KK

Automated Evaluation and Feedback of Participant Online Testing

A system and method automatically evaluate participant online testing and automatically provide feedback regarding test results of a participant. Automatic evaluation includes processing video frames, audio, timestamped screenshots, and initial test results of the participant answering multiple test questions and transcribing audio. The system and method utilize optical character recognition on the screenshots to generate a set of text for each screenshot and compare the set of text from each timestamped screenshot to determine which screenshots are associated with each test question. The system and method further determine time periods for each question, segment the transcribed audio, the screenshots, and the video frames by question and select a screenshot and a video frame for each test question from the segmented screenshots and video frames. The system and method utilize a first AI tool and a second AI tool to determine whether the participant demonstrated mastery for each test question.
Owner:2HR LEARNING INC

Hospital adverse event risk identification method and system based on causal reasoning

The invention provides a hospital adverse event risk identification method and system based on causal reasoning, and the method comprises the steps: responding to an operation request of a user for reporting a hospital adverse event based on the optimization of medical advice data processing and a medical quality management process in an operation risk identification scene, inputting the detailed information of the event through a formatted form, and carrying out the operation of the event. Through multi-node, multi-auditor and multi-department cooperative auditing, an adverse event form reported by a user is subjected to approval, flow and summary analysis, medical advice data is subjected to structured storage by adopting visual enhancement OCR, dynamic knowledge graph verification and a time sequence causal reasoning model, a time window division algorithm is adopted, and medical advice data is obtained. The causal discovery algorithm constructs an unplanned re-operation association network, early warning of adverse events of operation complications is carried out, and compared with a traditional mode, the medical advice processing efficiency, the transcription error rate and the unplanned operation detection rate are all greatly improved.
Owner:THE AFFILIATED SIR RUN RUN SHAW HOSPITAL OF SCHOOL OF MEDICINE ZHEJIANG UNIV

Detecting and remediating anomalies in institutional financial instruments using image processing

Methods and systems for processing a financial instrument, such as a check, are provided. An example method includes performing a text extraction process on a check image to obtain check textual information, and decoding an intelligent mail barcode (IMB) in the check image to obtain a character code. The method further includes utilizing the character code to extract encoded check information including a payee ZIP code from the IMB. The method also includes performing a fraud detection process on the check image based on at least one of (1) scores generated during the text extraction process, or (2) inconsistency between the check textual information and the encoded check information.
Owner:US BANK NATIONAL ASSOCIATION

Mobile phone interface automatic testing method based on ADB and YOLO

The invention relates to the technical field of mobile phone testing, and provides a mobile phone interface automatic testing method based on ADB and YOLO, which comprises the following steps: acquiring a mobile phone screenshot by using an ADB tool; carrying out image preprocessing on the obtained mobile phone screenshot; a pre-trained YOLO model is loaded, the preprocessed mobile phone screenshot is input into the YOLO model for target detection, and a detection result is obtained; executing OCR character recognition in a detection result of the YOLO model, and extracting prompt text content on a mobile phone interface; the text content recognized by the OCR is matched with predefined keywords, rules or regular expressions, whether prompt information on a mobile phone interface meets the expectation or not is judged, and the current operation state of the mobile phone is judged by recognizing the state of a specific icon or button on the mobile phone interface; and recording prompt information of a mobile phone interface into a log according to an identification result, and when further operation is needed, simulating key operation by using an ADB command to complete a corresponding automatic task.
Owner:BEIJING HONGSHAN INFORMATION TECH RES CO LTD