Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1459results about "Character recognition" patented technology

Methods for Automatically Generating a Training Dataset for Training an Optical Recognition Model for Reading Street Signs

Various embodiments include methods for generating image datasets for training an artificial intelligence machine learning (AI / ML) optical character recognition (OCR) model. Image processing may be performed on a plurality of roadway images to identify street signs within the images and generate a dataset of sign images categorized into sign variants of the same shape, color, pictogram, and characters. An OCR model may process sign images to obtain OCR results for images of each sign variant. An aggregation process may be performed on the OCR results for all sign images within each sign variant to identify a ground truth OCR result for each sign variant. The ground truth OCR result may be used to automatically label all sign images of each sign variant to produce an OCR model training dataset. The produced training dataset may then be used to retrain the initial AI / ML OCR model and / or train other AI / ML OCR models.
Owner:QUALCOMM INC

Power grid wiring diagram automatic analysis system and method based on deep learning

The invention provides a power grid wiring diagram automatic analysis system and method based on deep learning. The method comprises the steps of S1, data preprocessing and model training; s2, carrying out equipment identification and association; step S3, topology reasoning and verification are carried out; s4, standardized output: converting the verified power grid network model into a regulation and control cloud platform standard graphic file, and optimizing graphic arrangement through an adaptive layout algorithm; step S5, performing closed-loop optimization; and S6, performing manual intervention and iteration. According to the method, the wiring diagram analysis efficiency and accuracy can be remarkably improved, manual intervention is reduced, standardized graph generation of a power grid regulation and control cloud platform is supported, and a reliable data basis is provided for power system analysis and decision making.
Owner:HUBEI CENT CHINA TECH DEV OF ELECTRIC POWER

Systems and methods for generating and using semantic images in deep learning for classification and data extraction

Disclosed is a new document processing solution that combines the powers of machine learning and deep learning and leverages the knowledge of a knowledge base. Textual information in an input image of a document can be converted to semantic information utilizing the knowledge base. A semantic image can then be generated utilizing the semantic information and geometries of the textual information. The semantic information can be coded by semantic type determined utilizing the knowledge base and positioned in the semantic image utilizing the geometries of the textual information. A region-based convolutional neural network (R-CNN) can be trained to extract regions from the semantic image utilizing the coded semantic information and the geometries. The regions can be mapped to the textual information for classification / data extraction. With semantic images, the number of samples and time needed to train the R-CNN for document processing can be significantly reduced.
Owner:CROWDSTRIKE

System And Method For Using Artificial Intelligence (AI) To Analyze Social Media Content

Systems and methods for reducing the search space by processing media content to refine search parameters. A computing device may obtain the media content in response to receiving a request for inclusion of the media content in a media content knowledge repository, extract an audio component, a video component, and a text component of the media content, and determine attributes within the extracted components. The computing device may determine segment attributes based on a result of correlating the determined audio, video, and text attributes, integrate the segment attributes into the media content knowledge repository, and / or perform any of a variety of responsive actions.
Owner:SOCIAL VOICE LTD

Document remembrance and counterfeit detection

Disclosed herein are system, apparatus, device, method and / or computer program product embodiments for identifying, in a remote deposit system, a level of correspondence of a financial instrument with past deposit activity. The level of correspondence may be used to determine a funds availability schedule and / or identify counterfeit financial instruments. In some embodiments, the level of correspondence may be determined by comparing a date associated with the financial instrument to a date parameter, an amount of the financial instrument to an amount parameter, and / or one or more financial instrument characteristics with one or more payer-specific financial instrument characteristics.
Owner:CAPITAL ONE SERVICES LLC

Detecting triggering conditions for video game help sessions

The disclosed concepts relate to automatically identifying conditions in a video game to trigger a help session. When a help session is triggered, another video game player or machine learning model can temporarily take over for the current video game player until an ending condition is reached. Help session triggering can be designated by evaluation of prior gameplay data of other video game players to identify in-game conditions that may tend to cause user disengagement, such as in-game conditions that are associated with difficult in-game goals or negative in-game consequences.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Invoice automatic auditing and intelligent settlement system for automobile road rescue

The invention discloses an invoice automatic auditing and intelligent settlement system for automobile road rescue, and relates to the technical field of intelligent settlement systems, a correction module integrates a machine learning model to analyze historical bills, automatically corrects unreasonable settlement amount and generates a correction suggestion report, a prompt module builds an intelligent priority matrix, and the intelligent priority matrix is sent to a server. The reminding frequency is dynamically adjusted according to parameters such as the historical settlement scale and the service life of a supplier, the settlement module recognizes an abnormal invoicing terminal by combining an invoicing equipment fingerprint recognition technology, a payment process is automatically triggered after verification is passed, a supplier credit portrait system is established, and a dynamic settlement scheme is supported. The settlement system establishes an intelligent priority matrix to realize important supplier high-frequency reminding and low-risk project low-frequency intervention, so that the whole process timeliness is guaranteed, the defects of a traditional rule engine in coping with complex scenes are effectively overcome, and the response capability and the autonomous decision-making capability of the system to abnormal behaviors are improved.
Owner:HANGZHOU WANGLAN TECH CO LTD

Traffic object recognition systems and methods

This disclosure provides systems, methods, and devices for vehicle driving assistance systems that support image processing. In a first aspect, the methods addresses traffic sign recognition as a language model-based image reasoning task by utilizing a multi-modal transformer architecture that combines the strength of vision and language machine learning (ML) models. The transformer architecture recognizes traffic signs based on visual features and associated taxonomy of the traffic signs. In a second aspect, the methods leverage context surrounding an autonomous vehicle through a graph-based modeling framework that fuses outputs from multiple perception modules to construct a semantic scene graph representation of an intersection, which consolidates processing diverse data types for traffic light relevancy detection. Other aspects and features are also claimed and described.
Owner:QUALCOMM INC

Generating unsupervised adversarial examples for machine learning

A trained machine learning model and a training dataset used to train the trained machine learning model can be received. Based on the training dataset, unsupervised adversarial examples can be generated. Robustness of the trained machine learning model can be determined using the generated unsupervised adversarial examples. The training dataset can be augmented with the generated unsupervised adversarial examples. The trained machine learning model can be retrained using the augmented training dataset.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION +1

Rejection of impermissible documents

A computer implemented method, system and non-transitory computer-readable device for a remote deposit environment activating, on a client device, a financial application, wherein the financial application is configured to instantiate a customer interface (UI) on the client device. Upon receiving a customer request, based on interactions with the UI, the method implements an electronic deposit of a financial instrument by generating a live stream of image data of a field of view of at least one camera, wherein the live stream of image data includes imagery of at least a portion of the financial instrument, determining, based on the live stream of image data and a machine learning model (ML), an impermissibility score of the financial instrument and, based on exceeding a selectable threshold of the impermissibility score, modifying a status of the remote deposit to pause or terminate the remote deposit.
Owner:CAPITAL ONE SERVICES LLC

Systems and methods for trigger-based updates to camograms for autonomous checkout in a cashier-less shopping

Systems and methods for tracking inventory items in an area of real space are disclosed. The method includes receiving a signal generated in dependence on sensors. The signal indicates a change to a portion of an image of an area of real space. The method includes, in response to receiving the signal, implementing a trained location detection model to determine, based on inputs, whether an inventory item identified in the portion of the image has changed a position in the area of real space. The method includes implementing a trained item classification model to determine a classification of the inventory item. The method includes updating an inventory database with inventory item data determined in dependence on the classification of the inventory item to provide an updated map of the area of real space as a result of the received signal indicating the change to the portion of the image.
Owner:STANDARD COGNITION CORP

Real-time document image evaluation

Disclosed herein are system, apparatus, device, method and / or computer program product embodiments for determining, in a remote deposit system, whether a deposit attempt is illegitimate (e.g. fraudulent). Whether the deposit attempt is illegitimate may be assessed based on one or more of the following processes: comparing location data to a location parameter determined from past deposits, comparing an image capture location with a deposit location, and analyzing image-of-image characteristics obtained through image processing to identify whether an image associated with the deposit attempt is an image of an image. In some embodiments, a remote deposit status related to acceptance of the deposit attempt may be provided in real-time
Owner:CAPITAL ONE SERVICES LLC

Customer agent recording systems and methods

Automatic analyses of customer-agent interactions provide valuable, actionable feedback for managers, agents, and customers. To provide such analyses, methods for recording and analysis of customer-agent interactions using a customer relationship management (CRM) system are disclosed. A recorder application records the customer-agent interaction, and sensitive information may be identified. Sensitive portions of the recording may then be redacted and removed from the recording. The redacted recording is then analyzed to generate useful summary and analytics information.
Owner:LEVEL AI

Artificial intelligence cyber security analyst

An analyzer module forms a hypothesis on what are a possible set of cyber threats that could include the identified abnormal behavior and / or suspicious activity with AI models trained with machine learning on possible cyber threats. The Analyzer analyzes a collection of system data, including metric data, to support or refute each of the possible cyber threat hypotheses that could include the identified abnormal behavior and / or suspicious activity data with the AI models. A formatting and ranking module outputs supported possible cyber threat hypotheses into a formalized report that is presented in 1) printable report, 2) presented digitally on a user interface, or 3) both.
Owner:DARKTRACE HLDG LTD

Generating event commentary in videos using ai models

Disclosed are apparatuses, systems, and techniques for automatically generating commentary to videos that capture sporting activities, computer games, artistic events, political rallies, security-sensitive scenes, and / or any other actions. The techniques include processing a video segment that includes a plurality of video frames, to obtain a description of one or more objects pictured in the video segment and generating, using the obtained description, a prompt for a language model (LM). The techniques further include causing the LM to process the prompt to generate a commentary about an action performed by the one or more objects over a time interval associated with the plurality of video frames.
Owner:NVIDIA CORP

Language-agnostic OCR extraction

Technologies for language agnostic OCR extraction include identifying a word region of an image using optical character recognition, applying a language agnostic machine learning model to the word region, where the language agnostic machine learning model is trained on training data including a set of image-text pairs and a set of multilingual text translation pairs, receiving, from the language agnostic machine learning model, a word region embedding that is associated with the word region, searching a multilingual index for a text embedding that matches the word region embedding, receiving, from the multilingual index, text associated with the text embedding; and outputting at least one of the text or the text embedding to at least one downstream process, application, system, component, or network.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

System and method for trade finance operations and sanctions screening process

PendingUS20250378488A1Digital data information retrievalFinanceRisk ControlDocument representation
The present invention discloses a system and method for processing trade finance documents and performing automated compliance screening. The system comprises a computing device, and a database for storing trade finance documents. The system processes documents using OCR to extract text and positional data, generating structured document representations via a layout-aware AI model. An AI classifier module categorizes documents based on content, layout, and domain-specific roles, while a semantic verification module aligns document data with master Letter of Credit templates. A rule management module validates compliance against international trade standards, and a financial crime risk control module performs real-time checks against external sanctions, vessel, and dual-use goods databases. The system further determines and reports discrepancies, anomalies, and compliance issues. The system supports heterogeneous layouts, multi-language documents, and integration with banking APIs.
Owner:CLEARTRADE AI INC

Real-time image validity assessment

A computer implemented method, system, and non-transitory computer-readable device for a remote deposit environment. In some embodiments, a predictive machine learning (ML) model may be trained to determine a likelihood an image will be successfully processed via OCR. The predictive ML model may determine the likelihood prior to the image being uploaded to a remote server and / or processed via OCR, allowing the image to be rejected and replaced in real time. In some embodiments, the predictive ML model may be implemented on a mobile device. Optionally, the predictive ML model may be supported by a deep learning model operating to refine the predictive ML model.
Owner:CAPITAL ONE SERVICES LLC

Work order generation method, work order processing method and work order management system

The invention discloses a work order generation method, a work order processing method and a work order management system. The work order generation method comprises the following steps: performing text extraction on multi-modal problem data reported by a user to obtain corresponding target text data; analyzing the multi-modal problem data by using the problem classification model to obtain a corresponding problem type, inputting the target text data into a work order template matched with the problem type to obtain a work order to be processed, and determining a generation timestamp of the work order to be processed; analyzing the multi-modal problem data by using an emotional tendency recognition model to obtain a corresponding emotional tendency; and packaging the emotional tendency, the question type, the work order to be processed and the corresponding generation timestamp corresponding to the multi-modal question data into a work order information packet, sending the work order information packet to a server, and receiving a question result fed back by the server. The technical problem that the work order generation accuracy is poor due to the fact that the multi-modal feedback data of the user cannot be fully understood in the related technology is solved.
Owner:CHINA TELECOM CORP LTD

Court area identification method and device, and storage medium

The invention discloses a court area identification method and device and a storage medium, and relates to the technical field of data processing. According to the method, the depth map and the RGB map of the court acquired by the image acquisition device are acquired, the point cloud data are generated based on the depth map, the point cloud fluctuation amplitude corresponding to the court is determined according to the point cloud data, then the color distribution characteristics and the texture distribution characteristics of the court are determined based on the RGB map, and finally, the color distribution characteristics and the texture distribution characteristics of the court are determined. The area type of the court is determined based on the color distribution characteristics and the texture distribution characteristics of the court and the point cloud fluctuation amplitude corresponding to the court, so that court area identification is more comprehensive and accurate, and the risk that a golf cart enters a forbidden area by mistake is reduced.
Owner:RUICHI LASER (SHENZHEN) CO LTD

Context aware document augmentation and synthesis

A method includes obtaining a document structure from a data repository. The document structure includes multiple structured sections. A table is detected in a first structured section. A table representation of the table is processed by a general large language model (LLM) to generate a natural language description of the table. An image is detected in the first structured section. The image is processed by an image-processing LLM to generate a natural language description of the image. A form is detected in the first structured section. The form is processed by the general LLM to generate a natural language description of the form. The natural language descriptions of the table, image and form are inserted into the first structured section to obtain a modified first structured section. A modified document structure including the modified first structured section is outputted.
Owner:INTUIT INC

Methods and systems for generating textual outputs from images

Embodiments of the present disclosure provide systems and methods for performing text extraction from an image including textual data. The method performed by a processor includes extracting machine-readable textual data from the image. The machine-readable textual data includes one or more words. The method includes comparing each of the one or more words with a dataset including a domain lexicon database and a language dictionary database to determine a first set of words and a second set of words. The first set of words is words successfully matching with words available in the dataset, and the second set of words is words with no successful match with words available in the dataset. Further, the method includes splitting at least one word of the second set of words into two or more words to determine a third set of words and generating a textual output associated with the image.
Owner:MAERSK AS

Information extraction from unstructured documents with hybrid retrieval augmentation using multi-modal language models

A system for extracting a number of data elements from one or more data sources. Image-based documents are indexed using optical character recognition and a text embedding model to convert the document text to vector embeddings. Relevant portions of the document are identified by comparing the vector embedding of the documents to a vector embedding of a prompt or a request to extract information. The relevant text is mapped to a corresponding page of the documents. The page may be provided to a multi-modal language model for information extraction. The multi-modal language model can process contextual information included in the layout, figures, markings, etc. of the document to extract the information. The system populates an ontological data store based on the response from the language model. Extraction accuracy is improved without significant increases in computations performed by the system.
Owner:AMERICAN INTERNATIONAL GROUP INC

Personalized document field prediction based on learning from user feedback

Particular embodiments relate to personalized document field prediction based on user behavior and feature generation. Specifically, various embodiments have the technical effect of improved accuracy with respect to field / entity value prediction (e.g., predicting that the amount due is X via a Gradient Boosting Model) relative to document processing technologies by learning through user behavior data or feedback (e.g., through continuous reinforcement learning from human feedback (RLHF)). This is at least partially because of the technical solution of accessing or generating unique features from one or more documents previously used by a user.
Owner:BILL OPERATIONS LLC

Systems and methods for automated vehicle recognition at service facilities

Technical solutions are directed to a system including one or more processors, coupled with memory. The one or more processors can maintain a plurality of vehicle profiles for a plurality of vehicles. The one or more processors can capture one or more images of at least a portion of a vehicle, identify, using the one or more images, a first sequence of characters of a license plate of the vehicle and determine one or more characters of the first sequence of characters that satisfy a character replacement schema. The one or more processors can generate a second sequence of characters including one or more corresponding replacement characters in the first sequence of characters, identify, a vehicle profile of the vehicle based on executing a query using the generated second sequence of characters, and transmit, to a device, data included in the vehicle profile based on identifying the vehicle profile.
Owner:MW POS INC

Semantic matching between a source screen or source data and a target screen using semantic artificial intelligence

Semantic matching between a source screen or source data and a target screen using semantic artificial intelligence (AI) for robotic process automation (RPA) workflows is disclosed. The source data or source screen and the target screen are selected on a matching interface, semantic matching is performed between the source data / screen and the target screen using an artificial intelligence / machine learning (AI / ML) model, and matching graphical elements and unmatched graphical elements are highlighted, allowing the developer to see which graphical elements match and which do not. The matching interface may also provide a confidence score of the individual matches, provide an overall mapping score, and allow the developer to hide / unhide the matched / unmatched graphical elements. Activities of an RPA workflow may be automatically created based on the semantic mapping that can be executed to perform the automation.
Owner:UIPATH INC

Character recognition model training method and apparatus, character recognition method and apparatus, device and storage medium

The present disclosure provides a character recognition model training method and apparatus, a character recognition method and apparatus, a device and a medium, relating to the technical field of artificial intelligence, and specifically to the technical fields of deep learning, image processing and computer vision, which can be applied to scenarios such as character detection and recognition technology. The specific implementing solution is: partitioning an untagged training sample into at least two sub-sample images; dividing the at least two sub-sample images into a first training set and a second training set; where the first training set includes a first sub-sample image with a visible attribute, and the second training set includes a second sub-sample image with an invisible attribute; performing self-supervised training on a to-be-trained encoder by taking the second training set as a tag of the first training set, to obtain a target encoder.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Extracting images and determining their meaning for semantic image retrieval and training a transformer-based multi-modal large language model to generate domain-aware images based on image meanings

The disclosure relates to systems and methods automatically extracting an image and related image components, computationally determining an understanding of the image, and generating mathematical vector embeddings via sentence encoders based on the computationally determined understanding. The mathematical vector embeddings may be used for semantic image retrieval that enables image searching based on a semantic understanding of input images and / or input text. The mathematical vector embeddings may be used for training and executing generative Artificial Intelligence (AI) models to create new content that includes retrieved images and / or generate new images.
Owner:ROHIRRIM INC