Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

627results about "Character recognition" patented technology

Detecting triggering conditions for video game help sessions

The disclosed concepts relate to automatically identifying conditions in a video game to trigger a help session. When a help session is triggered, another video game player or machine learning model can temporarily take over for the current video game player until an ending condition is reached. Help session triggering can be designated by evaluation of prior gameplay data of other video game players to identify in-game conditions that may tend to cause user disengagement, such as in-game conditions that are associated with difficult in-game goals or negative in-game consequences.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Information extraction from unstructured documents with hybrid retrieval augmentation using multi-modal language models

A system for extracting a number of data elements from one or more data sources. Image-based documents are indexed using optical character recognition and a text embedding model to convert the document text to vector embeddings. Relevant portions of the document are identified by comparing the vector embedding of the documents to a vector embedding of a prompt or a request to extract information. The relevant text is mapped to a corresponding page of the documents. The page may be provided to a multi-modal language model for information extraction. The multi-modal language model can process contextual information included in the layout, figures, markings, etc. of the document to extract the information. The system populates an ontological data store based on the response from the language model. Extraction accuracy is improved without significant increases in computations performed by the system.
Owner:AMERICAN INTERNATIONAL GROUP INC

Method and system of converting unstructured digital documents to a structure format using a secure API

In one aspect, a computerized method for document extraction workflow for unstructured documents includes the steps of implementing a text mining operation on a set of digital documents the incoming documents. This is done by defining a document type of each digital document. Based on the document type, the method defines a set of data dictionaries to extract any data from each digital document. The method uses the defined set of data dictionaries to extract any data from each digital document.
Owner:YERRAMSETTY VENKATA SAI RAMAN +1

Methods and system for image-based analysis for intelligent item identification and utilization

Techniques described herein are directed to image-based analysis for intelligent item identification and utilization. Image data corresponding to an image of a list of items may be analyzed to determine what the items on the list are, and thereafter interaction data and machine-trained models may be utilized to determine merchant offerings for the items to display to a user. Various payment options may be presented to the user and utilized in association with a payment service.
Owner:AFTERPAY PTY LTD

Guided content capture

A computing system can process incident information corresponding to a vehicle incident involving a vehicle of a user. The system can initiate a guided content capture process with a user to capture images of a vehicle of the user. The system generates a sequential set of vehicle outlines of the vehicle on an image interface presented on a computing device of the user to implement the guided content capture process. The system performs computer vision on image data captured by a camera of the computing device to determine when the vehicle is aligned with a particular vehicle outline of the sequential set of vehicle outlines.
Owner:ASSURED INSURANCE TECH INC

Two-layered image compression for text content

Coding an image that includes text content and a background is disclosed. Text portions are identified in the image. The text portions are extracted from the image to obtain a background image, where the background image includes holes corresponding to respective areas of the text portions within the image. A filled-in background image is obtained based on the background image. The filled-in background image is encoded into a compressed bitstream using a block-based encoder. The text portions is also encoded into the compressed bitstream. Encoding the text portions includes encoding respective high quality text binarization upscaled binary maps.
Owner:GOOGLE LLC

Pipeline Architecture for Road Sign Detection and Evaluation

The technology provides a sign detection and classification methodology. A unified pipeline approach incorporates generic sign detection with a robust parallel classification strategy. Sensor information such as camera imagery and lidar depth, intensity and height (elevation) information are applied to a sign detector module. This enables the system to detect the presence of a sign in a vehicle's external environment. A modular classification approach is applied to the detected sign. This includes selective application of one or more trained machine learning classifiers, as well as a text and symbol detector. Annotations help to tie the classification information together and to address any conflicts with different outputs from different classifiers. Identification of where the sign is in the vehicle's surrounding environment can provide contextual details. Identified signage can be associated with other objects in the vehicle's driving environment, which can be used to aid the vehicle in autonomous driving.
Owner:WAYMO LLC

Cross-modality neural network transform for semi-automatic medical image annotation

A cross-modality neural network transform for semi-automatic medical image annotation is provided. In various embodiments, an input medical image is mapped to a first vector in a text vector space. The first vector corresponds to the features of the medical image. A set of predetermined vectors is searched for a closest one of the predetermined vectors to the first vector. From the closest one of the predetermined vectors, one or more keywords is determined describing the input medical image.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Automatically monitoring retail products based on captured images

A system for acquiring images of products in a retail store is disclosed. The system may include at least one first housing configured for location on a retail shelving unit, and at least one image capture device included in the at least one first housing and configured relative to the at least one first housing such that an optical axis of the at least one image capture device is directed toward an opposing retail shelving unit when the at least one first housing is fixedly mounted on the retail shelving unit. The system may further include a second housing configured for location on the retail shelving unit separate from the at least one first housing, the second housing may contain at least one processor configured to control the at least one image capture device and also to control a network interface for communicating with a remote server. The system may also include at least one data conduit extending between the at least one first housing and the second housing, the at least one data conduit being configured to enable transfer of control signals from the at least one processor to the at least one image capture device and to enable collection of image data acquired by the at least one image capture device for transmission by the network interface.
Owner:TRAX TECH SOLUTIONS

Automated captioning of augmented reality effects in videos

Described herein are techniques for generating captions for augmented reality (AR) effects in videos using an adapted multimodal large language model (MLLM). The technique involves sampling frames from base and AR-applied videos, combining them into concatenated frames, and processing them with a hybrid vision encoder. Visual tokens are projected into a language model token space, reshaped, downsampled, and interleaved with text tokens. A fine-tuned large language model processes this input sequence to generate AR effect captions. Optical character recognition extracts text from AR frames, which is combined with generated captions and metadata to produce merged captions and content tags. This approach enables accurate description of temporal AR effects and facilitates downstream applications like search and ranking.
Owner:SNAP INC

Blockchain-Based Web3 Subdomain System for Tokenization, Certification, and Income Management of Real-World, Digital, and Identity Assets

The invention presents a blockchain-based system for the tokenization, certification, and management of real-world assets, digital assets, and verified identities. It integrates Web3 domain-subdomain tokens, distributed ledger technology (DLT), decentralized identifiers (DIDs), and AI-driven verification processes to provide secure and transparent transactions. The system employs KYC subdomain tokens linked to users' blockchain wallets for identity verification, enabling compliance with regulatory standards. Human-readable Web3 tokens enhance asset traceability. AI and manual verification ensure efficient and accurate processing. Scalable and flexible, the system supports diverse asset types and includes time-based voiding mechanisms to address incomplete transactions, ensuring effective lifecycle management.
Owner:SHEPLEY MATTHEW SMITH

Extended reality understanding through multimodal constrained decoding

Examples relate to processing extended reality (XR) content. A system obtains an unmodified image and a modified image with an XR effect applied. A trained multimodal generative language model generates visual difference text describing differences between the images. Additional text data associated with the XR effect is obtained. The additional text data can include visual text displayed by the XR effect and / or metadata associated with the XR effect. A trained generative language model processes the visual difference text and additional text data to generate output text data descriptive of the XR effect. The output text data may include content tags, location information, and a merged caption. Constrained decoding ensures the output adheres to a predefined structure. The system enables automated understanding and categorization of XR effects for applications like content discovery, recommendations, and moderation.
Owner:SNAP INC

Unattended wagon balance monitoring method and system

The invention discloses an unattended wagon balance monitoring method and system, and relates to the field of wagon balance monitoring, and the method specifically comprises the following steps: S1, constructing a data processing architecture; s2, triggering and receiving multi-source data; s3, carrying out data verification and credibility evaluation; s4, executing a graded early warning decision; and S5, realizing data archiving and model iteration. The unattended wagon balance monitoring method and the unattended wagon balance monitoring system provided by the invention can accurately identify behaviors influencing data authenticity, such as incomplete loading of a vehicle and cargo cheating, and guarantee the accuracy of commercial measurement data from the source. Meanwhile, for a localized dynamic correction mechanism of slightly abnormal data, a complete traceability mark is reserved while data validity is ensured, a traceable verification evidence chain is formed, a reliable basis is provided for solving commercial transaction disputes, legal rights and interests of two transaction parties are practically maintained, meanwhile, the manual intervention requirement can be greatly reduced, and the transaction efficiency is improved. Management personnel are liberated from tedious remote verification work.
Owner:SHENZHEN WASTE TO GOLD TECH CO LTD

Image generation with legible scene text

Systems and methods for generating images with legible scene text are described. Embodiments are configured to obtain a prompt describing a scene, where the prompt includes scene text indicating text that is intended to be shown in a generated image; encode, using a prompt encoder, the prompt to generate a prompt embedding; encode, using a character-level encoder, the scene text to generate a character-level embedding; and generate, using an image generation network, an image that includes the scene text based on the prompt embedding and the character-level embedding.
Owner:ADOBE INC

System and method for confirming the identity of an item based on item height

A device captures an image of a first item and generates a first encoded vector for the image. The device identifies a set of items that have at least one attribute in common with the first item. The device determines the identity of the first item based at least on attributes of the first item. The device determines that a confidence score associated with the identity of the first item is less than a threshold percentage. In response, the device determines a height of the first item. The device identifies item(s) with average heights within a threshold range from the height of the first item. The device compares the first encoded vector with a second encoded vector associated with a second item from the identified item(s). If the first encoded vector corresponds to the second encoded vector, the device determines that the first item corresponds to the second item.
Owner:7-ELEVEN INC

Guided content capture

A computing system can process incident information corresponding to a vehicle incident involving a vehicle of a user. The system can initiate a guided content capture process with a user to capture images of a vehicle of the user. The system generates a sequential set of vehicle outlines of the vehicle on an image interface presented on a computing device of the user to implement the guided content capture process. The system performs computer vision on image data captured by a camera of the computing device to determine when the vehicle is aligned with a particular vehicle outline of the sequential set of vehicle outlines.
Owner:ASSURED INSURANCE TECH INC

Multimodal multitask machine learning system for document intelligence tasks

Multimodal multitask machine learning system for document intelligence tasks includes a feature extractor processing token values obtained from a document to obtain features, and a token extraction head classifying, using the features, the token values to obtain classified tokens. The classified tokens are aggregated into entities. A document classification model is executed on the features to classify the document and obtain a document label prediction. Further a confidence head model applying the document label prediction processes the entities to obtain a result.
Owner:INTUIT INC

Intelligent Cascade Auto-Review System

A cascade auto-review system for automated classification and annotation of input is provided. An example system is structure adaptive and task oriented and includes a communication module configured to receive the input including images, videos, and metadata. The system further includes a plurality of subsystems. Each subsystem has a series of successive classifier stages configured to detect tags in the input and approve or reject the tags based on the images, the videos, and the metadata. The system further includes a database to store results of the classification and annotation. The results are used to train computer vision and machine learning algorithms.
Owner:BLINKFIRE ANALYTICS INC

Optical information reading device

To provide an optical information reading device capable of narrowing down to the character strings to be output from multiple character strings attached to a target object. The imaging unit 31 of the optical information reading device 10 captures an image of the target object 11 with multiple character strings attached as a symbol 20, and generates an input image containing the character strings. The character string processing unit of the optical information reading device 10 obtains multiple character string candidates from the input image. Then, the character string processing unit narrows down to the output character string from the multiple character string candidates based on at least one of the character string likelihood indicating the plausibility as a character string, the distance from the aiming position, and the degree of matching with a predetermined format pattern for each of the multiple character string candidates.
Owner:KEYENCE CORP

Material tray information input method, system and equipment and storage medium

The invention relates to a material tray information input method, system and device and a storage medium, and relates to the field of material trays. Each material disc comprises an information area engraved with material disc information and a slot area used for containing chips. The method comprises the steps that material disc images of a plurality of stacked material discs in an overlook state are collected; segmenting the material tray image to obtain an information image corresponding to the information area; segmenting the material tray image to obtain a chip total image corresponding to the slot position area; analyzing the information image to obtain material tray information; analyzing the chip total image to obtain chip information; and inputting material tray information and chip information, and respectively associating and storing the information image and the chip total image corresponding to the material tray information and the chip information. The technical effect of the invention is that the accuracy of information input of the material tray is improved.
Owner:WUXI V-TEST SEMICON CO LTD

Systems and methods for face annotation

Systems and methods for face annotation are described. One or more of the systems and methods include receiving a plurality of annotated images, wherein each annotated image of the annotated images comprises a caption; cropping the annotated image based on a face detection algorithm to obtain a face crop; comparing the face crop to the caption corresponding to the annotated image to obtain a caption similarity score; and filtering the plurality of annotated images based on the caption similarity score to obtain a plurality of annotated face images.
Owner:ADOBE INC

Spatially aligned string concatenation systems and methods for improved optical character recognition

A spatial alignment computer system for string alignment within a document processed using an optical character recognition (OCR) tool is provided. The computer system includes a processor in communication with a memory, wherein the processor is programmed to receive a plurality of bounding boxes of a document scanned using an OCR tool, identify a centroid of each bounding box of the plurality of bounding boxes, calculate coordinates for each centroid of each bounding box of the plurality of bounding boxes using a weighted Euclidean distance approach, sort the weighted Euclidean distance of the centroid of each bounding box in ascending order to obtain a sorting index, and based upon the sorting index, align one or more output strings associated with each bounding box of the plurality of bounding boxes.
Owner:STATE FARM MUTAL AUTOMOBILE INSURANCE COMPANY

Importing structured prescription records from a prescription label on a medication package

A system comprises one or more processors and one or more non-transitory computer-readable storage devices storing computing instructions configured to run on the one or more processors and cause the one or more processors to perform operations comprising: receiving a set of images of respective different portions of a label on a package. The operations also can include: determining, using contrast, one or more respective locations in the images of the set of images that are associated with the respective different portions of the label on the package. The operations further can include: providing, for display on a user interface, a flattened reconstruction of the respective different portions of the label on the package based at least in part on the one or more respective locations that have been determined. Other embodiments and features are also disclosed.
Owner:WALMART APOLLO LLC

A method for processing video (images) acquired from a camera linked to a computing device, and a system utilizing the same.

According to some embodiments of the present disclosure, a method for processing an image, executed by a computing device including a processor, may include: acquiring an image; acquiring analysis information corresponding to an object included in the image from the image using an object analysis model; and acquiring characters included in the object from the analysis information corresponding to the object using an OCR model. A representative diagram may be shown in FIG. 2.
Owner:イ·チュンヨル

Mobile supplementation, extraction, and analysis of health records

A method includes displaying a cohort report; receiving a request to determine a therapy identification; causing information to be transmitted to a remote cloud server; receiving clinical trial data; and updating the cohort report. A computing system includes a processor; and a memory having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: display a cohort report; receive a request to determine a therapy identification; cause information to be transmitted to a remote cloud server; receive clinical trial data; and update the cohort report. A computer-readable medium having stored thereon a set of computer-executable instructions that, when executed by one or more processors, cause a computer to: display a cohort report; receive a request to determine a therapy identification; cause information to be transmitted to a remote cloud server; receive clinical trial data; and update the cohort report.
Owner:TEMPUS AI INC

Systems and methods for handwriting recognition using optical character recognition

The present disclosure relates to systems, software, and computer-implemented methods for automatically identifying handwritten tips. An example method includes obtaining an image of a receipt, where the image includes a reference region in a pre-set format, a plurality of category identifiers, and at least one set of handwritten characters. The reference region and the plurality of category identifiers can be printed. The plurality of category identifiers can be located at pre-set positions relative to the reference region. The method further includes obtaining optical character recognition (OCR) information of the image and identifying the reference region from the OCR information based on the pre-set format of the reference region. The method further includes determining a plurality of amounts based on the OCR information and the reference region, where at least one of the plurality of amounts is associated with a tip category.
Owner:TU HARRY +1

Information processing apparatus, information processing method, and storage medium

An information processing apparatus recognizes multiple character strings in a scanned image of a form on which a piece of identification information for identifying a form issuer is written by performing character recognition processing on the scanned image, and inquires of an external system, in which multiple pieces of identification information are registered, whether the character string recognized from the piece of identification information among the recognized multiple character strings is registered in the external system. In the case where the character string is registered, the information processing apparatus displays information that the character string is registered. In the case where the character string is not registered, the information processing apparatus displays information that the character string is not registered and displays a similar character string similar to one of the multiple character strings as a correction candidate for the one character string.
Owner:CANON KK

Methods, systems, articles of manufacture and apparatus to label data

PendingEP4738281A1Character recognition
Systems, apparatus, articles of manufacture, and methods are disclosed to label data. An example apparatus includes interface circuitry, machine-readable instructions, and at least one processor circuit to be programmed by the machine-readable instructions to separate label data in a first data set from portions of an image, generate candidate labeled data based on associated ones of unlabeled portions of the image and optical character recognition (OCR) data, generate key performance indicator (KPI) metric values based on a comparison between the candidate labeled data and a second data set, and adjust weights of a model based on the KPI metric values.
Owner:NIELSEN CONSUMER LLC

An intelligent pathological section label identification method, system and medium

The application provides a kind of intelligent pathological section label identification method, system and medium, the method includes: receiving the input image of pathological section label, based on the identification mode of pre-set identification mode, the image format corresponding to input image and OCR type configuration string are identified;Call identification mode to execute identification operation to input image, analyze whether the identification result is valid;If the identification result is valid, then based on the identification mode of pre-set identification mode, the identification result is output;If the identification result is invalid, then switch identification mode, secondary identification is carried out to input image, until the identification result is valid, improve the identification efficiency of intelligent pathological section label;Wherein the priority strategy under automatic mode and intelligent back mechanism, combined with image enhancement technology, ensure the normal operation of system under the condition of multiple angles, low definition, information loss etc. Efficient identification of OCR and two-dimensional code is realized through multi-strategy joint identification mechanism.
Owner:HANGZHOU DEEP INFORMATICS TECH CO LTD